# MiMo-V2.6 Pro versus Flash: reported RL spending and results

**Result.** Xiaomi reports an estimated **$2.62 million** for Pro's RL run and **$0.85 million** for Flash's: **3.08×** as much, or **$1.77 million** more. Pro has substantial reported leads on the two selected exploit benchmarks and a smaller lead on Terminal Bench 4.0; Flash is close on several other tasks and scores higher on CyberGym. These are associations across two different released models, not an estimate of the return on extra RL spending. [Xiaomi release](https://mimo.mi.com/docs/en-US/news/latest/v2-6) and [Xiaomi model-card comparison](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL#3-evaluation-results), accessed **2026-09-29**.

## Cost and training volume

Xiaomi's September 22 release gives approximately $850,000 for Flash and $2.62 million for Pro, with 30 RL steps each. Its [live RL dashboard](https://mimo.xiaomi.com/rl/) was publicly reachable on **2026-09-29**, but its numeric panels are client-rendered and did not yield a stable extract in this inspection. The release is therefore the verifiable primary source for the final cost estimates. These estimates cover the **reported RL phase**, not pretraining, earlier research, evaluation, deployment, or total development cost. The ratio is 2.62/0.85 = **3.082**.

Both [Pro](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) and [Flash](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL) model cards state **1,568 prompts × 16 rollouts per step**. With 30 steps each, that gives **752,640 nominal rollout slots per model**, a **1.00× Pro/Flash nominal volume ratio**. This is configuration arithmetic, not a count of completed trajectories or processed tokens. The release's phrase “approximately 750,000 trajectories” is ambiguous about whether it refers to each run or both together; the batch arithmetic would imply about 1.51 million slots across both. A final, matched pair of RL token totals was not verifiable from the accessible primary pages, so no token ratio is reported. Sources accessed **2026-09-29**.

## Shared released-model evaluations

The eight rows below are from the **same Xiaomi model-card table**, using the same named benchmark and version within each row. Scores are on each benchmark's reported 0–100 scale; gaps are **Pro minus Flash in score points**. Scores should not be averaged across benchmarks. [Pro card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL#3-evaluation-results) and [Flash card](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL#3-evaluation-results), accessed **2026-09-29**.

| Benchmark | Pro | Flash | Gap |
|---|---:|---:|---:|
| DeepSWE v1.1 | 71.9 | 67.9 | +4.0 |
| ProgramBench | 26.5 | 26.0 | +0.5 |
| AutomationBench v1.0.6 | 53.1 | 52.3 | +0.8 |
| Terminal Bench 4.0 | 34.9 | 28.8 | +6.1 |
| Terminal Bench 2.1 | 89.9 | 87.6 | +2.3 |
| CyberGym | 94.0 | 95.1 | −1.1 |
| ExploitGym | 17.8 | 6.0 | +11.8 |
| ExploitBench | 47.9 | 25.3 | +22.6 |

**Interpretation.** Pro's largest observed gaps are on ExploitBench and ExploitGym. Its Terminal Bench 4.0 and DeepSWE leads are also visible, though much smaller. Flash comes within 0.5–0.8 points on ProgramBench and AutomationBench, within 2.3 points on Terminal Bench 2.1, and exceeds Pro by 1.1 points on CyberGym. The data do not provide repeated-run variability, so even small differences cannot be classified as statistically reliable. “Clear” here describes the magnitude of a reported score gap, not a significance test.

The [DeepSWE benchmark-hosted listing](https://huggingface.co/datasets/datacurve/deep-swe) repeats 71.9 and 67.9, and the [Terminal Bench 2.1 listing](https://huggingface.co/datasets/harborframework/terminal-bench-2.1) repeats 89.9 and 87.6 (accessed **2026-09-29**). These listings corroborate the **posted values**, but the available entries do not establish an independent, matched rerun of both models. Accordingly, all eight pairs remain Xiaomi-reported results for this analysis.

## What the comparison cannot establish

The [release](https://mimo.mi.com/docs/en-US/news/latest/v2-6) also reports live/checkpoint DeepSWE values of 72.6 for Pro and 65.7 for Flash, which differ from the released-model table; those values were excluded to avoid mixing evaluation snapshots. GDPval-AA 2.1 was excluded because the table has no Flash value. The model cards describe different backbones—**1.02T total/42B active parameters** for [Pro](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) versus **309B/15B** for [Flash](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL)—and do not fully specify every evaluation harness in the comparison table. Thus the score differences cannot be attributed to RL spending alone. Nor can these vendor-reported scores determine total development cost, real-world reliability, or a single overall value-for-money ranking. Sources accessed **2026-09-29**.

## Reproduction

The companion [cleaned data](mimo_v26_cleaned.csv) records each input, unit, evidence type, source URL, and access date. Run [the notebook](mimo_v26_analysis.ipynb) with Python 3 and Matplotlib to regenerate the CSV, ratios, gaps, [cost chart](rl_cost.png), and [gap chart](benchmark_gaps.png).