MiMo-V2.6: Which Way In Fits You, and What the “Cheap” Numbers Mean

Last updated: October 4, 2026

Evidence record: Starter Guide · Tier C — Document-based Research

Built from Xiaomi’s public docs, blog, X post, pricing pages, terms and privacy policies, its Hugging Face model cards and GitHub issues, the vLLM and SGLang deployment guides and Artificial Analysis’s model page. For this guide, Future Stack Reviews (FSR) did not sign in, call the API, run the Desktop app, buy a plan or host the weights.

Starter Guide is the format; Tier is depth of evidence, not a rating.

Read
Last read October 4, 2026Asia/Tokyo
Plan figures
US dollar list prices as shown signed out, from Japan, on October 4. Taxes and checkout weren’t checked.
Links
No affiliate or referral codes in this page’s links.

MiMo-V2.6 is the AI model family Xiaomi announced and open-sourced on September 22, 2026. There are five ways in: Xiaomi’s Desktop app, its pay-as-you-go API, its Token Plan subscription for coding tools, a third-party host, and the open weights. This Starter Guide covers what each requires, what the “cheap” numbers measure, and what Xiaomi said, and what FSR couldn’t find, about its September 25 API update for repeated tool calls. FSR used none of the five for this guide.

Verdict

No way in suits everyone, and comparing the five costs nothing: Xiaomi’s rates and terms are public. Highlighted: If you then try the API, start with a small pay-as-you-go top-up, not a subscription, and use the safeguards in section 08. Pro’s $0.0036 input price applies only to input that hits the cache; a cache miss costs about 121 times as much. Token Plan is for coding tools only and can’t be refunded, and the Desktop app is licensed for “non-commercial purposes,” which its agreement doesn’t define. Wait if you need an API model that can’t change under you, personal data that stays in the EU, or independent evidence that the repeated tool calls are fixed.

Best for, not for, and when to wait

Best for · Not for · Wait

Best for

If any applies

First-time MiMo-V2.6 buyers, personal or company

  1. Highlighted: developers who can prepay a small amount and test on public or made-up material
  2. heavy coding-tool users: at monthly list prices, only the Token Plan Pro and Max tiers can beat pay-as-you-go, past about 91% and 84% of their Credits on the Pro model (later still on Flash); Xiaomi’s time-limited discounts lower that share for every tier
  3. teams with server-class GPUs who can check the license and runtime themselves

Not for

If any applies

  1. The Desktop app in South Korea, the UK or the EU, where Xiaomi says it isn’t available yet
  2. The Desktop app for commercial use, until you know what “non-commercial purposes” covers
  3. Token Plan for scripts, app backends or other non-coding API use

Wait

Any one overrides Best for

  1. You need an API model that can’t change under you: FSR found no snapshot ID (a model name tied to a date or build) for the V2.6 models in Xiaomi’s API docs
  2. You need personal data to stay in the EU: the API privacy policy says such data may be transferred to Singapore
  3. You need independent evidence that tool-call repetition is fixed on your route

See sections 02, 07 and 08

Tier is depth of evidence, not a rating.

At a glance

Key facts · checked October 4, 2026

What it is
Three API models, a Desktop app and five open checkpoints (downloadable weights) on Hugging Face, two added after launch.
Launched September 22, 2026. Five ways in (section 02).
API list prices
Real-time API, per 1M tokens, cache-miss input and output: Flash $0.14 and $0.28, Pro $0.435 and $0.87, Pro UltraSpeed $4.35 and $8.70.
Xiaomi’s “overseas” (outside China) price table, prepaid in US dollars. Pro’s $0.0036 input price applies only to cache hits (section 05). Taxes weren’t checked.
Biggest conditions
Desktop: licensed for “non-commercial purposes” (undefined), and not yet available in South Korea, the UK or the EU. Token Plan: limited to coding tools.
Red note: No refunds for Token Plan or an activated Desktop membership (exceptions in section 08). Pay-as-you-go balances are refundable on request, with exclusions that leave the amount unclear.
September 25 update
Xiaomi says that on September 25 (06:00 UTC+8) it updated mimo-v2.6-pro and mimo-v2.6-flash on its API, under the same names, to curb repeated tool calls (the model requesting the same action again and again).
Red note: FSR found no snapshot ID to lock a V2.6 model version in Xiaomi’s API docs.

01 What MiMo-V2.6 is, in Xiaomi’s words

Xiaomi’s launch post describes two natively multimodal models, Pro and Flash. The API sells three model IDs, mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed, each listed with a 1M-token context window. Xiaomi pitches Pro at complex, long-running work, Flash at frequent, large-scale calls, and UltraSpeed as Pro “up to 20x faster,” a figure its release note gives without test conditions. The launch post also concedes a gap to the strongest closed-source models. These are Xiaomi’s claims, not FSR findings.

02 Five ways in, side by side

Your job is to pick a route or decide to wait. Six questions sort most readers: where are you, is the use commercial, would you put client or personal data in, do you need a model version that can’t change, do you use a supported coding tool, and do you have server-class GPUs? Mind the names: “Pro” is a model, a Token Plan tier and a Desktop tier, and the API’s model IDs differ from the checkpoint names (section 06).

The five ways in
Way inWhat you need firstThe catch to check
Desktop appWindows or an Apple-silicon Mac, and either a membership ($20 to $400 a month, tied to a Xiaomi account) or your own API keyNot in South Korea, the UK or the EU yet; licensed for “non-commercial purposes” (section 03)
Pay-as-you-go APIA Xiaomi account, a prepaid balance, an API key and an OpenAI- or Anthropic-format clientThe balance can fall below zero (section 04)
Token PlanA supported coding tool, plus the plan’s own key and endpoint addressCoding tools only; no refunds (section 04)
Third-party hostAn account with the host, on its termsWhich checkpoint it serves, and its data terms
Open weightsFor Pro or Flash, server-class GPUs on the published setupsPublished setups, for the RL checkpoints, list 208 GB of GPU memory for Flash and 680 GB for Pro (section 06)

Xiaomi’s model cards name OpenRouter as another place to use the models.

FSR read no host’s terms for this guide. Sources: Hugging Face, MiMo-V2.6-Pro-MOPD model card

03 The Desktop app: regions, plans and two clauses to read first

Xiaomi’s Desktop page pitches the app, for Windows and Apple-silicon Macs, at office work such as slides, documents, spreadsheets and scheduling. It says the app is not yet available in South Korea, the UK or EU member states, and names no countries where it is, so check from where you are.

Desktop membership on monthly billing, as shown October 4, 2026
PlanMonthly priceUsage quota
Starter$201×
Plus$606×
Pro$20025×
Ultra$40060×

The quota is a multiple, not an amount: the pricing page doesn’t say how much use 1× buys. Yearly billing is marked “Save 10%.” UltraSpeed is listed on Pro and Ultra only.

Xiaomi’s docs FAQ says yearly Desktop billing is 12% off, “subject to the actual display on the page,” and uses other tier names (Basic, Intermediate, Premium, Elite); this guide follows the pricing page. Sources: Xiaomi MiMo Desktop Membership Plans (Monthly and Yearly views) · Xiaomi MiMo Docs, FAQ, “Token Plan and Desktop Membership”

Highlighted: Before you pay, read two clauses, because an activated membership is non-refundable apart from listed exceptions (section 08). First, the Desktop Service Agreement licenses the software for “non-commercial purposes,” a phrase it doesn’t define. The same agreement counts “individuals or organizations” as users, and the membership agreement lets a member be “a single entity.” FSR draws no legal conclusion; ask Xiaomi or your own counsel what the clause covers before you subscribe for commercial use.

Second, the Desktop Privacy Policy says Xiaomi may use your inputs and outputs to train and improve its models, with an opt-out under Settings, General, Privacy, “Experience Optimization Plan.” It doesn’t state the default, so look before you type anything that matters. Your own API key doesn’t take you outside these two documents, which govern the app. The policy does say that for “a self-configured third-party large model” Xiaomi does “not actively collect or store” inputs or outputs, with listed exceptions, but not whether a MiMo API key counts as one.

That both documents still apply with your own key is FSR’s reading. Sources: Xiaomi MiMo Desktop Privacy Policy (last updated September 16, 2026), section 3.1 · Desktop Service Agreement

04 The API: pay-as-you-go or Token Plan

Pay-as-you-go is prepaid: you top up a balance, and each call draws it down at the listed rate.

Pay-as-you-go list prices per 1M tokens: real-time API, overseas table, as shown October 4, 2026
ModelInput: cache hit / cache missOutput
mimo-v2.6-flash$0.0028 / $0.14$0.28
mimo-v2.6-pro$0.0036 / $0.435$0.87
mimo-v2.6-pro-ultraspeed$0.036 / $4.35$8.70

Batch runs of Pro and Flash are listed at half these rates; web search is billed separately, at $5 per 1,000 calls. The payment FAQ says billing lags, so the balance can go negative; calls then stop, and the next top-up first covers the shortfall. A top-up is not a hard spending cap.

The last sentence is FSR’s inference. Sources: Xiaomi MiMo Docs, “API Pricing” · FAQ, “Payment”

Token Plan is a subscription with a monthly Credit allowance. On the Pro model, a cache-miss input token uses 300 Credits, an output token 600 and a cache-hit input token 2.5. Xiaomi limits the plan to “programming tools” such as OpenClaw and OpenCode and prohibits API calls in “obvious non-coding scenarios” such as automated scripts and custom application backends; violations can bring suspension and a blocked key. Yet Xiaomi’s FAQ also lists the Desktop app among the tools a Token Plan key runs, without saying whether non-coding work there breaks the rule.

Token Plan, Individual, monthly list prices, before discounts
PlanPrice and monthly CreditsSame Credits on pay-as-you-goShare to use before the plan wins
Lite$6, 4.1 billion$5.945100.9%
Standard$16, 11 billion$15.95100.3%
Pro$50, 38 billion$55.1090.7%
Max$100, 82 billion$118.9084.1%

The last two columns are FSR’s arithmetic for the Pro model at real-time list prices: the full quota’s cost on pay-as-you-go, and the plan’s price as a share of it. Above 100%, pay-as-you-go is cheaper even if you spend every Credit. On the Flash model the shares run higher: 104.5%, 103.9%, 94.0% and 87.1%. Xiaomi lists three discounts, which it calls time-limited, and each lowers the share for every tier. A first Individual monthly purchase, or an annual plan, is 12% cheaper, which puts the Pro-model shares between 88.8% (Lite) and 74.0% (Max). Credits are also spent at 0.8 times the usual rate from 16:00 to 24:00 UTC; use confined to those hours would put the list-price shares between 80.7% and 67.3%. Either way, a plan wins only if you use most of its quota.

FSR’s arithmetic from Xiaomi’s prices and Credit rates, not a measurement of anyone’s usage. Cache-miss input (300 Credits a token) and output (600) both work out to $1.45 per billion Credits; at the cache-hit rate (2.5 Credits a token, $1.44 per billion) the shares come out up to 0.7 points higher. Flash’s rates (100, 200 and 2 Credits a token) all work out to $1.40 per billion. With 12% off, the four Pro-model shares are 88.8%, 88.3%, 79.9% and 74.0%; with all use in the off-peak hours, 80.7%, 80.3%, 72.6% and 67.3%. Annual plans are listed at twelve monthly payments less 12%, for twelve times the monthly Credits. Sources: Xiaomi MiMo Docs, “Token Plan” · “API Pricing” · FAQ, “Plans and Pricing” · FAQ, “Usage and Quota”

A paid plan can’t be refunded, and when the quota runs out, service stops instead of drawing on your pay-as-you-go balance. The shares assume a plan runs to expiry, after which Xiaomi’s FAQ says unused Credits don’t carry over. Renewing or upgrading a current Individual plan before then turns unused Credits into a “remaining value” that offsets part of the new payment; the table leaves that out, and FSR hasn’t seen it applied in an account. UltraSpeed isn’t supported for now. Team plans cost the same per seat, need at least two seats, and have no Lite tier or first-purchase discount.

05 What the “cheap” numbers actually measure

Four numbers make MiMo-V2.6 look cheap. Highlighted: Each measures something specific, and none is the cost of your own task.
Four numbers and what they measure
The numberWhat it measuresWhat it doesn’t tell you
$0.0036 per 1M input tokens (Pro)The cache-hit price, about 1/121 of the $0.435 cache-miss priceHow much of your input will qualify
$0.13 per task (Artificial Analysis, October 4)That site’s weighted average cost of running its own Intelligence Index tasks on ProYour task’s cost, or which model version was tested
4.1 to 82 billion Credits a monthA prepaid Token Plan quota that expiresWhether you save (section 04)
Open weights with an MIT tagPublic downloads tagged MITThe hardware bill: 208 GB of GPU memory for Flash, 680 GB for Pro on published setups (section 06)

Sources by row. Cache-hit price: Xiaomi MiMo Docs, “API Pricing”, FSR’s arithmetic (0.435 ÷ 0.0036 = 120.8) · Cost per task: Artificial Analysis, “MiMo-V2.6-Pro”, where FSR found no test date and no tested model version · Credits: “Token Plan”, FAQ, “Validity and Expiry” · Open weights: Hugging Face, MiMo-V2.6-Pro-MOPD model card, vLLM Recipes, MiMo-V2.6-Flash-RL, MiMo-V2.6-Pro-RL

For the first, work out your own blend: effective input price = h × cache-hit price + (1 − h) × cache-miss price, where h is the share of input tokens billed at the cache-hit price. For Pro, h = 0 gives $0.435 per 1M input tokens, h = 0.5 gives $0.219, and even h = 0.925 gives $0.036, ten times the cache-hit price. The pricing page says only that a cache hit is when “the requested prefix content hits the Prompt Cache,” and lists cache writes as “Limited-time Free” with no end date; a keyword search of Xiaomi’s docs found nothing on how long a cache lasts or how much must match. The Chat Completions format’s usage fields report cached tokens and reasoning tokens (the model’s thinking), so you can measure yours. The pricing page lists no separate rate for reasoning tokens.

FSR’s arithmetic from Xiaomi’s list prices, not a measurement. Sources: Xiaomi MiMo Docs, “API Pricing” · “OpenAI Chat Completions API Compatibility” · full-text file

For the second, mind the class. On October 4, Artificial Analysis ranked Pro 16th-cheapest of 118 on cost per task and first of 118 on its Intelligence Index, but the site says it compares an open-weights model only with other open-weights models of the same size class. First of 118 is not first among all models: the site’s wider chart shows closed models scoring higher.

06 Open weights: what running them takes

Xiaomi’s Hugging Face collection holds five checkpoints: Pro-RL, Flash-RL, Pro-MOPD, Flash-MOPD and Distill-Qwen-9B. The cards put Pro at 1.02 trillion total parameters and Flash at 309 billion. RL (reinforcement learning) and MOPD (a later distillation stage) are training stages. Each MOPD card calls itself “the MOPD upgrade” of the RL checkpoint and says it “mitigates” the tool-call repetition covered in section 07. FSR saw no UltraSpeed weights there.

The vLLM project’s recipes, written for the RL checkpoints, list Pro at eight H200 GPUs, four MI355X GPUs, or an equivalent total of at least 680 GB of GPU memory, and Flash at four H200s or at least 208 GB. These are published configurations, not FSR measurements and not minimums for every setup. vLLM had no recipe for the MOPD checkpoints on October 4, and the MOPD cards state no memory requirement.

SGLang’s cookbook lists the same GPU counts; a vLLM pull request adding MOPD recipes was open on October 4. Sources: vLLM Recipes, MiMo-V2.6-Pro-RL · MiMo-V2.6-Flash-RL · pull request 1059 · SGLang cookbook, MiMo-V2.6 · Hugging Face, MiMo-V2.6-Pro-MOPD model card · MiMo-V2.6-Flash-MOPD model card

Distill-Qwen-9B is not a small Pro: its card describes a 9B fine-tune of Qwen3.5-9B on MiMo-generated data.

All five repositories carry an MIT tag, and FSR found no separate LICENSE file in them. Confirm the terms before you build a product on the weights.

The last sentence is FSR’s advice. Sources: Hugging Face, MiMo-V2.6 collection and the five repositories it lists

07 The September 25 update: what Xiaomi said and what FSR couldn’t find

After launch, Xiaomi’s model card says, the model would sometimes repeat the same or nearly the same tool call (a request to run a tool, such as a search), using up time and context without making progress; on a metered route, those are tokens you pay for. Xiaomi’s September 27 blog post says updated models have been on its API platform since September 25 at 06:00 UTC+8, or 22:00 UTC on September 24. <mark class=”fsr-mg-hi”><span class=”fsr-mg-sr”>Highlighted: </span>It adds: “The model names remain unchanged: mimo-v2.6-pro and mimo-v2.6-flash.”</mark> The post also announces the MOPD checkpoints and, as an apology, a quota reset for Desktop users.

The gloss on tool calls and the remark on cost are FSR’s. Sources: Hugging Face, MiMo-V2.6-Pro-MOPD model card · Xiaomi MiMo blog, “Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6” · FSR’s time conversion

Xiaomi’s wording on the effect varies by channel: its X post says “diagnosed & fixed,” its model card “mitigates,” and its blog that repetition “dropped substantially.” The blog’s metric counts exact repeats within one turn, which the blog itself calls “only a lower bound” on the repetition a user could observe.

In the blog, the X post, the four Pro and Flash model cards and a keyword search of Xiaomi’s full docs, FSR did not find:

  • a statement that the API models are the same weights as the public MOPD checkpoints (the blog announces both in one passage, and all four cards, RL and MOPD alike, list Xiaomi’s API as another place to use the model; neither says which weights it serves)
  • whether UltraSpeed, the Token Plan endpoints, the Desktop app or third-party hosts got the update
  • a snapshot ID to lock either model’s version, or a September 25 entry in the release notes
  • any compensation over the repetition for API or Token Plan users

Self-hosting pins the weights you run, though nothing FSR read says they match the API.

User reports settle little. FSR read in full, with comments, the four public issues opened on Xiaomi’s GitHub repository between the update and October 4: none reports the original repetition, though one describes a self-hosted Pro-MOPD refusing a valid call. In a September 28 comment on an earlier issue, a user says a third-party host was serving Pro-RL, the pre-MOPD checkpoint. That shows neither that the fix worked nor how common any fault is.

User reports, not FSR findings; FSR didn’t verify what any host serves. Sources: GitHub, XiaomiMiMo/MiMo, issues 100, 101, 102 and 103, and issue 98 (comment of September 28)

08 Before you commit: data, the way out and a first run

Training use differs by route. The API privacy policy says Xiaomi won’t use the content you provide for model training; the Desktop policy says it may, unless you opt out (section 03). Not used for training doesn’t mean not stored: both policies describe collecting what you submit (section 03 has the Desktop exception) and keep personal information “for the period necessary,” naming no period. Both say web search passes your query and IP address to Google.

Where data sits: the API privacy policy says personal data of EEA, UK and Swiss residents is stored in the Netherlands and may be transferred to Singapore under legal transfer mechanisms, and everyone else’s is stored in Singapore; the API agreement says you may ask Xiaomi to select storage regions, possibly for a fee. If data must stay in the EU, ask before you send any. The agreement also lets Xiaomi change or interrupt the service at any time without notice.

The way out, by route
RouteRefundHow it stops
Pay-as-you-goThe unused balance, on request, except invoiced or gifted amounts (see below). Once a refund is accepted, you can’t call the models or top upDelete the key. Batch jobs already running continue and are billed
Token PlanNone: “non-refundable and non-cancellable upon purchase,” in the API agreement’s wordsCancel auto-renewal on the Token Plan page; service ends when the plan expires
Desktop membershipOnce activated, “non-refundable,” except where a Xiaomi breach makes the service completely unusable, the law requires a refund, Xiaomi agrees, or the agreement says otherwise; a refund covers only the unused part“Cancel auto-renewal” under “Subscription & Invoices”

Menu names as Xiaomi’s pages give them; FSR didn’t open them in an account. The API agreement ends pay-as-you-go access earlier than the FAQ does: once a refund is initiated, not once it is accepted. Sources by row. Pay-as-you-go: Xiaomi MiMo Docs, FAQ, “Payment”, Xiaomi MiMo Service Agreement (API), section 3.12 · Token Plan: Xiaomi MiMo Service Agreement (API), section 3.13, FAQ, “Payment”, FAQ, “Plans and Pricing”, FAQ, “Validity and Expiry” · Desktop membership: Xiaomi MiMo Desktop Membership Service Agreement, section 4.7, Desktop Membership Plans (FAQ)

Don’t count on that first row. The payment FAQ excludes invoiced amounts from refunds yet says each overseas top-up is invoiced automatically, and the account FAQ, an older page, says the overseas version has no invoicing function for now. FSR can’t tell from these pages how much is refundable or what invoice you’d get: top up only what you’re prepared to spend, and ask Xiaomi first if you need invoices.

For developers who want one step more, FSR proposes a small evaluation it has not run. Use a client that lets you cap tool calls and output tokens; without one, stay with the documents. Top up a small amount (FSR found no minimum in Xiaomi’s docs), enable the balance alert Xiaomi’s FAQ describes, and run one small real-time job on public or made-up material, with a tool that only reads. Read each response’s usage figures and put your cache-hit share into section 05’s formula. Stop on repeated identical tool calls, a repetition_truncation stop reason (the docs’ signal that the model detected repetition), repeated 429 errors, or spending beyond what you planned; then delete the key and wait, or try another route. A clean run shows only that this job didn’t trigger the problem.

An evaluation FSR proposes, not a test FSR ran. The caps and the read-only tool are controls to set in your own client; FSR didn’t check what limits MiMo’s API enforces. Use a real-time job because Xiaomi’s FAQ says deleting a key doesn’t stop batch jobs already running. FSR found repetition_truncation on the Chat Completions and Anthropic-format reference pages, not on the Responses page; error 429 covers both rate limits and an exhausted Token Plan quota. Sources: Xiaomi MiMo Docs, FAQ, “Payment” · “OpenAI Chat Completions API Compatibility” · “Anthropic Messages API Compatibility” · “OpenAI Responses API Compatibility” · “Error Codes”

09 FAQ

Is the API model the same as the open MOPD weights?

Xiaomi hasn’t said so in anything FSR read. Its blog announces the updated API models and the open MOPD checkpoints together, and the model cards, RL ones included, list the API as another place to use the model. Neither says which weights the API serves (section 07).

Xiaomi MiMo blog, “Diagnosing and Mitigating Tool-Call Repetition in MiMo-V2.6” · Hugging Face, MiMo-V2.6-Pro-MOPD model card · MiMo-V2.6-Pro-RL model card

Does Xiaomi train on what I type?

On the API, its privacy policy says no. In the Desktop app, its policy says Xiaomi may, unless you opt out, and doesn’t state the default (sections 03 and 08).

Xiaomi MiMo Privacy Policy (API), section 3.1 · Xiaomi MiMo Desktop Privacy Policy, section 3.1

Can I run MiMo-V2.6 on my own PC?

Not Pro or Flash on the setups FSR read, which list 208 GB of GPU memory for Flash and 680 GB for Pro; FSR read no guide to smaller builds. Distill-Qwen-9B is far smaller, but a different model (section 06).

vLLM Recipes, MiMo-V2.6-Flash-RL · MiMo-V2.6-Pro-RL · Hugging Face, MiMo-V2.6-Distill-Qwen-9B model card

10 Methodology and limits

Starter Guide, Tier C. For this guide, FSR did not sign in to any Xiaomi service, call the API, run the Desktop app, buy a plan or host the weights, so nothing here measures quality, speed, cost or how often tool calls repeat. The pay-as-you-go equivalents and shares in section 04 and the blended prices in section 05 are FSR’s arithmetic from Xiaomi’s published prices; section 08’s evaluation is a proposal.

FSR read the pages cited here on October 3 and 4, 2026 (Asia/Tokyo), signed out and from Japan; figures are as last read on October 4. The terms are the English versions served there; FSR didn’t check what other regions are shown. FSR did not read Xiaomi’s technical report, any host’s terms, any signed-in screen or any checkout. “Not found” covers only the pages and passages FSR read. Artificial Analysis’s figures move; its class count changed during October 3.

FSR researched, drafted and reviewed this guide with AI assistants, including Anthropic’s Claude and an OpenAI assistant, and checked their findings against the pages cited; agreement between AI tools is not independent verification. This page’s links carry no affiliate or referral codes.

FSR’s own disclosure. See also: FSR, Methodology · FSR, Disclosure

11 Verdict

Highlighted: Choose the route by your own conditions, not the headline prices, and compare Xiaomi’s published rates before you pay anything. If you then try the API, start with a small pay-as-you-go top-up, not a subscription, and use the safeguards in section 08.
  • Pay-as-you-go API: your cost depends on how much input hits the cache.
  • Token Plan: only for heavy coding-tool use. At monthly list prices, only the Pro and Max tiers can beat pay-as-you-go, past about 91% and 84% of their Credits on the Pro model (later still on Flash); Xiaomi’s time-limited discounts lower that share for every tier (section 04). No tier can be refunded.
  • Desktop app: only if it is offered where you live, and keep commercial use out of it until you know what “non-commercial purposes” covers.
  • Third-party host: first check which checkpoint it serves and its data terms.
  • Open weights: the published setups for Pro and Flash call for server-class GPUs.

Wait if you need an API model that can’t change under you (self-hosting pins the weights, if you have the hardware), personal data that stays in the EU, or independent evidence that the repeated tool calls are fixed.

In MiMo-V2.6’s favor: the prices are published in enough detail to do this arithmetic before you pay, Xiaomi published its own diagnosis of the repetition problem, and the weights are public and tagged MIT.

This verdict would change if Xiaomi said how its API models map to the public checkpoints, offered a model version you can pin, or defined “non-commercial purposes”: each would replace an unknown with something you can check. Solid evidence of how the updated API behaves, good or bad, would change the advice to wait. FSR used none of the five routes, so none of this says how good the models are.

Where you would start

First check
The six questions in section 02. For a first run, section 08 has a small evaluation that FSR proposes and hasn’t run.
Desktop app
Xiaomi MiMo Desktop page (opens in a new tab). Read the two clauses first (section 03).
API
Prepaid in US dollars. Token Plan is for coding tools only (section 04).
Open Xiaomi’s API pricing pagemimo.mi.com/docs/en-US/price/pay-as-you-go (opens in a new tab)

Official Xiaomi and Hugging Face links, with no affiliate or referral code.

12 Corrections and contact

Corrections and contact

When
A price, quota, region or term here is out of date, or Xiaomi publishes how its API models map to the open checkpoints.
Send
The public source URL, the passage it affects and its date, if shown.
Do not send
API keys, account details, invoices, or anything personal or confidential.
Send through the FSR contact pagefuture-stack-reviews.com/contact (opens in a new tab)

Verified errors are fixed or clearly updated, not quietly removed.

Ask whether it fits you

Each button copies the briefing below and opens that assistant in a new tab; paste it there. Don’t assume the assistant knows where you are, what you’d run or any employer’s rules. Leave out API keys, account details and client or confidential data. If nothing is copied, open “What gets copied” and copy it yourself.

What gets copied
FSR BRIEFING PACKET - MiMo-V2.6: Which Way In Fits You, and What the "Cheap" Numbers Mean
Publisher: Future Stack Reviews. Article: https://future-stack-reviews.com/xiaomi-mimo-v2-6-starter-tierc/
Format: Starter Guide. Evidence depth: Tier C, built from Xiaomi's public docs, blog, X post, pricing pages, terms and privacy policies, its Hugging Face model cards and GitHub issues, the vLLM and SGLang deployment guides and Artificial Analysis's model page; for this guide FSR did not sign in, call the API, run the Desktop app, buy a plan or host the weights.
Read: October 3 and 4, 2026 (Asia/Tokyo); figures as last read on October 4.
The article's links carry no affiliate or referral codes. FSR used AI assistants, including Anthropic's Claude and an OpenAI assistant, to research, draft and review the article.

WHAT IT IS
MiMo-V2.6 is the AI model family Xiaomi announced and open-sourced on September 22, 2026: three API models (mimo-v2.6-pro, mimo-v2.6-flash and mimo-v2.6-pro-ultraspeed), a Desktop app, and open checkpoints (downloadable weights) on Hugging Face. Three checkpoints were there at launch (Pro-RL, Flash-RL and Distill-Qwen-9B); Pro-MOPD and Flash-MOPD were added on September 27.

FIVE WAYS IN (public pages, read October 4, 2026)
- Desktop app: Windows or an Apple-silicon Mac. Xiaomi's page says it is not yet available in South Korea, the UK or EU member states, and lists no countries where it is. Membership is $20, $60, $200 or $400 a month for a 1x, 6x, 25x or 60x usage quota; the pricing page does not say how much use 1x buys. Once activated, membership is non-refundable, with listed exceptions. The service agreement licenses the app for "non-commercial purposes" and does not define the phrase. The Desktop privacy policy says inputs and outputs may be used for training unless you opt out, and does not state the default.
- Pay-as-you-go API: prepaid in US dollars. Real-time list prices per 1M tokens (cache-hit input / cache-miss input / output): Flash $0.0028 / $0.14 / $0.28; Pro $0.0036 / $0.435 / $0.87; Pro UltraSpeed $0.036 / $4.35 / $8.70. Batch runs of Pro and Flash are listed at half those rates. The payment FAQ says the balance can go negative because billing lags. Refunds of an unused balance exclude invoiced amounts, and the same FAQ says each overseas top-up is invoiced automatically, so do not count on one. The API privacy policy says Xiaomi will not use submitted content for model training, and that personal data of EEA, UK and Swiss residents is stored in the Netherlands and may be transferred to Singapore. The API agreement says data is stored on servers in Europe and Singapore and that you may ask Xiaomi to select storage regions, possibly for a fee.
- Token Plan (Individual, monthly): Lite $6 for 4.1 billion Credits, Standard $16 for 11 billion, Pro $50 for 38 billion, Max $100 for 82 billion. Team plans cost the same per seat (no Lite) and need at least two seats. Xiaomi limits it to programming tools; non-coding API use can get the subscription suspended and the key blocked, yet its FAQ also lists the Desktop app among the tools a Token Plan key runs. No refunds. Unused Credits do not carry over once a plan expires. Before expiry, Xiaomi's FAQ says, renewing or upgrading a current Individual plan offsets the new payment by the unused Credits' "remaining value"; FSR's shares below leave that offset out, and FSR has not seen it applied in an account.
- Third-party host: Xiaomi's model cards name OpenRouter. A host's own prices and terms apply; FSR read none of them. Check which checkpoint the host serves.
- Open weights: vLLM's recipes, written for the RL checkpoints, list four H200 GPUs or 208 GB of GPU memory for Flash and eight H200 GPUs or 680 GB for Pro; the model cards put Pro at 1.02 trillion total parameters and Flash at 309 billion. The repositories carry an MIT license tag and no separate LICENSE file. Distill-Qwen-9B is a different, 9B model.

WHAT THE "CHEAP" NUMBERS MEASURE
- $0.0036 per 1M input tokens is Pro's cache-hit price, about 1/121 of its cache-miss price. Effective input price = h x cache-hit price + (1 - h) x cache-miss price, where h is the share of input tokens billed at the cache-hit price. For Pro, h = 0 gives $0.435, h = 0.5 gives $0.219 and h = 0.925 gives $0.036 (FSR's arithmetic, not a measurement).
- $0.13 per task is Artificial Analysis's cost of running its own Intelligence Index tasks on Pro. Its ranks are 16th-cheapest of 118 on cost per task and first of 118 on the index; by the site's own note, it compares an open-weights model only with other open-weights models of the same size class.
- Token Plan on the Pro model, at monthly list prices and before discounts (FSR's arithmetic, assuming unused Credits are lost at expiry): Lite and Standard cost slightly more than their full quota would on pay-as-you-go (100.9% and 100.3%), so they do not beat it; the Pro tier beats it only past 90.7% of its Credits and Max past 84.1%. On the Flash model the four shares are higher: 104.5%, 103.9%, 94.0% and 87.1%. Xiaomi lists three discounts, which it calls time-limited, and each lowers the share for every tier: 12% off a first Individual monthly purchase or an annual plan (Pro-model shares 88.8%, 88.3%, 79.9% and 74.0%), and Credits spent at 0.8 times the usual rate from 16:00 to 24:00 UTC (80.7%, 80.3%, 72.6% and 67.3% if all use falls in those hours). Either way, a plan wins only if you use most of its quota.

THE SEPTEMBER 25 UPDATE
Xiaomi's blog says updated models have been on its API since September 25, 2026, 06:00 UTC+8, under unchanged names (mimo-v2.6-pro and mimo-v2.6-flash), and that it open-sourced MOPD checkpoints that address tool-call repetition. Its wording on the effect ranges from "diagnosed & fixed" (X post) to "mitigates" (model card), and the blog calls its own metric only a lower bound on observable repetition.

WHAT IS NOT SETTLED
- Whether the API models are the same weights as the public MOPD checkpoints. Xiaomi's blog announces both in one passage, and the model cards, RL and MOPD alike, list the API as another place to use the model; neither says which weights the API serves.
- Whether UltraSpeed, the Token Plan endpoints, the Desktop app or third-party hosts received the update.
- Any way to lock a model version on the API: FSR found no snapshot ID (a model name tied to a date or build) for the V2.6 models in Xiaomi's API docs.
- What "non-commercial purposes" covers for the Desktop app, and the default of its training setting.
- Whether non-coding work in the Desktop app on a Token Plan key breaks the plan's coding-only rule.
- How much of a pay-as-you-go balance is refundable for overseas users.
- How much use the Desktop quota buys; taxes and checkout; your own cache-hit rate and cost per task.

IF YOU TRY THE API (section 08 of the article; FSR's proposal, not a test FSR ran)
Comparing the routes needs no payment. If you then try the API, use a client that lets you cap tool calls and output tokens; without one, stay with the documents. Top up a small amount (FSR found no minimum in Xiaomi's docs), turn on the balance alert, run one small real-time job on public or made-up material with a tool that only reads, and read each response's usage figures. Stop on repeated identical tool calls, a repetition_truncation stop reason, repeated 429 errors, or spending beyond what you planned; then delete the key and wait, or try another route. A clean run shows only that this job did not trigger the problem.

FSR VERDICT
Choose the route by your own conditions, not the headline prices, and compare Xiaomi's published rates before you pay anything. If you then try the API, start with a small pay-as-you-go top-up, not a subscription, and use the safeguards in section 08.
- Pay-as-you-go API: your cost depends on how much input hits the cache.
- Token Plan: only for heavy coding-tool use. At monthly list prices, only the Pro and Max tiers can beat pay-as-you-go, past about 91% and 84% of their Credits on the Pro model (later still on Flash); Xiaomi's time-limited discounts lower that share for every tier (section 04). No tier can be refunded.
- Desktop app: only if it is offered where you live, and keep commercial use out of it until you know what "non-commercial purposes" covers.
- Third-party host: first check which checkpoint it serves and its data terms.
- Open weights: the published setups for Pro and Flash call for server-class GPUs.
Wait if you need an API model that can't change under you (self-hosting pins the weights, if you have the hardware), personal data that stays in the EU, or independent evidence that the repeated tool calls are fixed.

WHAT WOULD CHANGE IT
Xiaomi saying how its API models map to the public checkpoints, offering a model version you can pin, or defining "non-commercial purposes"; or solid evidence, good or bad, of how the updated API behaves.

SOURCES
- API Pricing: https://mimo.mi.com/docs/en-US/price/pay-as-you-go
- Payment FAQ: https://mimo.mi.com/docs/en-US/quick-start/faq/payment
- Token Plan: https://mimo.mi.com/docs/en-US/price/token-plan
- API Service Agreement: https://mimo.mi.com/docs/quick-start/terms/user-agreement
- API Privacy Policy: https://privacy.mi.com/XiaomiMiMoPlatformos/en_GB/
- Desktop page: https://mimo-ai.xiaomimimo.com/desktop/
- Desktop Membership Plans: https://mimo-ai.xiaomimimo.com/pricing/
- Desktop Membership Service Agreement: https://mimo-ai.xiaomimimo.com/legal/membership-service-terms
- Desktop Service Agreement: https://mimo.xiaomi.com/legal/service-agreement
- Desktop Privacy Policy: https://mimo.xiaomi.com/legal/privacy-policy
- Blog on the update: https://mimo.xiaomi.com/blog/mimo-v2-6-tool-call-repetition
- Hugging Face collection: https://huggingface.co/collections/XiaomiMiMo/mimo-v26
- vLLM recipes: https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.6-Pro-RL and https://recipes.vllm.ai/XiaomiMiMo/MiMo-V2.6-Flash-RL
- Artificial Analysis: https://artificialanalysis.ai/models/mimo-v2-6-pro

LIMITS
Prices, plans, terms and rankings change; verify on the official pages before you act. FSR used none of these routes for this guide, so nothing here measures quality, speed or cost, or shows what a particular account can do. Where this packet and the live pages disagree, the live pages win.

Help me decide which way into MiMo-V2.6 fits me, or whether to wait. Open the article linked above if you can; if you can't, say so and work from this briefing. Ask where I am, whether the use is commercial, whether I would put client or personal data in and whether it must stay in the EU, whether I need a model version that cannot change, whether I already use a supported coding tool, and whether I have server-class GPUs. Do not ask me to paste API keys, account details or confidential data.

Your assistant’s reply is its own answer, not FSR’s. Check it against the article and the Xiaomi pages it cites.