# MiMo V2.6 API: Flash, Pro, UltraSpeed and the V2.5 migration

> MiMo V2.6 uses separate model and price IDs for Flash, Pro and UltraSpeed. Migrate with a complete tool-loop and usage check.

_Published 2026-10-03 · Updated 2026-10-03 · [KeepRouter Editorial](https://keeprouter.com/editorial-policy#editorial-team) · 7 minute read_

![Cache pricing audit diagram: a short green bar beside a tall accent bar, measured against the same baseline, with a ledger grid on the right](https://keeprouter.com/editorial/blog/llm-cache-pricing-audit.png)

_Compare the complete request cost with and without cache reuse. Illustration, not measured savings._

**Use an exact MiMo V2.6 model ID with KeepRouter Chat Completions, then verify your tool history and billed usage.** Flash, Pro and Pro UltraSpeed are separately priced choices. Xiaomi's pricing notice schedules MiMo V2.5 and V2.5 Pro retirement for October 21, 2026 at 10:00 Beijing time, so an integration still using those IDs needs a planned migration.

This guide's source review is dated October 3, 2026. It separates maker specifications, KeepRouter prices and evaluation recommendations. It does not infer production latency from a variant's name.

## Which MiMo ID should you evaluate?

| Exact ID | Evaluation to start with | Decision evidence |
|---|---|---|
| [mimo-v2.6-flash](/models/mimo-v2.6-flash) | Short extraction or support classification | Correct fields and total spend |
| [mimo-v2.6-pro](/models/mimo-v2.6-pro) | A multi-step repository or document task | Accepted result after tools |
| [mimo-v2.6-pro-ultraspeed](/models/mimo-v2.6-pro-ultraspeed) | The same Pro task under a latency target | Measured latency and price premium |

These are suggested starting tasks, not a ranking. Xiaomi documents multimodal understanding and a one-million-token context for the V2.6 family. That does not turn these IDs into audio-generation endpoints. If your product needs spoken output, choose a separately documented speech model and its speech route.

## Make a bounded text request

```python
import os
from openai import OpenAI
client = OpenAI(base_url="https://keeprouter.com/v1",
                api_key=os.environ["KEEPROUTER_KEY"], max_retries=0)
r = client.chat.completions.create(
    model="mimo-v2.6-flash", max_tokens=1024,
    messages=[{"role": "user", "content": "Classify: 'I cannot reset my password'. Reply with one category."}])
print(r.choices[0].message.content)
print(r.usage)
```

Use a server-side key and inspect the finish reason. A request that stops during reasoning needs a budget review rather than an endless retry. Keep the first test small so input, output and request cost can be matched to the ledger.

For an application that processes images or audio, test the documented input encoding separately. Do not infer media compatibility from a text success. Store format-specific validation errors so a bad asset is not reported as a model outage.

## The migration risk inside a tool loop

MiMo's deep-thinking guidance requires preserving `reasoning_content` in tool-call history. An agent framework that converts every assistant response to `role` and `content` only can drop that state. Inspect the framework's serialized second request, not just its first response.

Use a harmless `read_config` function returning one setting. Save the complete assistant message with tool calls and reasoning, append the matching tool result, and ask the model to state the setting. A passing sequence must preserve the tool ID and finish with the returned value. Add a missing-setting case to check that the answer reports uncertainty instead of inventing configuration.

Your logs can record whether a reasoning field was retained without storing its text. This makes it easier to catch an adapter regression while keeping unnecessary payloads out of analytics.

## Compare prices without mixing regions or units

The model pages publish KeepRouter customer rates in USD per million input, output and cached-input tokens. Xiaomi also publishes native-currency and international rates. Those are useful source references, not amounts to copy directly into a KeepRouter invoice. Currency reference rates can differ from settlement rates, and a quota-plan entitlement is not the same as a pay-as-you-go token price.

For a task with 20,000 input tokens, 5,000 cached tokens and 1,000 output tokens, the worksheet quantities are 0.015 million ordinary input, 0.005 million cached input and 0.001 million output. Multiply by the current page rates and add every billable retry. These are hypothetical quantities, not measured MiMo usage or a promised cache ratio. Use the [calculator](/tools/api-cost-calculator) to change them.

UltraSpeed is a separate serving variant. Compare it against ordinary Pro with identical payloads, output budgets and acceptance rules. Report median and tail latency only after collecting enough application samples, and record the number of observations. One quick response cannot establish a sustained speed advantage.

## Migrate before the retirement deadline

Inventory production code, environment variables, agent settings and queued jobs for `mimo-v2.5` and `mimo-v2.5-pro`. Select one V2.6 candidate and replay representative tasks without changing tool definitions. Check output schemas, non-English values, reasoning history, streaming termination and usage reporting.

Move a small cohort first and keep its ledger separate. Set an explicit rollback threshold for accepted results and tool-loop failures. Do not leave a fallback pointing to a model scheduled for retirement after that deadline. Model catalog reads are useful for discovering lifecycle changes, but they do not prove that generation or a long workflow succeeds.

[Create a KeepRouter key](/docs/quickstart), run the text example, then validate the read-only tool loop. The [agent context budgeting guide](/blog/context-window-exceeded-fix) helps reduce repeated prompt growth during migration.

Primary documentation checked for this guide：[Xiaomi · MiMo V2.6 release](https://mimo.mi.com/docs/en-US/news/latest/v2-6) · [Xiaomi · prices and retirement](https://mimo.mi.com/docs/pricing) · [Xiaomi · deep thinking](https://mimo.mi.com/docs/en-US/quick-start/usage-guide/other/deep-thinking).

## Frequently asked questions

### When do MiMo V2.5 and V2.5 Pro retire?

Xiaomi announces October 21, 2026 at 10:00 Beijing time. Recheck its lifecycle notice before migration.

### Is UltraSpeed the same pricing ID as Pro?

No. Use its exact ID and live price, and measure any latency benefit on your workload.

### Do these IDs generate audio?

They provide multimodal understanding with text output. Speech generation needs its own model and endpoint.

## Sources reviewed

_Article last reviewed 2026-10-03_

1. [Xiaomi · MiMo V2.6 release](https://mimo.mi.com/docs/en-US/news/latest/v2-6)
2. [Xiaomi · prices and retirement](https://mimo.mi.com/docs/pricing)
3. [Xiaomi · deep thinking](https://mimo.mi.com/docs/en-US/quick-start/usage-guide/other/deep-thinking)

## Related guides

- [mimo v2.6 flash](https://keeprouter.com/models/mimo-v2.6-flash.md)
- [mimo v2.6 pro](https://keeprouter.com/models/mimo-v2.6-pro.md)
- [mimo v2.6 pro ultraspeed](https://keeprouter.com/models/mimo-v2.6-pro-ultraspeed.md)
- [api cost calculator](https://keeprouter.com/tools/api-cost-calculator)
- [GLM 5.3 vs Flash vs FlashX: inputs, reasoning and API costs](https://keeprouter.com/blog/glm-5-3-flash-flashx-api-guide.md)
- [Grok 4.7, Qwen 3.8, Gemma 4 and Doubao 2.1: choose by task](https://keeprouter.com/blog/grok-qwen-gemma-doubao-model-selection.md)

## Check this model's price and API

See the current customer rate, supported endpoint and setup example for the model discussed here.

[View model and pricing](https://keeprouter.com/models/mimo-v2.6-pro)

[Create a key to test the free model](https://keeprouter.com/login?returnTo=%2Fconsole%2Fkeys%3Fmodel%3Dfree)

The free test uses the free model. Other paid models require sufficient prepaid credit.

[All posts](https://keeprouter.com/blog.md) · [Models & pricing](https://keeprouter.com/models.md)
