Novita AI LLM

MiMo V2.6 Pro on Novita AI: API Quickstart and Pricing

MiMo V2.6 Pro on Novita AI: API Quickstart and Pricing

MiMo V2.6 Pro is available on Novita AI as a serverless, OpenAI-compatible chat model. Use the model ID xiaomimimo/mimo-v2.6-pro, send requests through https://api.novita.ai/openai, and budget $4.35 per 1 million input tokens, $0.036 per 1 million cached input tokens, and $8.70 per 1 million output tokens at the time of writing.

What is MiMo V2.6 Pro?

MiMo V2.6 Pro is Xiaomi’s flagship reasoning model in the MiMo V2.6 family. The Novita AI catalog lists it as a serverless chat model with reasoning, function calling, and structured-output support. It accepts text, image, video, and audio input and returns text output.

The model is a practical fit for applications that need more than short text completion: long-horizon agent loops, repository or document analysis, multimodal research, tool-driven workflows, and structured responses. That does not mean every request should use the maximum context or output allowance. For short prompts and high-volume classification, a smaller model can be easier to budget and faster to operate.

The verified Novita AI model ID is:

xiaomimimo/mimo-v2.6-pro

MiMo V2.6 Pro specifications and pricing

The following values come from the live Novita AI model catalog checked on September 30, 2026.

PropertyMiMo V2.6 Pro on Novita AI
Model IDxiaomimimo/mimo-v2.6-pro
AccessServerless API
Context window1,048,576 tokens
Maximum output131,072 tokens
Input modalitiesText, image, video, audio
Output modalityText
Endpoint familiesChat completions, Anthropic, Responses
FeaturesReasoning, function calling, structured outputs
Input price$4.35 per 1M tokens
Cached input price$0.036 per 1M tokens
Output price$8.70 per 1M tokens

Novita reports prices in U.S. dollars per million tokens. A simple uncached request with 10,000 input tokens and 2,000 output tokens would cost about $0.0609: $0.0435 for input plus $0.0174 for output. Actual usage depends on the tokens sent and generated, so use the usage data returned by the API for accounting rather than estimating from character count.

The cached-input rate is much lower than the regular input rate. It matters when an application repeatedly sends a stable system prompt, tool schema, document prefix, or other context that the API can read from the prompt cache. Design caching around repeated prefixes, and verify the response usage fields before assuming that a request received the cached rate.

Prices, limits, and availability can change. Check the MiMo V2.6 Pro model page and Novita AI pricing before committing to a production budget.

How to call MiMo V2.6 Pro with Python

The OpenAI Python client can call the Novita AI compatibility endpoint by changing the base URL, API key, and model ID. Install the client and set the key in your shell:

pip install openai
export NOVITA_API_KEY="YOUR_NOVITA_API_KEY"

Then send a chat completion:

import os

from openai import OpenAI


client = OpenAI(
    api_key=os.environ["NOVITA_API_KEY"],
    base_url="https://api.novita.ai/openai",
)

response = client.chat.completions.create(
    model="xiaomimimo/mimo-v2.6-pro",
    messages=[
        {
            "role": "system",
            "content": "You are a precise software architecture assistant.",
        },
        {
            "role": "user",
            "content": "Review this service boundary and list the three highest-risk failure modes.",
        },
    ],
    max_tokens=2048,
    temperature=0.2,
)

print(response.choices[0].message.content)

Start with a modest max_tokens value. The catalog maximum is an upper limit, not a recommendation for every request. Keeping the requested output close to the task also makes cost and latency easier to control.

The compatibility layer uses the standard chat-completions shape. For the full request and response fields, see the Novita AI chat completions API reference.

Migrating from MiMo V2.5

If you already call MiMo V2.5 through Novita AI, the first migration test is small: keep the same OpenAI-compatible base URL, replace the model ID with xiaomimimo/mimo-v2.6-pro, and run your existing prompt suite. Do not assume that a larger context limit or additional input modality automatically improves your application. Compare answer quality, tool-call arguments, structured-output validity, latency, and token usage on representative requests.

Keep the old model available while you evaluate the new one. A configuration-level model switch makes it easier to roll back if a prompt, tool schema, or response parser behaves differently. Recheck price assumptions as well: the V2.6 Pro rates in this article are tied to the live catalog checked on September 30, 2026, not a permanent price guarantee.

How to call MiMo V2.6 Pro with cURL

For a direct HTTP request, post to /openai/v1/chat/completions and keep the model ID in the JSON body:

curl --request POST \
  --url https://api.novita.ai/openai/v1/chat/completions \
  --header "Authorization: Bearer $NOVITA_API_KEY" \
  --header "Content-Type: application/json" \
  --data '{
    "model": "xiaomimimo/mimo-v2.6-pro",
    "messages": [
      {
        "role": "system",
        "content": "You are a concise research assistant."
      },
      {
        "role": "user",
        "content": "Turn these notes into a numbered implementation plan."
      }
    ],
    "max_tokens": 1024,
    "temperature": 0.2
  }'

Keep NOVITA_API_KEY in an environment variable or secret manager. Do not put it in source control, browser code, or a client-side application bundle.

What should you build with it?

MiMo V2.6 Pro is worth evaluating when the task benefits from a large working context or deliberate multi-step reasoning:

  • Repository-scale coding assistants: provide related files, constraints, and test failures in one context, then ask for a plan or patch outline.
  • Research and document workflows: combine long source material with a structured extraction schema instead of splitting every document into isolated prompts.
  • Tool-using agents: use function calling to let the model request actions such as search, ticket lookup, or deployment checks while your application retains execution control.
  • Multimodal analysis: pass supported image, video, or audio inputs when the application needs to reason over more than text.
  • Structured business workflows: request JSON-shaped output for routing, extraction, or evaluation, then validate it in application code before using it.

The strongest reason to choose this model is not the headline context size by itself. It is the combination of long context, reasoning, tool support, and multimodal input behind one hosted API. If your workload only needs short text responses, compare cost and latency with smaller options in the Novita AI model library.

Important limits and implementation notes

Treat the context window as a budget

The 1,048,576-token context window is a ceiling. Large prompts still cost money and can increase latency. Retrieve only the files or passages relevant to the current step, summarize stale conversation history, and reserve output space when constructing requests.

Validate structured output

Structured-output support helps constrain the response format, but your application should still parse and validate the returned data. Reject missing required fields, unexpected values, and invalid JSON before triggering an external action.

Keep tool execution outside the model

Function calling describes the action the model wants to take; it does not grant the model access to your systems. Your application should authenticate the tool call, validate arguments, apply authorization checks, execute the operation, and return only the necessary result.

Measure cache behavior

Prompt caching can materially change input cost for repeated prefixes, but it should be measured from response usage rather than assumed. Track input, cached-input, output, and total token counts for representative traffic before changing your budget model.

Start with a small production slice

Before routing all traffic, test your own prompts for tool-call accuracy, structured-output validity, multimodal quality, timeout behavior, and cost. A model’s catalog features tell you what the endpoint supports; they do not replace an evaluation against your application’s failure cases.

FAQ

What is the MiMo V2.6 Pro model ID on Novita AI?

Use xiaomimimo/mimo-v2.6-pro.

How much does MiMo V2.6 Pro cost on Novita AI?

As checked on September 30, 2026, Novita lists $4.35 per 1M input tokens, $0.036 per 1M cached input tokens, and $8.70 per 1M output tokens. Confirm the live model page before production use because pricing can change.

What context window does MiMo V2.6 Pro have?

Novita lists a 1,048,576-token context window and a 131,072-token maximum output.

Does MiMo V2.6 Pro support images, video, and audio?

The Novita catalog lists text, image, video, and audio as input modalities, with text as the output modality. Test the exact content format your application needs before relying on a multimodal workflow in production.

Which API endpoint should I use?

Use the OpenAI-compatible base URL https://api.novita.ai/openai with the Python client, or call https://api.novita.ai/openai/v1/chat/completions directly with HTTP.

Is MiMo V2.6 Pro a good default for every request?

No. It is a better candidate for long-context, reasoning, tool-use, structured-output, or multimodal work. For short and repetitive requests, compare smaller models for lower cost and faster responses.

Sources

Related Posts