AIREITER

AI Image

Grok Imagine Image 2.0Midjourney V8.1Midjourney V7Z-Image TurboKrea 2 TurboSeedream 5.0 Pro LayerizeQwen Image 3.0 ProMore

AI Video

HappyHorse 1.0HappyHorse 1.1Gemini Omni FlashFLUX 3 VideoVeo 3.1Veo 3.1 FastSeedance 2.0 Fast FaceMore

LLM

MiniMax M3GLM 5.2Doubao Seed 2.1 TurboKimi K2.7 CodeDeepSeek V4 FlashDeepSeek V4 ProClaude Opus 5More
Coming soonaa
API DOCSPRICING
TEMPLATES
minimaxText Chat

MiniMax M3 API: long-context text model for documents and agent memory

Use MiniMax M3 for long-context document review, agent memory inspection, knowledge workflows, and production text automation.

InputAIReiter $0.60 per 1M tokensOutputAIReiter $2.40 per 1M tokensCache readAIReiter $0.12 per 1M tokens
Run with API
PlaygroundReadmeAPI

INPUT

1
2
3
4
5
6
7
8
9
10
11
12

Install the official Anthropic client — AIReiter speaks the same protocol, so only the base URL changes:

npm install @anthropic-ai/sdk

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import Anthropic from "@anthropic-ai/sdk";

const client = new Anthropic({
  apiKey: process.env.AIREITER_API_KEY,
  baseURL: "https://aireiter.com/api",
});

Run minimax-m3:

const message = await client.messages.create({
    "model": "minimax-m3",
    "max_tokens": 2048,
    "messages": [
      {
        "role": "user",
        "content": "Read this long project note and extract decisions, open questions, and owners."
      }
    ],
    "temperature": 0.7,
    "top_p": 1
  });

console.log(message.content);

Stream the response instead:

const stream = client.messages.stream({
    "model": "minimax-m3",
    "max_tokens": 2048,
    "messages": [
      {
        "role": "user",
        "content": "Read this long project note and extract decisions, open questions, and owners."
      }
    ],
    "temperature": 0.7,
    "top_p": 1
  });

stream.on("text", (text) => process.stdout.write(text));
const message = await stream.finalMessage();

Install the official Anthropic client — AIReiter speaks the same protocol, so only the base URL changes:

pip install anthropic

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Point the client at AIReiter:

import os
import anthropic

client = anthropic.Anthropic(
    api_key=os.environ["AIREITER_API_KEY"],
    base_url="https://aireiter.com/api",
)

Run minimax-m3:

message = client.messages.create(
      model = "minimax-m3",
      max_tokens = 2048,
      messages = [
        {
          role = "user",
          content = "Read this long project note and extract decisions, open questions, and owners."
        }
      ],
      temperature = 0.7,
      top_p = 1
)

print(message.content)

Stream the response instead:

with client.messages.stream(
      model = "minimax-m3",
      max_tokens = 2048,
      messages = [
        {
          role = "user",
          content = "Read this long project note and extract decisions, open questions, and owners."
        }
      ],
      temperature = 0.7,
      top_p = 1
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

Set the AIREITER_API_KEY environment variable:

export AIREITER_API_KEY=<paste-your-key-here>

Run minimax-m3 against AIReiter's API:

curl -s -X POST \
  -H "x-api-key: $AIREITER_API_KEY" \
  -H "Content-Type: application/json" \
  "https://aireiter.com/api/v1/messages" \
  -d '{
  "model": "minimax-m3",
  "max_tokens": 2048,
  "messages": [
    {
      "role": "user",
      "content": "Read this long project note and extract decisions, open questions, and owners."
    }
  ],
  "temperature": 0.7,
  "top_p": 1
}'

Add "stream": true to the body to receive the response as server-sent events.

OUTPUT

Example

A rate limit caps how many requests an API accepts from you in a given window. Once you exceed it, the server stops doing work for you and answers 429 Too Many Requests instead.

Handling 429

  1. Read the Retry-After response header. When present it tells you exactly how long to wait, in seconds.
  2. When it is absent, back off exponentially with jitter so retries from many clients do not line up.
  3. Cap the number of retries, then surface the failure instead of looping forever.
async function withRetry(request, maxRetries = 4) {
  for (let attempt = 0; ; attempt++) {
    const response = await request();
    if (response.status !== 429 || attempt === maxRetries) return response;
    const retryAfter = Number(response.headers.get("retry-after"));
    const backoff = Number.isFinite(retryAfter) ? retryAfter * 1000 : 2 ** attempt * 500 + Math.random() * 250;
    await new Promise((resolve) => setTimeout(resolve, backoff));
  }
}

Treat the limit as a budget you plan around, not an error you retry your way out of: batch requests where you can, cache repeated reads, and spread bulk work over time.

{
  "model": "minimax-m3",
  "input": {
    "model": "minimax-m3",
    "max_tokens": 2048,
    "messages": [
      {
        "role": "user",
        "content": "Read this long project note and extract decisions, open questions, and owners."
      }
    ],
    "temperature": 0.7,
    "top_p": 1
  },
  "output": "A rate limit caps how many requests an API accepts from you in a given window. Once you exceed it, the server stops doing work for you and answers **429 Too Many Requests** instead.\n\n## Handling 429\n\n1. Read the `Retry-After` response header. When present it tells you exactly how long to wait, in seconds.\n2. When it is absent, back off exponentially with jitter so retries from many clients do not line up.\n3. Cap the number of retries, then surface the failure instead of looping forever.\n\n```js\nasync function withRetry(request, maxRetries = 4) {\n  for (let attempt = 0; ; attempt++) {\n    const response = await request();\n    if (response.status !== 429 || attempt === maxRetries) return response;\n    const retryAfter = Number(response.headers.get(\"retry-after\"));\n    const backoff = Number.isFinite(retryAfter) ? retryAfter * 1000 : 2 ** attempt * 500 + Math.random() * 250;\n    await new Promise((resolve) => setTimeout(resolve, backoff));\n  }\n}\n```\n\nTreat the limit as a budget you plan around, not an error you retry your way out of: batch requests where you can, cache repeated reads, and spread bulk work over time.",
  "metrics": {
    "input_tokens": 26,
    "output_tokens": 214,
    "generated_in_seconds": 4.1
  },
  "example": true
}
Generated in
4.1 seconds
Input tokens
26
Output tokens
214
Tokens per second
52.20 tokens / second
Time to first token
-

Model details

Use the same model key in Playground, API requests, and internal workflows.

Model ID
minimax-m3
Provider
minimax
Protocol
Anthropic Messages
Context window
1,000,000 tokens
Max output
-
Input tokens
60 credits / 1M tokens
Output tokens
240 credits / 1M tokens
Cache read
12 credits / 1M tokens
Cache write
-

A long-context route for documents and agent memory

MiniMax M3 is useful when a workflow must keep more evidence in one request: policy sets, research packs, meeting history, or long agent memory.

MiniMax M3 API cover

Should you choose MiniMax M3?

Position it as a practical long-context model for document-heavy teams and agent workflows that need continuity.

Choose it when

You need long document review, memory inspection, knowledge-base synthesis, or a cheaper long-context route before escalating to a flagship model.

Use another model when

The task is a small chat turn, simple extraction, or high-risk code reasoning that needs a more specialized coding model.

Public API protocol

Call POST https://aireiter.com/api/v1/messages with model "minimax-m3". Streaming is supported through the same Messages-compatible endpoint.

Token and cache usage

Pricing is based on input, cache-read, and output usage. For long prompts, cache-read fields matter because reused context can materially change cost.

MiniMax M3 production workloads

Best for long inputs where continuity and summarization quality matter.
01

Document packs

Review policies, reports, transcripts, and research notes while keeping cross-document references visible.

02

Agent memory

Inspect long histories, tool calls, and state updates to explain why an agent made a decision.

03

Knowledge workflows

Turn long internal material into summaries, briefs, requirements, and action lists.

04

Cost-aware long context

Evaluate whether a practical long-context model is enough before using a higher-cost flagship route.

How MiniMax M3 fits your model stack

Do not route every request to the newest model. Pick the cheapest model that still passes your quality bar, then reserve deeper models for failures or high-risk tasks.

For fast batches

Use Doubao or DeepSeek V4 Flash for small fast tasks; MiniMax M3 is for context-heavy inputs.

For deeper reasoning

Use GLM 5.2 or DeepSeek V4 Pro when reasoning depth is more important than context size.

For long context

Compare MiniMax M3 with Kimi K2.7 Code when the workload mixes documents, code, and agent traces.

For production rollout

Track cache-read usage on repeated long prefixes; this is where long-context workflows can become more economical.

MiniMax M3 API questions

Questions developers usually check before moving a text model from playground testing to production API traffic.

/ 01

What model ID should I send for MiniMax M3?

Use "minimax-m3" in the API request body. The internal DB key is only used by AIReiter routing.

/ 02

Which endpoint should MiniMax M3 use?

Use POST https://aireiter.com/api/v1/messages for public API calls. Keep x-api-key / Authorization authentication consistent with your AIReiter API key setup.

/ 03

Does MiniMax M3 support streaming?

Yes. Send stream=true and read server-sent events until the message completes. Test non-streaming first when debugging authentication or model ID issues.

/ 04

How do I confirm token and cache billing for MiniMax M3?

Check the usage object returned by the API. Input, output, and cache-read token fields are the source of truth for settlement; a repeated prompt alone does not prove a cache hit.

/ 05

Should I always set max_tokens for MiniMax M3?

Set max_tokens according to the expected summary length. Long input does not always require huge output, but review tasks need enough room for findings and evidence.

AIREITER

Questions? Contact us at
support@aireiter.com

新速率有限公司NEWRATE LIMITED香港九龍花園街 2-16 號好景商業中心 2304 室Room 2304, Haojing Commercial Center, 2-16 Garden Street, Kowloon, Hong Kong

LLM

MiniMax M3GLM 5.2Doubao Seed 2.1 TurboKimi K2.7 CodeDeepSeek V4 Flash

AI Video

HappyHorse 1.0HappyHorse 1.1Gemini Omni FlashFLUX 3 VideoVeo 3.1

AI Image

Grok Imagine Image 2.0Midjourney V8.1Midjourney V7Z-Image TurboKrea 2 Turbo

Blog

View All →

Company

Privacy PolicyTerms of ServiceRefund Policy

© 2026 AIReiter. All rights reserved.