LLM Reference

MAI-Thinking-1

Released
2026-06-02
Last refreshed
2026-08-24
Status
Researched 7d ago
ProprietaryCommercial use: conditionalCodingRAGAgentsLong contextClassificationJSON / Tool useHighlight

MAI-Thinking-1 is worth evaluating for coding, rag, and agents when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, rag, and agents
  • Workloads that can use a 256k context window
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Vision or document-understanding workloads
Specifications
Family
MAI
Released
2026-06-02
Context
256k
Max output
64,000
Parameters
1T total / 35B active
Architecture
Mixture of Experts
Specialization
reasoning
Openness
Proprietary
License
ProprietaryCommercial use: conditional
Weights
Not released
Code
Unknown
Training
Pretrained
Created by

Applied AI products and platforms from Microsoft

Redmond, Washington, United States
Website
Pricing
Output / 1M
-
Input / 1M
-

Cheapest of 1 route · Microsoft Foundry

About

MAI-Thinking-1 is Microsoft AI's flagship reasoning model, built from scratch on enterprise-grade commercially licensed data without third-party distillation. The sparse mixture-of-experts model activates about 35B parameters from roughly 1T total parameters, supports a 256K-token context window, and targets frontier reasoning and software engineering work at a mid-weight price point. Microsoft reports 97% on AIME 2025, 94.5% on AIME 2026, 84.2% on GPQA Diamond, 87.7% on LiveCodeBench v6, 73.5% on SWE-bench Verified, and 52.8% on SWE-bench Pro. In a 1,276-task Surge blind side-by-side evaluation, it narrowly beat Claude Sonnet 4.6 but trailed Claude Opus 4.6.

Top use-case fit: coding, agents, and build tasks

Coding

3 relevant benchmarks in the decision map.

RAG

Included by capability and metadata signals in the decision map.

Agents

2 relevant benchmarks in the decision map.

Capabilities

ReasoningJSON / Tool use

Benchmark peer barsfor Coding

Benchmark scores(10)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
AIME 202597.0AIME 2025Observed 2026-06-02Source
AIME 202694.5AIME 2026Observed 2026-06-02Source
HMMT February 202684.9HMMT Feb 2026Observed 2026-06-02Source
Google-Proof Q&A84.2GPQA DiamondObserved 2026-06-02Source
LiveCodeBench87.7v6Observed 2026-06-02Source
Terminal-Bench 2.046.0Terminal-Bench 2.0Observed 2026-06-02Source
SWE-bench Verified73.5SWE-bench VerifiedObserved 2026-06-02Source
SWE-bench Pro52.8Public datasetObserved 2026-06-02Source
MMLU PRO85.0MMLU-Pro (accuracy)Observed 2026-06-07Source
MultiChallenge53.0Multi-Challenge leaderboard rank 15 of 28 (accuracy%)Observed 2026-06-07Source

Migration checks

No linked migration route is available for this model yet.