All Models

Mercury 2.5

mercuryReasoningTool CallingStructured Output

Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.

Providers6
ReleasedSep 1, 2026
Input Modalitiestext
Output Modalitiestext
Tarsk Usecoding

Available Providers (6)

ProviderModel IDInput CostOutput CostContextMax OutputDocs
NanoGPTinception/mercury-2.5-preview$0.04/MTok$0.15/MTok260K65.5K
Vercel AI Gatewayinception/mercury-2.5$0.04/MTok$0.15/MTok260K65.5K
Inceptionmercury-2.5$0.04/MTok$0.15/MTok260K65.5K
OpenRouterinception/mercury-2.5$0.04/MTok$0.15/MTok260K65.5K
Venice AImercury-2-5$0.05/MTok$0.19/MTok260K65.5K
Kilo Gatewayinception/mercury-2.5$0.20/MTok$0.75/MTok260K65.5K

Capabilities

Reasoning
Tool Calling
Attachments
Open Weights
Structured Output