All Models
Mercury 2.5
Mercury 2.5 Preview is Inception's latest and most intelligent diffusion language model. Instead of generating tokens strictly one at a time, it produces and refines multiple tokens in parallel, reaching up to 1,107 tokens per second on standard GPUs. It delivers a 10+ point intelligence gain over Mercury 2, with tunable reasoning, parallel tool calls, schema-aligned JSON output, and a 260K context window. It is built for latency-sensitive production work such as search agents, voice pipelines, customer support, rapid coding iteration, and coding subagents.
Available Providers (6)
| Provider | Model ID | Input Cost | Output Cost | Context | Max Output | Docs |
|---|---|---|---|---|---|---|
inception/mercury-2.5-preview | $0.04/MTok | $0.15/MTok | 260K | 65.5K | ||
inception/mercury-2.5 | $0.04/MTok | $0.15/MTok | 260K | 65.5K | ||
mercury-2.5 | $0.04/MTok | $0.15/MTok | 260K | 65.5K | ||
inception/mercury-2.5 | $0.04/MTok | $0.15/MTok | 260K | 65.5K | ||
mercury-2-5 | $0.05/MTok | $0.19/MTok | 260K | 65.5K | ||
inception/mercury-2.5 | $0.20/MTok | $0.75/MTok | 260K | 65.5K |
Capabilities
Reasoning
Tool Calling
Attachments
Open Weights
Structured Output