All Models

pixtral-12b-2409

Reasoning Tool Calling Attachments Open Weights Structured Output

Pixtral 2409 12B is a state-of-the-art multimodal model with 12B parameters and a 400M vision encoder, natively trained on interleaved text and image data. It excels in tasks spanning vision-language reasoning, instruction following, and pure text understanding, making it highly effective for real-world multimodal applications.

Providers 2
Released Sep 25, 2024
Input Modalities text, image
Output Modalities text
Tarsk Use coding

Available Providers (2)

Provider Model ID Input Cost Output Cost Context Max Output Docs
Scaleway pixtral-12b-2409 $0.20/MTok $0.20/MTok 128K 4.1K
Cortecs pixtral-12b-2409 $0.22/MTok $0.22/MTok 128K 128K

Capabilities

Reasoning
Tool Calling
Attachments
Open Weights
Structured Output