All Models
Llama 3.2 11B Vision Instruct
Open Llama multimodal model for image understanding and text reasoning
Available Providers (5)
| Provider | Model ID | Input Cost | Output Cost | Context | Max Output | Docs |
|---|---|---|---|---|---|---|
| | meta/llama-3.2-11b-vision-instruct | $0/MTok | $0/MTok | 128K | 4.1K | |
| | @cf/meta/llama-3.2-11b-vision-instruct | $0.05/MTok | $0.68/MTok | 128K | 128K | |
| | workers-ai/@cf/meta/llama-3.2-11b-vision-instruct | $0.05/MTok | $0.68/MTok | 128K | 128K | |
| | meta/llama-3.2-11b-vision-instruct | $0.06/MTok | $0.06/MTok | 16K | 4.1K | |
| | deepinfra/meta-llama/Llama-3.2-11B-Vision-Instruct | $0.34/MTok | $0.34/MTok | 131.1K | 4.1K |
Capabilities
Reasoning
Tool Calling
Attachments
Open Weights
Structured Output