All Models
Vision Small
Fast, low-cost multimodal model for understanding text, images, audio, video, and PDFs, with tool calling and a 1M-token context window.
Available Providers (1)
| Provider | Model ID | Input Cost | Output Cost | Context | Max Output | Docs |
|---|---|---|---|---|---|---|
vispark/vision-small | $1.05/MTok | $3.16/MTok | 1M | 65.5K |
Capabilities
Reasoning
Tool Calling
Attachments
Open Weights
Structured Output