Skip to content

Model ecosystem

Models & prices.

Prices are in US dollars per million tokens, charged from your credits. The struck-through figure is the reference list price we measure savings against; its source is listed below. Latency is the median of the last 24 hours of streamed requests, shown once there are enough of them.

9 of 9 models

Models, prices per million tokens and measured latency
ModelContextInput / 1MCached input / 1MOutput / 1MSavingLatency (24 h)
Claude Fable 5.1
claude-fable-5-1

Anthropic's model for demanding reasoning and long-horizon agentic work, served by Anthropic at 40% below its list price. Thinking is always on and billed as output. Anthropic's OpenAI-compatible API, which serves it, does not return the thinking and offers neither prompt caching nor response_format.

ToolsVisionReasoning
1M
out 128K
$6.00
$10.00
—
$30.00
$50.00
−40%—
Claude Opus 5
claude-opus-5

Anthropic's Claude Opus 5 (a legacy model; Anthropic's current Opus is 5.5), served by Anthropic at 40% below its list price. Adaptive thinking is on by default and billed as output. Anthropic's OpenAI-compatible API, which serves it, does not return the thinking and offers neither prompt caching nor response_format.

ToolsVisionReasoning
1M
out 128K
$3.00
$5.00
—
$15.00
$25.00
−40%—
Claude Sonnet 5
claude-sonnet-5

Anthropic's Claude Sonnet 5 (a legacy model; Anthropic's current Sonnet is 5.5), served by Anthropic at 40% below its list price. Adaptive thinking is on by default and billed as output; non-default temperature or top_p values are rejected. Anthropic's OpenAI-compatible API, which serves it, does not return the thinking and offers neither prompt caching nor response_format.

ToolsVisionReasoning
1M
out 128K
$1.20
$2.00
—
$6.00
$10.00
−40%—
GPT-6 Astra
gpt-6-astra

OpenAI's GPT-6 Astra reasoning model, served by OpenAI at 40% below its list price. A request whose prompt exceeds 272,000 tokens is billed at 2x input and 1.5x output prices in full, as OpenAI bills it. Prompt caching is off through Avenro, so every input token is billed at the input price. reasoning_effort accepts low, medium, high, xhigh and max.

  • Prompts over 272,000 tokens: the whole request at $12.00 input and $45.00 output per 1M.
ToolsJSON modeVisionReasoning
1.05M
out 128K
$6.00
$10.00
—
$30.00
$50.00
−40%—
GPT-4.1 mini
gpt-4.1-mini

OpenAI's proprietary GPT-4.1 mini with a 1M-token context window, served through OpenAI at 40% below OpenAI's list price.

ToolsJSON modeVision
1.05M
out 33K
$0.24
$0.40
$0.06
$0.10
$0.96
$1.60
−40%—
GPT-4o mini
gpt-4o-mini

OpenAI's proprietary GPT-4o mini, served through OpenAI at 40% below OpenAI's list price.

ToolsJSON modeVision
128K
out 16K
$0.09
$0.15
$0.045
$0.075
$0.36
$0.60
−40%—
Gemini 3.8 Flash
gemini-3.8-flash

Google's Gemini 3.8 Flash, served by the Gemini API's paid tier (which does not use prompts to improve Google's products) at 40% below its list price. Thinking tokens are billed as output. Google raises this price on 2027-01-01; Avenro withdraws the model at that moment until it posts the new price.

  • Listed at this price until Jan 1, 2027, 12:00 AM UTC, when the provider's price changes.
ToolsJSON modeVisionReasoning
1.05M
out 66K
$0.45
$0.75
$0.045
$0.075
$2.25
$3.75
−40%—
Grok 4.6
grok-4.6

xAI's Grok 4.6 reasoning model, served by xAI at 40% below its list price. A request whose prompt reaches 200,000 tokens is billed at double prices in full, as xAI bills it. Reasoning tokens are billed as output.

  • Prompts of 200,000 tokens or more: the whole request at $2.40 input, $0.60 cached input and $7.20 output per 1M.
ToolsJSON modeVisionReasoning
500K
out 500K
$1.20
$2.00
$0.30
$0.50
$3.60
$6.00
−40%—
DeepSeek V4 ProOpen weights
deepseek-v4-pro

DeepSeek-V4-Pro (0813, open weights under the MIT licence), served by DeepSeek's own API, whose servers are in China. Avenro charges 40% below DeepSeek's peak-hour list price at all hours (DeepSeek halves its own price off-peak). Thinking mode is enabled with "thinking": {"type": "enabled"}; its reasoning arrives in reasoning_content and is billed as output.

ToolsReasoning
1M
out 384K
$0.792
$1.32
$0.0264
$0.044
$2.376
$3.96
−40%—

Sources

Reference prices

For each model, the public list price the savings are measured against. List prices change; each one shows the date it was last checked.

The same list is available to your code: GET https://api.avenro.tech/v1/models.