DocsAPI reference
Chat completions
POST https://api.avenro.tech/v1/chat/completions takes a conversation and returns the model's reply, in OpenAI's Chat Completions format.
Request
| Field | Type | Description |
|---|---|---|
| model* | string | A model id from GET /v1/models, for example deepseek-r1. |
| messages* | array | The conversation so far: system, user, assistant and tool messages, as in OpenAI's API. |
| stream | boolean | Send the reply as server-sent events. Default false. |
| stream_options | object | {"include_usage": true} adds a final chunk with token usage. |
| max_completion_tokens | integer | Most tokens to generate (alias max_tokens). Must not exceed the model's maximum output. It also sets how much is reserved while the request runs; without it the model's maximum output is reserved. |
| n | integer | Number of choices to generate, 1 to 8. Each choice is billed. |
| temperature, top_p, stop, seed | … | Sampling controls, passed to the model. |
| tools, tool_choice | … | Function calling, on models whose capabilities include tools. |
| response_format | object | JSON mode or structured output, on models that support it. |
Other OpenAI fields are passed to the model unchanged; whether a model honours them depends on the model. The store and metadata fields are never forwarded, because Avenro does not keep conversations.
Response
{
"id": "chatcmpl-6f1c2a8e9b4d4f0e8a3b5c7d9e1f2a3b",
"object": "chat.completion",
"created": 1791370800,
"model": "llama-3.1-70b-instruct",
"choices": [
{
"index": 0,
"message": { "role": "assistant", "content": "Hamlet seeks revenge…" },
"finish_reason": "stop"
}
],
"usage": { "prompt_tokens": 21, "completion_tokens": 48, "total_tokens": 69 }
}modelis always the id you requested, whichever deployment served it.usagereports the exact tokens billed.prompt_tokens_details.cached_tokensappears when part of the prompt came from a cache and was billed at the cached-input price.
Response headers
x-request-id: the request id, also shown in your dashboard. Quote it when you contact support.x-avenro-execution-path:self-hostedwhen Avenro's own clusters served the request,fallbackwhen an outside provider did (the maker of a proprietary model, for example).x-avenro-cost-usd: what the request cost, in dollars (non-streaming responses).x-ratelimit-limit-requests,x-ratelimit-remaining-requests,x-ratelimit-reset-requests: see rate limits.
Streaming
With stream: true the reply arrives as text/event-stream: one data: line per chunk, ending with data: [DONE]. During long pauses (a reasoning model thinking, for example) the stream carries : keep-alive comment lines, which SSE clients ignore.
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{"content":"Hello"}}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14}}
data: [DONE]If the model fails partway through, the stream ends with a chunk holding an error object (code upstream_stream_error) instead of [DONE]. You are billed for the prompt and the output delivered before that point, and nothing if no output arrived. If you close the connection yourself, generation stops and you are billed for the prompt and the output delivered.
Reasoning models
Reasoning models such as deepseek-r1 return their chain of thought separately from the answer, always in reasoning_content on the message (or on each streamed delta), whichever deployment serves the request. Reasoning tokens are output tokens and are billed as such. They can be long: set max_completion_tokens generously, or the answer may be cut off after the reasoning.
How a request is billed
Before the request is sent to a model, its largest possible cost is reserved: the prompt (estimated generously from its size) plus the output limit, times n. If your available balance cannot cover that, the request is refused with 402 insufficient_credits and nothing is charged; lowering max_completion_tokens lowers the reservation. When the request finishes, the reservation is released and the tokens actually used are charged. Requests the model rejects (400) and requests that fail before any output cost nothing.