Skip to content

DocsAPI reference

Chat completions

POST https://api.avenro.tech/v1/chat/completions takes a conversation and returns the model's reply, in OpenAI's Chat Completions format.

Request

FieldTypeDescription
model*stringA model id from GET /v1/models, for example deepseek-r1.
messages*arrayThe conversation so far: system, user, assistant and tool messages, as in OpenAI's API.
streambooleanSend the reply as server-sent events. Default false.
stream_optionsobject{"include_usage": true} adds a final chunk with token usage.
max_completion_tokensintegerMost tokens to generate (alias max_tokens). Must not exceed the model's maximum output. It also sets how much is reserved while the request runs; without it the model's maximum output is reserved.
nintegerNumber of choices to generate, 1 to 8. Each choice is billed.
temperature, top_p, stop, seed…Sampling controls, passed to the model.
tools, tool_choice…Function calling, on models whose capabilities include tools.
response_formatobjectJSON mode or structured output, on models that support it.

Other OpenAI fields are passed to the model unchanged; whether a model honours them depends on the model. The store and metadata fields are never forwarded, because Avenro does not keep conversations.

Response

200 OK
{
  "id": "chatcmpl-6f1c2a8e9b4d4f0e8a3b5c7d9e1f2a3b",
  "object": "chat.completion",
  "created": 1791370800,
  "model": "llama-3.1-70b-instruct",
  "choices": [
    {
      "index": 0,
      "message": { "role": "assistant", "content": "Hamlet seeks revenge…" },
      "finish_reason": "stop"
    }
  ],
  "usage": { "prompt_tokens": 21, "completion_tokens": 48, "total_tokens": 69 }
}
  • model is always the id you requested, whichever deployment served it.
  • usage reports the exact tokens billed. prompt_tokens_details.cached_tokens appears when part of the prompt came from a cache and was billed at the cached-input price.

Response headers

  • x-request-id: the request id, also shown in your dashboard. Quote it when you contact support.
  • x-avenro-execution-path: self-hosted when Avenro's own clusters served the request, fallback when an outside provider did (the maker of a proprietary model, for example).
  • x-avenro-cost-usd: what the request cost, in dollars (non-streaming responses).
  • x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests: see rate limits.

Streaming

With stream: true the reply arrives as text/event-stream: one data: line per chunk, ending with data: [DONE]. During long pauses (a reasoning model thinking, for example) the stream carries : keep-alive comment lines, which SSE clients ignore.

Event stream
data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{"role":"assistant","content":""}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{"content":"Hello"}}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[{"index":0,"delta":{},"finish_reason":"stop"}]}

data: {"id":"chatcmpl-…","object":"chat.completion.chunk","model":"deepseek-r1","choices":[],"usage":{"prompt_tokens":12,"completion_tokens":2,"total_tokens":14}}

data: [DONE]

If the model fails partway through, the stream ends with a chunk holding an error object (code upstream_stream_error) instead of [DONE]. You are billed for the prompt and the output delivered before that point, and nothing if no output arrived. If you close the connection yourself, generation stops and you are billed for the prompt and the output delivered.

Reasoning models

Reasoning models such as deepseek-r1 return their chain of thought separately from the answer, always in reasoning_content on the message (or on each streamed delta), whichever deployment serves the request. Reasoning tokens are output tokens and are billed as such. They can be long: set max_completion_tokens generously, or the answer may be cut off after the reasoning.

How a request is billed

Before the request is sent to a model, its largest possible cost is reserved: the prompt (estimated generously from its size) plus the output limit, times n. If your available balance cannot cover that, the request is refused with 402 insufficient_credits and nothing is charged; lowering max_completion_tokens lowers the reservation. When the request finishes, the reservation is released and the tokens actually used are charged. Requests the model rejects (400) and requests that fail before any output cost nothing.

Balances can go slightly below zero when a request turns out to cost more than was reserved; new requests are then refused until you add credits. See Credits & billing.