Skip to content

OpenAI-compatible gateway

AI credits for less.

Unified access to leading AI models, powered by intelligent inference optimization.

Endpoint
api.avenro.tech/v1
Models
9, from Anthropic, OpenAI, Google, xAI and DeepSeek
Pricing
40% below list price, per token
Billing
Prepaid credits, no subscriptions

412 in · 186 out tokens, at the routed model's posted pricesexample request

Diagram: your app sends a request to the Avenro gateway, which checks the API key, routes the request to Anthropic, OpenAI, Google, xAI and DeepSeek, streams the answer back and meters it by the token.
Request lifecycle

One request, start to finish.

Every call is authenticated, routed to the model you named, streamed back and metered to the token. Nothing is swapped, and prompts are never stored.

Request

Trace3f9c2a7e-58d1-4b6e-9a40-c2e7d81f05b3

complete
  1. Request

    POST /v1/chat/completions

    model "claude-fable-5-1" · stream true

    “Write a two-line release note: exports now run in the background.”

  2. Auth

    Authorization: Bearer avenro_live_••••3fq1

    Key active, within its rate and spending limits.

  3. Route

    → Anthropic · claude-fable-5-1

    The model you named, never a substitute.

  4. Stream

    Example answer: Exports now run in the background, so you can keep working while large files are prepared. You'll get a notification as soon as yours is ready.

    31 tokens streamed

  5. Meter

    22 in · 31 out → $0.001062

    Exact to the token, rounded up to the next $0.000000001.

  6. Response

    200 OK · x-request-id 3f9c2a7e-58d1…

    Logged to Usage with its tokens and cost.

Model ecosystem

5 providers. One endpoint.

Every model is served through its provider's own API, 40% below list price. Pick a model to see what it costs.
Compare all models
Avenro/v1/chat/completionsmodel "claude-fable-5-1"
  • Anthropic3

  • OpenAI3

  • Google1

  • xAI1

  • DeepSeek1

Output price per million tokensOne endpoint, one key, one balance

Inference optimization, priced in.

  1. Aggregate

    Requests from every customer are pooled, so capacity runs at high, steady utilization.

  2. Optimize

    Open-weights models can run on Avenro's own vLLM clusters instead of behind a retail API's markup.

  3. Route

    Each request goes to the lowest-cost healthy deployment and fails over when one is busy or down.

  4. Reduce

    Higher utilization lowers the infrastructure cost behind every token.

  5. Pass it on

    You pay a posted price below the reference list price, and every request reports its saving.

At reference prices
$100.00
Through Avenro
$60.00
You save
40%
Metering & control

You control the spend.

A prepaid balance, a limit on every key, and every request metered to the token.
  • Prepaid balanceRequests draw on credits you have bought, so there is no invoice to surprise you.
  • Per-key limitsGive each key its own cap; a key stops at its limit, whatever the balance.
  • Usage by model and keySpend by day, and every request with its tokens and cost.

Controlsample account

Balance
$24.94
of $50.00 prepaid
Spend
$25.06
30 days · all keys
Requests
1,370
30 days
Tokens
2.8M
in + out

Spend trace$25.06 over 30 days

Model usage

  • Claude Fable 5.1$11.03

  • Claude Opus 5$6.49

  • Claude Sonnet 5$4.37

  • Other models$3.17

API keys

  • production$18.29 / $40.00

  • eval-runner$5.26 / $10.00

  • local-dev$1.51 · no limit

A key stops at its limit, whatever the balance.

Request logtail

Integration

Your code stays. Two lines change.

Avenro speaks OpenAI's Chat Completions format: point the client you already use at Avenro and use your key.
  1. Create a key

    Sign up, then create a key under API keys. It is shown once, so keep it as AVENRO_API_KEY.

  2. Add credits

    Top up from Billing with USDG or USDC on Robinhood Chain, from $1. Credits don't expire.

  3. Swap two values

    Point your client at https://api.avenro.tech/v1 and use your key.

1import OpenAI from "openai";
2
3const client = new OpenAI({
4 baseURL: "https://api.openai.com/v1",
5 apiKey: process.env.OPENAI_API_KEY,
6});
7
8const stream = await client.chat.completions.create({
9 model: "gpt-4.1-mini",
10 messages: [{ role: "user", content: "Summarize this ticket." }],
11 stream: true,
12});

Your OpenAI configurationEverything else, the model name included, stays as it is

Reference

The details.

How credits, billing and your data work. Anything else, ask @avenrotech on X.

What does one credit buy?
One US dollar of usage. Requests spend credits per token at the posted price of the model you call, and the models page lists every price.
How is a single request billed?
When it starts, its largest possible cost is set aside from your balance. When it ends, you pay for the tokens it used, rounded up to the next billionth of a dollar, and the rest is released. Non-streaming responses state the charge in the x-avenro-cost-usd header.
Which models are behind the key?
Each model is served through its provider's own API (Anthropic, OpenAI, Google, xAI and DeepSeek). GET /v1/models lists them with their prices.
Is it really the model I asked for?
Yes. Each request goes to the model you name and is never swapped for another. The x-avenro-execution-path response header says whether Avenro's own GPUs or the model's provider served it.
Are my prompts stored?
No. Prompts and outputs pass through memory and are never written to Avenro's databases or logs; only billing records are kept. The company serving a model handles the request under its own policy. Data handling
What changes in my code?
Two values: the base URL and the API key. The API follows OpenAI's Chat Completions format, so the OpenAI SDKs, the Vercel AI SDK and plain HTTP clients work unchanged.
How do I pay?
From Billing in the dashboard, in USDG or USDC on Robinhood Chain, sent from your own wallet, from $1. How top-ups work
Can I cap what a key spends?
Yes. Give any key a spending limit, and requests that would take it past the limit are refused, so a leaked key can only spend what its limit allows.
Do credits expire?
No, not while your account is open. There's no subscription, plan or minimum spend.
What does a failed request cost?
Nothing, if it fails before producing output. If a stream breaks partway through, you pay for the input and the output already delivered.
Where do I see what I've spent?
In the dashboard: spend by day and by model, and every request with its tokens and cost, filterable by key and model.
Can I get a refund?
Unused purchased credits can be refunded, less payment fees we can't recover. Ask @avenrotech on X; the Terms have the details.
Build

Build on one endpoint.

Create an account, add credits and send your first request. Keys, limits and usage live in your dashboard.

Pay with USDG or USDC on Robinhood Chain. One credit is one US dollar of usage, and credits don't expire while your account is open.

$ curl https://api.avenro.tech/v1/models
{  "object": "list",  "data": [    {      "id": "claude-fable-5-1",      "object": "model",      "owned_by": "Anthropic",      "pricing": { "input": "6", "output": "30", … }    },    {      "id": "claude-opus-5",      "object": "model",      "owned_by": "Anthropic",      "pricing": { "input": "3", "output": "15", … }    },    … 7 more  ]}