True streaming
Server-sent events with keep-alive pings and hour-long connections, so long agent sessions never get cut off mid-task.
Claude, GPT, Gemini, DeepSeek and 29 more models through one OpenAI-compatible API. Start with 10M free tokens every month, or go unlimited for $50.
10M tokens / mo
No card required
100M tokens / mo
All 33 models included
∞ no monthly cap
For heavy agent workloads
One API for the models you already know
Quickstart
Use the official OpenAI SDKs for Python, Node, Go, or .NET — or any tool that accepts a custom endpoint. Switching models is a one-string change.
from openai import OpenAI
client = OpenAI(
base_url="https://radoxai.com/v1",
api_key="sk-xxxxxxxxxxxxxxxxxxxxxx",
)
stream = client.chat.completions.create(
model="claude-opus-4.8",
messages=[{"role": "user", "content": "Hello!"}],
stream=True,
)
for chunk in stream:
print(chunk.choices[0].delta.content or "", end="")
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://radoxai.com/v1",
apiKey: "sk-xxxxxxxxxxxxxxxxxxxxxx",
});
const stream = await client.chat.completions.create({
model: "claude-opus-4.8",
messages: [{ role: "user", content: "Hello!" }],
stream: true,
});
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
curl https://radoxai.com/v1/chat/completions \
-H "Authorization: Bearer sk-xxxxxxxxxxxxxxxxxxxxxx" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4.8",
"messages": [{"role": "user", "content": "Hello!"}],
"stream": true
}'
Integrations
Any coding agent or app that accepts a custom OpenAI endpoint works out of the box. Paste the base URL and your key, pick a model, and you're done.
See setup guidesHow it works
No SDK to learn, no migration. If your code talks to OpenAI, it already talks to us.
Sign up and generate an API key from your dashboard. Scope it to inference, models, or embeddings.
Set the base URL to our endpoint and paste the key. Everything else in your code stays the same.
Track every request, token, and dollar in the dashboard. Swap models without touching your integration.
Platform
Everything a production integration needs — not a thin proxy.
Server-sent events with keep-alive pings and hour-long connections, so long agent sessions never get cut off mid-task.
Full function-calling round trips, including parallel calls and streamed argument deltas — what coding agents run on.
Send screenshots, mockups, and diagrams as base64 or URLs to any vision-capable model, with per-model capability discovery.
Chain-of-thought output arrives in a separate field, so thinking never contaminates your answer. Token budgets handled for you.
Every request, token, latency figure, and cent — per key and per model. Real numbers, no estimates.
Top up with Paddle, Binance Pay, or direct cryptocurrency. Pay smoothly from anywhere in the world.
Models
Lightweight and fast, or deeply capable — same API either way.
gemma-2-2b
Lightweight open model by Google. Fast and efficient for everyday tasks.
deepseek-v4-flash-0731
DeepSeek fast open model. Very affordable with solid reasoning, 304B params.
gpt-5.6-luna
OpenAI efficient model. Great value with strong reasoning, 290c/s speed.
deepseek-v4-pro-0813
DeepSeek open-source pro model. Excellent reasoning and coding, 1.6T params.
qwen3.8-27b
Alibaba Qwen compact open model. Efficient 27.8B params with solid capability.
seed-2.1-turbo
ByteDance Seed turbo model. Fast with solid reasoning and coding.
Pricing
Spend your allowance on Claude Fable 5 or anything else in the catalog — all 33 models are on every plan. Bigger package, cheaper per token.
| Plan | Tokens / month | Price | Per 1M tokens |
|---|---|---|---|
| Free | 10M | Free | — |
| Starter | 100M | $4.50/mo | $0.045 |
| Basic | 200M | $8.50/mo | $0.0425 |
| Pro Best value | 350M | $12.50/mo | $0.0357 |
| Scale | 500M | $15.50/mo | $0.031 |
| Unlimited No cap | Unlimited | $50/mo | — |
FAQ
Can't find what you're looking for? The documentation covers setup, limits, and billing in detail.
Browse the docsStart free with 10M tokens every month. Paid plans begin at $4.50/month for 100M tokens and scale up to unlimited tokens at $50/month. Every plan unlocks all 33 models including Claude Fable 5 — you are only buying token volume, never feature access.
No. Every plan, including the free tier, reaches all 33 models — Claude Fable 5 included. Plans differ only in monthly token volume and per-minute rate limits.
Yes — completely. Point any official OpenAI SDK (Python, Node, Go, .NET) at our base URL and pass your key. Request and response shapes match the OpenAI spec, including streaming chunks, tool calls, and error envelopes.
Anything that accepts a custom OpenAI-compatible endpoint: Kilo Code, Cline, Cursor, opencode, Claude Code, Continue, and Aider are all verified. Set the provider to "OpenAI Compatible", paste the base URL and key, and pick a model.
Tokens cover both your prompt and the model's reply, combined. Your plan allowance is spent first and resets monthly; anything beyond it bills against your account balance at each model's published per-token rate.
Yes. Streaming uses server-sent events with keep-alive pings and connections that stay open for up to an hour — suited to long agent sessions. Tool calling supports full round trips, parallel calls, and streamed argument deltas.
Requests fall back to your balance. With an empty balance you get HTTP 402 and a clear error before the request reaches a model, so you are never billed unexpectedly and never cut off mid-response. Top up or upgrade and you resume immediately.
We accept Credit/Debit Cards, Apple Pay, and Google Pay (via Paddle), Binance Pay, and direct Cryptocurrency (USDT, etc.). Orders are processed and verified automatically.
Create a key, change your base URL, and start shipping. 10M free tokens every month — no card required.
# Point any OpenAI client at
https://radoxai.com/v1
10M
free / month
$4.50
for 100M
33
models