Free tier 10M tokens every month · no card required

Stop overpaying for AI.
100M tokens for $4.50.

Claude, GPT, Gemini, DeepSeek and 29 more models through one OpenAI-compatible API. Start with 10M free tokens every month, or go unlimited for $50.

Free $0

10M tokens / mo

No card required

Starter $4.50/mo

100M tokens / mo

All 33 models included

Unlimited $50/mo

∞ no monthly cap

For heavy agent workloads

ChatGPT Claude Gemini Llama DeepSeek Kimi
33 models on every plan — one key

One API for the models you already know

Quickstart

Swap your base URL. Keep your code.

Use the official OpenAI SDKs for Python, Node, Go, or .NET — or any tool that accepts a custom endpoint. Switching models is a one-string change.

  • Streaming over SSE with hour-long connections
  • Tool calling, including parallel calls
  • Vision input and reasoning models
  • One key for all 33 models
Read the full quickstart guide
from openai import OpenAI

client = OpenAI(
    base_url="https://radoxai.com/v1",
    api_key="sk-xxxxxxxxxxxxxxxxxxxxxx",
)

stream = client.chat.completions.create(
    model="claude-opus-4.8",
    messages=[{"role": "user", "content": "Hello!"}],
    stream=True,
)

for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://radoxai.com/v1",
  apiKey: "sk-xxxxxxxxxxxxxxxxxxxxxx",
});

const stream = await client.chat.completions.create({
  model: "claude-opus-4.8",
  messages: [{ role: "user", content: "Hello!" }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}
curl https://radoxai.com/v1/chat/completions \
  -H "Authorization: Bearer sk-xxxxxxxxxxxxxxxxxxxxxx" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4.8",
    "messages": [{"role": "user", "content": "Hello!"}],
    "stream": true
  }'
Streaming Tool calling Vision Reasoning

Integrations

Drop-in for the tools you already use

Any coding agent or app that accepts a custom OpenAI endpoint works out of the box. Paste the base URL and your key, pick a model, and you're done.

See setup guides

How it works

Live in three steps

No SDK to learn, no migration. If your code talks to OpenAI, it already talks to us.

  1. 1

    Create a key

    Sign up and generate an API key from your dashboard. Scope it to inference, models, or embeddings.

    API keysk-••••••••••••3f9a
  2. 2

    Point your client

    Set the base URL to our endpoint and paste the key. Everything else in your code stays the same.

    base_urlhttps://radoxai.com/v1
  3. 3

    Ship

    Track every request, token, and dollar in the dashboard. Swap models without touching your integration.

    status● 200 OK · streaming

Platform

Built for real applications

Everything a production integration needs — not a thin proxy.

Read the documentation →

True streaming

Server-sent events with keep-alive pings and hour-long connections, so long agent sessions never get cut off mid-task.

Tool calling

Full function-calling round trips, including parallel calls and streamed argument deltas — what coding agents run on.

Vision input

Send screenshots, mockups, and diagrams as base64 or URLs to any vision-capable model, with per-model capability discovery.

Reasoning models

Chain-of-thought output arrives in a separate field, so thinking never contaminates your answer. Token budgets handled for you.

Usage analytics

Every request, token, latency figure, and cent — per key and per model. Real numbers, no estimates.

Global payments

Top up with Paddle, Binance Pay, or direct cryptocurrency. Pay smoothly from anywhere in the world.

Models

Models to start with

Lightweight and fast, or deeply capable — same API either way.

Browse all 33 models
Gemma 2 2B

Gemma 2 2B

gemma-2-2b
chat

Lightweight open model by Google. Fast and efficient for everyday tasks.

⚡ Streaming 📄 JSON
Context
8K
In /1M
$0.115
Out /1M
$0.23
DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731

deepseek-v4-flash-0731
chat

DeepSeek fast open model. Very affordable with solid reasoning, 304B params.

🧠 Reasoning 🔧 Tools ⚡ Streaming 📄 JSON
Context
1049K
In /1M
$0.115
Out /1M
$0.115
GPT-5.6 Luna

GPT-5.6 Luna

gpt-5.6-luna
chat

OpenAI efficient model. Great value with strong reasoning, 290c/s speed.

🧠 Reasoning 🖼️ Vision 🔧 Tools ⚡ Streaming 📄 JSON
Context
1050K
In /1M
$0.357
Out /1M
$0.357
DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813

deepseek-v4-pro-0813
chat

DeepSeek open-source pro model. Excellent reasoning and coding, 1.6T params.

🧠 Reasoning 🔧 Tools ⚡ Streaming 📄 JSON
Context
1049K
In /1M
$0.552
Out /1M
$0.552
Qwen3.8-27B

Qwen3.8-27B

qwen3.8-27b
chat

Alibaba Qwen compact open model. Efficient 27.8B params with solid capability.

🧠 Reasoning 🔧 Tools ⚡ Streaming 📄 JSON
Context
262K
In /1M
$0.575
Out /1M
$0.575
Seed 2.1 Turbo

Seed 2.1 Turbo

seed-2.1-turbo
chat

ByteDance Seed turbo model. Fast with solid reasoning and coding.

🧠 Reasoning 🖼️ Vision 🔧 Tools ⚡ Streaming 📄 JSON
Context
262K
In /1M
$0.92
Out /1M
$0.92

Pricing

100M tokens for $4.50. Or unlimited for $50.

Spend your allowance on Claude Fable 5 or anything else in the catalog — all 33 models are on every plan. Bigger package, cheaper per token.

Plan Tokens / month Price Per 1M tokens
Free 10M Free —
Starter 100M $4.50/mo $0.045
Basic 200M $8.50/mo $0.0425
Pro Best value 350M $12.50/mo $0.0357
Scale 500M $15.50/mo $0.031
Unlimited No cap Unlimited $50/mo —
All 33 models on every plan Allowance resets monthly Overage billed at published rates Pay by card or crypto

FAQ

Common questions

Can't find what you're looking for? The documentation covers setup, limits, and billing in detail.

Browse the docs

Start free with 10M tokens every month. Paid plans begin at $4.50/month for 100M tokens and scale up to unlimited tokens at $50/month. Every plan unlocks all 33 models including Claude Fable 5 — you are only buying token volume, never feature access.

No. Every plan, including the free tier, reaches all 33 models — Claude Fable 5 included. Plans differ only in monthly token volume and per-minute rate limits.

Yes — completely. Point any official OpenAI SDK (Python, Node, Go, .NET) at our base URL and pass your key. Request and response shapes match the OpenAI spec, including streaming chunks, tool calls, and error envelopes.

Anything that accepts a custom OpenAI-compatible endpoint: Kilo Code, Cline, Cursor, opencode, Claude Code, Continue, and Aider are all verified. Set the provider to "OpenAI Compatible", paste the base URL and key, and pick a model.

Tokens cover both your prompt and the model's reply, combined. Your plan allowance is spent first and resets monthly; anything beyond it bills against your account balance at each model's published per-token rate.

Yes. Streaming uses server-sent events with keep-alive pings and connections that stay open for up to an hour — suited to long agent sessions. Tool calling supports full round trips, parallel calls, and streamed argument deltas.

Requests fall back to your balance. With an empty balance you get HTTP 402 and a clear error before the request reaches a model, so you are never billed unexpectedly and never cut off mid-response. Top up or upgrade and you resume immediately.

We accept Credit/Debit Cards, Apple Pay, and Google Pay (via Paddle), Binance Pay, and direct Cryptocurrency (USDT, etc.). Orders are processed and verified automatically.

Ready in about a minute

Create a key, change your base URL, and start shipping. 10M free tokens every month — no card required.

# Point any OpenAI client at

https://radoxai.com/v1

10M

free / month

$4.50

for 100M

33

models