Live · auto-refresh 8s
OpenAI-compatible · /v1/chat/completions

One endpoint.
A growing AI catalog.

Chat, reason, call tools, and generate images through Cloudflare Workers AI. Explore the latest GLM, Kimi, Qwen, and DeepSeek models with a shared API key and usage dashboard.

Loading model catalog…
Base URL
https://…
Auth
Bearer token — ask the admin for a key

Live usage

Spend against the shared budget

Estimated lifetime usage. Requests stop when the recorded total reaches the cap; concurrent requests can exceed it. This is separate from Cloudflare billing.

Spent
Remaining
of $ total
Requests
across all users
Status
last activity —

Spend over time

Last 24 hours · hourly buckets

Spend by model

Lifetime · top consumers

Catalog

Available models

Use a short alias or full model ID. Chat prices are USD per million tokens; image prices are per 1024 × 1024 output at the listed steps.

Checking catalog version…

Model / provider Capabilities / context $ / 1M in $ / 1M cached $ / 1M out Spent Requests
loading…

Cloudflare platform notes

Know what each model needs

“Paid access” means Workers Paid or prepaid AI Gateway credits are required on the hosting account. A configured provider is not a guarantee of quota or model availability.

Daily allocation

Workers AI includes 10,000 Neurons per day. Paid overage is $0.011 per 1,000 Neurons. This dashboard estimates model usage before account credits and free allocations.

Official pricing ↗

Rate limits

Limits depend on the model and billing method. Cloudflare can return HTTP 429 for quota or capacity, and HTTP 403 when paid access is required.

Current limits ↗

Model migrations

Older Llama 3 / 3.1 variants and Gemma 3 were retired on May 30, 2026. Kimi K2.5 moved to K2.6 with different pricing. Use explicit model IDs when migrating.

Deprecation details ↗

Quickstart

Plug in from any language

Anywhere you can configure an OpenAI base_url + API key, you can use this. Pick your stack: