Alternatives

Top Cloudflare AI Gateway Alternatives

Cloudflare AI Gateway adds logs, caching and routing in front of providers you already use. If what you need is the models themselves, especially GLM, Kimi, Qwen or MiniMax, a gateway is only half the answer. These are the options to compare.

Facts checked on

On this page

What is Cloudflare AI Gateway?

Cloudflare AI Gateway is a managed gateway on Cloudflare鈥檚 network. You send requests to gateway.ai.cloudflare.com instead of the provider, and Cloudflare adds analytics, logging, caching, rate limiting, retries and fallbacks. It is available on all Cloudflare plans.

It authenticates upstream in three ways, in this order: a provider key on the request, a key stored in Cloudflare (BYOK), or Unified Billing. Unified Billing lets you prepay Cloudflare credits for Workers AI, OpenAI, Anthropic, Google AI Studio, Google Vertex AI, xAI and Groq.

An OpenAI-compatible endpoint takes model names in provider/model form. Any HTTPS API can be added as a custom provider.

Why teams look for Cloudflare AI Gateway alternatives

Most providers still need your own key
Unified Billing covers six third-party providers. For everything else you open an account with the vendor, fund it and store its key in Cloudflare.
Chinese model coverage is thin
DeepSeek is the only Chinese vendor among the built-in providers of the OpenAI-compatible endpoint. GLM, Kimi, Qwen and MiniMax need custom providers and upstream keys you obtain yourself.
A fee on prepaid credit
Unified Billing charges 5% on credit purchases: a $100 top-up costs $105.
Log pricing changed
Accounts created on or after September 24, 2026 follow Workers Logs pricing for gateway logs, while older accounts keep the legacy limits.

How to choose

Start with what you want Cloudflare AI Gateway to do. If you need logs, caching and rate limits over providers you already pay, the closest replacements are other gateways: LiteLLM if you want to self-host, Portkey or Vercel for a managed product, Kong if you already run it.

If you mainly want access to models Cloudflare cannot bill for, a gateway does not solve that. GLM, Kimi, Qwen and MiniMax need an account with each vendor, or a hosted endpoint such as A2Agent that sells all five behind one key.

Compare how each option charges. Cloudflare adds 5% to Unified Billing credit purchases. OpenRouter adds 5.5% to card top-ups. Vercel states no token markup but passes on payment-processing fees. A2Agent charges no top-up fee.

You can also keep Cloudflare and add A2Agent as a custom provider. Cloudflare keeps the logs and limits, and A2Agent supplies the models. The last section shows the setup.

Top Cloudflare AI Gateway Alternatives at a glance

PlatformFocusKey featuresIdeal for
A2AgentA hosted API endpoint that sells access to Chinese models: GLM, Kimi, DeepSeek, Qwen and MiniMax. You call it with one A2Agent key and pay per token. There is no gateway to deploy and no provider account to open.OpenAI Chat Completions and Responses, Anthropic Messages and Gemini protocols; per-key spend limits, expiry and IP allowlists; usage records in the dashboardTeams that want GLM, Kimi, DeepSeek, Qwen and MiniMax behind one key without running a gateway or opening accounts with each vendor
LiteLLMAn open-source Python SDK and self-hosted proxy that give one OpenAI-style interface to many providers.Virtual keys, per-key and team budgets, rate limits, load balancing, fallbacks, caching, admin UITeams that want a free, MIT-licensed gateway in their own infrastructure and can run PostgreSQL and Redis
OpenRouterA hosted API that sells access to models from many vendors through one key and one bill, with an OpenAI-compatible endpoint.One key for many vendors, prepaid credits, usage analytics, optional BYOKTeams that want one hosted key across models from many vendors, including OpenAI, Anthropic and Google
PortkeyAn AI gateway and control platform: routing, observability, guardrails and prompt management in front of providers you already pay. Palo Alto Networks completed its acquisition of Portkey on May 29, 2026.Routing and fallback, load balancing, simple and semantic caching, guardrails, budgets, MCP gatewayTeams that bring their own provider keys and want guardrails, prompt management and observability in one managed product
Vercel AI GatewayA managed gateway that also sells model access through AI Gateway Credits. Your app does not need to run on Vercel.OpenAI Chat Completions and Responses, Anthropic Messages, AI SDK; fallbacks, provider ordering, request logs, budgetsTeams on Vercel or the AI SDK that want credits, logs and budgets in one place
Kong AI GatewayAI plugins for the Kong API gateway, self-hosted or run through Kong Konnect with data planes in your environment.Semantic caching, cost-based rate limiting, prompt guardrails, PII redaction, MCP support, OpenTelemetryOrganizations that already run Kong for API management

In-depth look at each alternative

01A2Agent

A hosted API endpoint that sells access to Chinese models: GLM, Kimi, DeepSeek, Qwen and MiniMax. You call it with one A2Agent key and pay per token. There is no gateway to deploy and no provider account to open.

Strengths

  • One key reaches five Chinese model families. You do not open an account, pass verification or hold a key with each vendor.
  • The same key works over OpenAI Chat Completions, OpenAI Responses, Anthropic Messages and Gemini, so existing SDKs and coding tools connect by changing the base URL.
  • Each key can carry a USD spend quota, 5-hour, daily and weekly spend limits, an expiry date and an IP allowlist.
  • Pay as you go from a $1 top-up with no top-up fee. Credits do not expire while your account is active.

Limitations

  • A2Agent is a model provider, not a gateway product. It has no guardrails, prompt management or semantic caching, and its catalog covers the five Chinese model families only.
  • You cannot bring your own vendor keys. Every request is billed from your A2Agent balance.
  • There is no free tier.

Pricing

Per-token billing from a prepaid balance. The minimum top-up is $1 with no top-up fee, and top-ups can earn bonus credit. Purchased credit is refundable within 7 days; bonus credit is not. Per-model rates are listed on the pricing page.

Ideal for

Teams that want GLM, Kimi, DeepSeek, Qwen and MiniMax behind one key without running a gateway or opening accounts with each vendor

02LiteLLM

An open-source Python SDK and self-hosted proxy that give one OpenAI-style interface to many providers.

Strengths

  • Free and MIT-licensed. You keep traffic in your own infrastructure.
  • Virtual keys, budgets, rate limits and guardrails are in the open-source proxy.
  • Deploy with Docker, Helm, or Terraform modules for AWS ECS and Google Cloud Run.

Limitations

  • Self-hosted only; the pricing page lists no hosted plan. Production needs PostgreSQL, and Redis once you run more than one instance.
  • SSO beyond 5 users, audit logs, RBAC, IP allowlists and several other controls need an Enterprise license.
  • You bring your own provider keys.

Pricing

Open source is $0 under the MIT license. Enterprise is custom-priced by deployment size and is also self-hosted.

Ideal for

Teams that want a free, MIT-licensed gateway in their own infrastructure and can run PostgreSQL and Redis

03OpenRouter

A hosted API that sells access to models from many vendors through one key and one bill, with an OpenAI-compatible endpoint.

Strengths

  • Nothing to deploy. Point an OpenAI SDK at https://openrouter.ai/api/v1.
  • Covers models from many vendors, not only Chinese ones.
  • Bring-your-own-key usage is free up to $25,000 a month on pay-as-you-go.

Limitations

  • The fee sits on credit purchases: 5.5% by card with a $0.80 minimum, or 5% in crypto.
  • BYOK usage above the monthly free allowance carries a 5% fee.

Pricing

Prepaid credits. OpenRouter states no markup on inference prices and charges the fee when you buy credits. Free models are limited to 50 requests a day without purchased credits.

Ideal for

Teams that want one hosted key across models from many vendors, including OpenAI, Anthropic and Google

04Portkey

An AI gateway and control platform: routing, observability, guardrails and prompt management in front of providers you already pay. Palo Alto Networks completed its acquisition of Portkey on May 29, 2026.

Strengths

  • Managed cloud, an MIT-licensed open-source gateway, and private deployment for Enterprise.
  • Simple and semantic caching, guardrails and prompt management are built in.
  • A free Developer plan to start.

Limitations

  • You bring your own provider keys. Portkey does not sell model access, so each vendor account and bill stays with you.
  • The free plan records 10,000 logs a month and keeps them for 3 days. Private deployment is Enterprise-only.
  • Portkey is now the core gateway of Palo Alto Networks Prisma AIRS. Check current plan terms before you commit.

Pricing

Developer is free. Production is $49 a month with 100,000 logs and $9 per extra 100,000. Enterprise is custom. The open-source gateway is free to self-host.

Ideal for

Teams that bring their own provider keys and want guardrails, prompt management and observability in one managed product

05Vercel AI Gateway

A managed gateway that also sells model access through AI Gateway Credits. Your app does not need to run on Vercel.

Strengths

  • Works without provider keys: requests use Vercel credentials and draw on your credits by default.
  • Vercel states no markup and no platform fee on tokens, including with BYOK.
  • Budgets per team, project, key or member.

Limitations

  • BYOK needs the paid tier and purchased credits. A failed BYOK request is retried on Vercel credentials and charged to your credits.
  • Budgets are soft caps and do not count BYOK spend.
  • Some features cost extra, such as custom reporting, team-wide provider allowlists and trace drains.

Pricing

Prepaid AI Gateway Credits with no token markup; you pay payment-processing fees. A monthly free credit covers a subset of models until you buy credits.

Ideal for

Teams on Vercel or the AI SDK that want credits, logs and budgets in one place

06Kong AI Gateway

AI plugins for the Kong API gateway, self-hosted or run through Kong Konnect with data planes in your environment.

Strengths

  • Handles AI traffic inside an existing Kong deployment.
  • The ai-proxy plugin is open source and has built-in providers including DeepSeek and Alibaba Cloud DashScope.
  • Governance features such as PII redaction and cost-based rate limiting.

Limitations

  • You bring your own provider keys and run the data plane.
  • Advanced load balancing and cross-provider failover (ai-proxy-advanced) are Enterprise-only and need Kong Gateway 3.8 or newer.

Pricing

Konnect Plus is billed per gateway, from $25 a month for serverless, with 1 million API requests included. The AI Gateway on Plus includes 5 LLM models, then $100 a month per extra model. Enterprise is custom, and a 30-day trial is available.

Ideal for

Organizations that already run Kong for API management

Why A2Agent is different

The models come with the endpoint
A gateway only forwards traffic. You still sign up with each vendor, pass its checks, fund it and store its key. A2Agent sells the model access itself, so one account replaces five vendor accounts.
Nothing to operate
There is no proxy to deploy, no database to back up and no Redis to size. A2Agent runs failover across its upstream accounts, health checks and rate limiting on its side.
Every major client protocol
OpenAI Chat Completions and Responses, Anthropic Messages and Gemini are all served, so Claude Code, Codex, OpenCode, Cursor and SDKs connect with a base URL change. The integrations page has setup guides for each tool.
Small, fee-free top-ups
Start with $1. There is no top-up fee and no subscription, and credits do not expire while your account is active.
Key-level spend control
Quotas, rolling spend limits, expiry and IP allowlists are set per key in the dashboard, without writing gateway policy.

See every model鈥檚 per-token price next to its official list price on the models page.

Use A2Agent with Cloudflare AI Gateway

If you keep Cloudflare AI Gateway, add A2Agent as a custom provider in Compute & AI > AI Gateway > Custom Providers, with the base URL below and slug a2agent. The path after custom-a2agent/ is appended to the base URL. Store your A2Agent key with BYOK rather than sending it on each request.

shell
# Custom provider base URL: https://a2agent.me/v1
curl https://gateway.ai.cloudflare.com/v1/$ACCOUNT_ID/$GATEWAY_ID/custom-a2agent/chat/completions \
  -H 'Authorization: Bearer YOUR_A2AGENT_API_KEY' \
  -H "cf-aig-authorization: Bearer $CF_AIG_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"model": "YOUR_MODEL_ID", "messages": [{"role": "user", "content": "Reply with OK."}]}'
Cloudflare AI Gateway documentation

Frequently asked questions

Can Cloudflare AI Gateway bill me for GLM, Kimi, Qwen or MiniMax?

Not through Unified Billing, which covers Workers AI, OpenAI, Anthropic, Google AI Studio, Google Vertex AI, xAI and Groq. For other providers you bring your own key, or add an endpoint such as A2Agent as a custom provider and pay that endpoint directly.

Can I keep Cloudflare AI Gateway in front of A2Agent?

Yes. Add A2Agent as a custom provider with an HTTPS base URL, store your A2Agent key with BYOK, and send requests through custom-a2agent/. Cloudflare keeps its logs, caching and rate limits, and A2Agent bills the model usage.

Is Cloudflare AI Gateway free?

Analytics, caching and rate limiting are free on all plans. Log storage follows Workers Logs pricing for accounts created on or after September 24, 2026. Unified Billing adds 5% to credit purchases, and guardrails are billed as Workers AI inference.

Is A2Agent a gateway like Cloudflare AI Gateway?

No. A gateway sits in front of providers and forwards requests with your keys. A2Agent is the provider side: it sells access to GLM, Kimi, DeepSeek, Qwen and MiniMax, and you call it with an A2Agent key. You can also put A2Agent behind a gateway as one upstream.

Do I need accounts with Zhipu, Moonshot, DeepSeek, Alibaba or MiniMax?

Not with A2Agent. One A2Agent account and key reach all five families. With a bring-your-own-key gateway you open and pay each vendor account yourself.

Which APIs does A2Agent support?

OpenAI Chat Completions, OpenAI Responses, Anthropic Messages and Gemini generateContent, plus a model list at /v1/models. Most OpenAI or Anthropic SDKs connect by changing the base URL and key.

How is A2Agent billed?

Per token from a prepaid balance. The minimum top-up is $1, there is no top-up fee, and credits do not expire while your account is active. There is no subscription requirement and no free tier.

Can I limit what each key spends?

Yes. Each key can have a USD quota, 5-hour, daily and weekly spend limits, an expiry date and an IP allowlist. Usage for every request shows in the dashboard.

Where do I compare model prices?

The pricing page lists every model鈥檚 input, output and cache rates next to the official list price, and compares them with OpenRouter model by model.

Get started

Create an account, top up from $1 and create an API key. Point any OpenAI, Anthropic or Gemini client at A2Agent and pick a model.

Sources

Checked on 2026-09-28. Vendors change plans and features often. Check each official page before you decide.