Free AI APIs in 2026: Quotas, Catches and Who Each One Fits

A2Agent Team · 2026-10-08T00:00:00Z

Almost every model provider now advertises a free tier. Sign up for a few and the differences show up fast. Some want a credit card. Some quotas are gone by lunchtime. Some terms change from month to month, and some free tiers train on whatever you send them.

This guide covers six international platforms, one local tool and two platforms based in China. For each one you'll find the quota, the trade-offs, how to make a first call and who it suits. We checked the quotas on September 30, 2026. Providers change them often, so trust each provider's console over any number here.

Five kinds of "free"

In this guide, "free" means you can call a model's API without paying. That covers five different arrangements:

  • Rate-limited. Free with no end date, capped by requests per minute, requests per day or tokens per day. The quota never runs out as long as you stay under the cap. Gemini, Groq, SambaNova, Cerebras, the OpenRouter free pool and Zhipu work this way.
  • Trial credit. A one-time balance billed per token, gone once it's spent. It suits a short burst of testing. Example: Alibaba Model Studio's 90-day credit.
  • Rotating pool. The quota is permanent, but the set of free models keeps changing. You can stay free for a long time as long as your project isn't tied to one specific model. Example: OpenRouter's :free models.
  • User-pays. Free for the developer, with no developer-side bill. AI usage is charged to each end user's own account, and new accounts usually come with starter credit. Example: Puter.
  • Local. The model runs on your own machine, so there are no request or token limits and it works offline. You pay for it in hardware. Example: LM Studio.

What you pay instead of money

The terms for Gemini's free tier and OpenRouter's free models allow your inputs and outputs to be used for training, so keep sensitive data out of them. Some platforms want personal details instead: Zhipu asks for a phone number, and Puter needs each of your users to sign in to a Puter account. LM Studio costs nothing in API fees, but you supply the GPU and memory.

Free quotas are sized for learning and development. They won't carry production traffic, and once a project grows you will pay.

The map at a glance

Platform Category Free quota (highlights) Best for Main catch
A2Agent Aggregator $5 credit at signup; first $10 top-up earns another $10; model prices up to 50% off Everyday calls, and cheap continued use after free quotas run out Apart from new-user offers, promotions for existing users change from time to time
Google Gemini Official developer platform No card; static limits no longer published, Flash models roughly 10 to 15 requests/min A first AI project, multimodal work Free-tier data used for training (except in the EEA, Switzerland and the UK)
Groq Fast inference 30 requests/min, 1,000 requests/day, 200K tokens/day, no card Low-latency apps, existing OpenAI code Few models, fast turnover; old tutorials break
SambaNova Fast inference Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day (per model), no card Running 405B-class open models for free Only 20 requests a day
Cerebras Fast inference About 1M tokens/day, no card; published request limits disagree Long documents, batch text jobs Free model lineup changes often
OpenRouter Aggregator Free models: 20 requests/min, 50 requests/day; a one-time $10 top-up raises that to 1,000/day for good Comparing models side by side Rotating pool; requires a training/logging opt-in
Puter Aggregator Free for developers; AI usage billed to end users' accounts, new accounts get starter credit Web apps, frontend developers Undisclosed daily cap; browser only
LM Studio Local No limits; capability depends on your hardware Sensitive data, offline or intranet work You need the GPU and RAM
Zhipu BigModel Official platform (China) GLM-4-Flash free permanently, plus new-user credit Learners in China, high-volume agent workloads Bonus-credit terms per Zhipu's official rules
Alibaba Model Studio Official platform (China) 1M tokens per model, over 70M in total, valid 90 days Trying Chinese models side by side Expires after 90 days; Beijing region only

Quotas come from each provider's documentation and community verification pages from September 2026. Google and Cerebras in particular adjust limits and free models often, so check the console before you plan around a number.

How the platforms differ

Official developer platforms: Gemini, Zhipu, Alibaba Model Studio

These are the model makers' own developer sites, where you register and call their models directly. The free tiers are aimed at individual developers who are trying things out. Quotas are fairly generous (Gemini's multimodal access, Zhipu's permanently free model), and the terms usually say your inputs and outputs can be used to train their models.

Fast inference platforms: Groq, SambaNova, Cerebras

These companies serve open-weight models on their own inference chips at several hundred to several thousand tokens per second, which helps when you're developing something where latency matters. They meter differently. Groq counts requests, SambaNova allows few requests but plenty of tokens per request, and Cerebras counts total tokens.

Aggregators: OpenRouter and Puter

Third-party platforms that give you one entry point to models from many vendors. OpenRouter puts 500+ models behind one key, 20+ of them free. Puter targets browser apps: developers integrate at no cost, and the AI bill lands on the end user's account.

Local: LM Studio

The model and the compute both live on your computer, and nothing goes through a server. You don't need an account, there's no quota, and your data stays on your machine. How strong a model you can run depends on your GPU and memory.

One API format across all nine

All nine free options accept the OpenAI API format. Once you have working openai library code, you can move between them by changing the base URL, the key and the model name, which is why the "keep a backup provider" tip further down works. LM Studio additionally speaks the Anthropic format (/v1/messages), so tools like Claude Code can talk to a local model directly. If you're new to switching providers this way, our OpenAI SDK integration guide shows the two-line change.

The platforms, one by one

Google Gemini (AI Studio)

Official developer platform · rate-limited

Gemini's official API, offered through Google AI Studio. The free tier needs no credit card and handles text, images and code, which makes it the most well-rounded free option for everyday use.

Free quota

  • No credit card; sign up and start
  • Flash models at roughly 10 to 15 requests per minute
  • Some accounts see per-model tokens-per-minute limits in the millions
  • Since September 2026 Google no longer publishes a static limits table; the console shows your live limits

Strengths

  • Handles text, images and code, with the strongest multimodal capability in the free pool
  • Lowest barrier: no card and no phone verification
  • Fairly generous limits, and Flash models handle volume well
  • Mature SDKs and a smooth path to paid

Drawbacks

  • Limits aren't transparent, and the figures quoted in the press contradict each other
  • Free-tier data may be used for training (except in the EEA, Switzerland and the UK), so don't send sensitive data
  • Pro models have mostly left the free tier, and enabling billing on a project replaces that project's free tier
  • Hardest to reach from mainland China (see the China section below)

Getting started

  1. Go to ai.google.dev and sign in to AI Studio with a Google account.
  2. Click "Get API key". You don't need a card, and it takes about a minute.
  3. Use the OpenAI-compatible endpoint https://generativelanguage.googleapis.com/v1beta/openai/ or Google's own SDK. Example model: gemini-2.5-flash.
  4. Check the rate-limit panel in AI Studio for your account's current limits before you start.

Best for

  • Students and career changers building a first AI project. No card or verification, and one platform covers text, images and code while you learn.
  • Indie developers who need mixed image and text processing. Nothing else in the free pool matches Gemini's vision capability.
  • Developers with reliable international connectivity building content or automation tools. The quota is large enough to keep a small tool running long term, which most rate-limited tiers can't do.

Groq

Fast inference · rate-limited

Groq serves open-weight models (OpenAI's gpt-oss series, Qwen variants and others) on its own LPU chips. The free tier runs at several hundred tokens per second, and for chat or voice apps it feels close to a paid product.

Free quota (per organization, no card)

  • 30 requests/min and 1,000 requests/day
  • 8,000 tokens/min and 200,000 tokens/day
  • Models: gpt-oss-120b, gpt-oss-20b, qwen3.6-27b

Strengths

  • The fastest of the free tiers here
  • Fully OpenAI-compatible: change base_url and the key in your existing code and it runs
  • 200K tokens a day is generous for a rate-limited tier, enough for a project in development
  • Email signup, no card

Drawbacks

  • Few models, and they rotate quickly. The Llama models were retired in August 2026, so code from older tutorials fails
  • Open-weight models only, nothing at the level of Gemini's closed flagships
  • The 200K daily token cap goes faster than you'd expect with long prompts and large contexts
  • Requests may queue at peak times; less stable than paid

Getting started

  1. Sign up with email at console.groq.com. No card needed.
  2. Create a key on the API Keys page, and check the Limits page for your live limits.
  3. Point your OpenAI code at https://api.groq.com/openai/v1 and use the model openai/gpt-oss-120b.
  4. The widely quoted "14,400 requests a day" is out of date. Plan for 1,000.

Best for

  • Indie developers building an MVP. 1,000 requests a day covers development, and debugging feels the same as on a paid tier.
  • Real-time chat and voice apps. Groq's free tier responds faster than any other here.
  • Anyone with OpenAI-format code who wants to run it for free, since only two config lines change.

SambaNova

Fast inference · rate-limited

SambaNova Cloud runs inference on its own RDU chips at speeds close to Groq's. Its free list includes Llama 3.1 405B, a flagship open model, while most free tiers stop at small and mid-size models.

Free quota (no card, counted per model)

  • Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day
  • Llama 3.1 405B: about 10 requests/min
  • DeepSeek-V3.1 and gpt-oss-120b are also free
  • New users get an extra one-time $5 credit

Strengths

  • Free access to 405B-class open models, which few free tiers offer
  • Fast inference on RDU chips
  • 200K tokens a day per model works out to about 10K tokens per request, which suits long tasks
  • No card, OpenAI-compatible, and rate-limit status comes back in the response headers

Drawbacks

  • Only 20 requests a day, which suits a few big jobs and rules out frequent chat
  • The lineup is mostly Llama, so the choice is narrow
  • Official documentation is vague about some limits, and sources disagree on the token cap

Getting started

  1. Register at cloud.sambanova.ai. No card needed.
  2. Create an API key in the console.
  3. Use the OpenAI-compatible endpoint https://api.sambanova.ai/v1.
  4. Try the model Meta-Llama-3.3-70B-Instruct, and read the rate-limit fields in the response headers to track what's left.

Best for

  • Anyone who wants to run a 405B-class open model for free. Hardly any other free tier has one.
  • Occasional heavy jobs such as long-document analysis, where about 10K tokens per request across 20 requests a day is a good fit.
  • A backup for Groq. Both are fast, but the models and quota structure differ, so each covers the other's gaps.

Cerebras

Fast inference · rate-limited

Cerebras runs inference on wafer-scale chips at several thousand tokens per second, so long outputs arrive almost instantly. Its free quota is measured in total tokens, about 1M a day, the largest token allowance of the nine.

Free quota

  • About 1,000,000 tokens/day
  • No card, free with no end date
  • Request limits are reported inconsistently: older sources say 30/min, 2026 sources say 5/min. Check the console
  • Current free models: gpt-oss-120b and GLM-4.7 (Llama has been phased out)

Strengths

  • Very fast on long outputs
  • The highest daily token allowance here, enough for real batch processing
  • 131K context for long documents
  • No card, OpenAI-compatible

Drawbacks

  • Published request limits contradict each other, so it's hard to plan around them
  • The free lineup changes quickly; Llama is gone, replaced by gpt-oss-120b and GLM-4.7
  • Unused tokens don't roll over, and the count resets daily

Getting started

  1. Register at cloud.cerebras.ai. No card needed.
  2. Create an API key in the console.
  3. Use the OpenAI-compatible endpoint https://api.cerebras.ai/v1.
  4. Check the current free model list before choosing a model, because it changes every few months. Examples: gpt-oss-120b, glm-4.7.

Best for

  • Long-document and batch text jobs. About 1M tokens a day is more than any other free tier here gives you.
  • Apps where users wait on long replies. At thousands of tokens per second, they barely wait at all.

OpenRouter

Aggregator · rate-limited + rotating pool

OpenRouter is a single gateway to 500+ models, so one account and one key reach almost every major model. More than 20 free models (IDs ending in :free) rotate through the catalog, including Llama 3.3 70B, Qwen3 Coder, GPT-OSS and Nemotron. It's most useful when you haven't decided which model to use yet.

Free quota

  • 20 requests/min and 50 requests/day
  • A one-time $10 top-up raises the daily cap to 1,000 requests permanently, even after the balance is spent
  • Free models cost $0 per token; the limit is the request count
  • The free router openrouter/free picks an available model for you

Strengths

  • One key for 500+ models, so comparing them costs almost nothing
  • A one-time $10 permanently raises the daily limit 20× (50 to 1,000 requests), the best-value one-off spend in this guide
  • The free router fails over automatically, so if one model is down it moves to another
  • Switching models means changing one parameter, so you aren't locked in

Drawbacks

  • Free models require you to enable training/logging, which puts the data terms in the same category as Gemini's
  • Failed requests count against the daily quota, so a retry loop can burn through 50 requests quickly
  • Free endpoints may have smaller context windows than paid ones, and upstream providers throttle at peak times
  • OpenRouter itself says free models aren't meant for production

Getting started

  1. Sign in at openrouter.ai with Google or GitHub.
  2. Create an API key on the Keys page.
  3. Use the endpoint https://openrouter.ai/api/v1 and add :free to a model ID to use the free quota, for example qwen/qwen3-coder:free.
  4. Without a top-up you get 50 requests a day. After one $10 top-up you get 1,000 a day for good.

Best for

  • Anyone who wants to compare models before paying. Send one prompt to 20+ free models and you'll know which fits your task before you buy a subscription.
  • Teams running model evaluations. With one interface and one meter, an evaluation script works across every model.
  • Developers who don't want to depend on one vendor. Your code talks to OpenRouter, and the vendors behind it can change.

Puter

Aggregator · user-pays

Puter (puter.com) is an open-source "internet operating system". Its Puter.js library lets a web app add AI in a few lines of code, with access to 400+ models from 70+ providers, including OpenAI, Anthropic, Gemini, Qwen and Nemotron as well as image, video and speech models. Developers integrate at no cost, and AI usage is charged to each end user's own Puter account.

Free quota

  • Developer side: free to integrate, no bill
  • User side: new accounts get starter credit
  • Community testing finds a hidden cap of about 100 requests per account per day (not officially published)
  • 400+ models, plus access to OpenRouter's :free pool

Strengths

  • No backend and no keys: you add a few lines of frontend code and never see a bill
  • The widest model coverage in this guide
  • User-pays fits products that give AI features to users for free
  • Open source and self-hostable

Drawbacks

  • The daily cap isn't published. Community tests put it around 100 requests per account per day, after which calls fail with a usage-limited error
  • Browser only; no server-side or scheduled jobs
  • Not a complete API surface: no Responses API, no structured outputs
  • Every end user has to register a Puter account, so the sign-up hassle moves to them

Getting started

  1. Read the developer docs at puter.com and include Puter.js in your page.
  2. Call puter.ai.chat() with a model name.
  3. Your users sign in with their own Puter accounts, and AI usage is charged to them.
  4. For zero-cost backend-style work, route through the OpenRouter free models it proxies.

Best for

  • Frontend developers building web apps and demos. You manage no backend, keys or bill, and users pay for their own AI usage.
  • Teams that want every model behind one entry point. Coverage is the widest here, and changing models takes one line.

LM Studio

Local · no limits

A desktop app for Mac, Windows and Linux. You download open models (Llama, Qwen, DeepSeek and others) from its catalog and run them on your own computer, and it includes an OpenAI-compatible server at http://localhost:1234/v1. Because the model and the compute are yours, there are no limits and no account, and your data never leaves your machine.

Free quota

  • Unlimited requests and tokens, offline
  • Real capability depends on your RAM and GPU
  • The app is free and the models are free to download

Strengths

  • Free and unlimited, so you can run automated tests and agent debugging as much as you like
  • Data stays on your machine, and it works on intranets and offline
  • Supports the OpenAI format, Responses and Anthropic Messages, so tools like Claude Code can connect to local models directly

Drawbacks

  • You need the hardware: 16GB of RAM is a sensible minimum, and large models need a GPU or lots of VRAM
  • Local inference speed depends on your hardware and is usually slower than the cloud
  • You're limited to the open models your machine can run (quantized versions help shrink them)

Getting started

  1. Download and install from lmstudio.ai.
  2. Search for and download a model in the app (GGUF format; quantized versions such as Q4 are available).
  3. Turn on the local server on the Developer page. The default address is http://localhost:1234/v1.
  4. In your code, set base_url to the local address, put anything as the key and use the name of the loaded model.

Best for

  • Developers with sensitive data or intranet environments. When data must stay on the machine, local deployment is the only option.
  • Anyone running automated tests or agent debugging at volume, since there's no quota to burn.
  • Work without reliable connectivity. It runs on a plane, or in a server room with the network down.

To point Claude Code at a local model, set ANTHROPIC_BASE_URL to http://localhost:1234.

Zhipu BigModel

Official platform (China) · rate-limited, permanently free

The developer platform of Chinese model maker Zhipu, serving the GLM model family. GLM-4-Flash is permanently free with a 128K context, and GLM-4.7-Flash is also free with a 200K context. Registration takes a phone number. For readers in mainland China it's the easiest place to start, and it needs no special network setup.

Free quota

  • GLM-4-Flash: free permanently
  • GLM-4.7-Flash: free permanently
  • Concurrency around 30 (sources differ)
  • New users get bonus tokens (amount and validity per Zhipu's official terms)

Strengths

  • Free for good, with no time limit and no dollar cap, as long as you stay within the rate limit
  • No card, and no special network setup in China
  • OpenAI-compatible, and the 128K and 200K contexts suit document work
  • Direct connection within China gives the best latency and stability for users there

Drawbacks

  • The free models are Flash-class, below the flagship open models in the international free pools
  • Claims of "tens of millions of free tokens at signup" vary between sources (5 million, 20 million, monthly allowances), so don't plan a long-term project around the bonus
  • Concurrency limits are also reported inconsistently; check the console before running volume

Getting started

  1. Register with a phone number at open.bigmodel.cn.
  2. Create an API key in the console. The bonus credit arrives automatically.
  3. Use the OpenAI-compatible endpoint https://open.bigmodel.cn/api/paas/v4/ with the model glm-4-flash.

Best for

  • Learners and students in China making their first API call. It costs nothing, needs no network workaround, and takes about ten minutes from signup to a working request.
  • Agent workflows, coding assistance and other high-volume use. The model is permanently free with long context, so repeated agent calls cost nothing.
  • Developers in China who want a stable direct connection without fiddling with network setup.

Flash models have limits. One way to work around them is to build and test the workflow on GLM-4-Flash, then use Alibaba Model Studio's new-user credit to try flagships such as Qwen and DeepSeek. Stronger Zhipu models are paid. An A2Agent key serves the same models with the same request format at a lower price, so you can keep going for less after you upgrade.

Alibaba Model Studio (Bailian)

Official platform (China) · trial credit

Alibaba Cloud's model platform brings together 70+ Chinese and third-party models, including the full Qwen family, DeepSeek, Kimi, MiniMax and GLM. New users get 1M tokens per model, over 70M in total. That makes it a trial-credit offer: a 90-day balance meant for trying Chinese models side by side. When the 90 days are up, usage moves to pay-as-you-go billing. If you already know how to call these models, an A2Agent key gives you the same models behind the same OpenAI-compatible interface at a lower price, and you can switch without changing your code.

Free quota

  • 1M tokens per model
  • Over 70M tokens in total
  • Valid for 90 days, with no extension after expiry
  • Beijing (China North 2) region only, and only for real-time inference

Strengths

  • 70+ models, including third-party flagships like DeepSeek and Kimi, under one signup
  • 70M tokens is enough for a serious comparison
  • Direct connection in China, and companies already on Alibaba Cloud can move to paid easily

Drawbacks

  • The 90-day expiry is firm, so it suits a burst of testing or a short project
  • Requires an Alibaba Cloud account with real-name verification, a higher bar than Zhipu
  • Per-model quotas aren't shared, and the same verified identity can't claim them twice
  • Once the credit runs out, billing switches to pay-as-you-go by default, so turn on the "stop when free quota is used up" setting

Getting started

  1. Create an Alibaba Cloud account and complete real-name verification.
  2. Activate Model Studio at bailian.console.aliyun.com. The free credit is issued automatically.
  3. Check each model's remaining quota and expiry in the model marketplace, then create an API key.
  4. Use the OpenAI-compatible endpoint https://dashscope.aliyuncs.com/compatible-mode/v1. Example models: qwen-max, deepseek-v3.
  5. Turn on "stop when free quota is used up" in settings to avoid surprise charges.

Best for

  • Anyone who wants to try every major Chinese model at once. With 1M tokens each for Qwen, DeepSeek, Kimi and GLM, 90 days is plenty of time to pick one.
  • Teams with a project or competition inside a 90-day window. Trial credit works best when you spend it in one push.
  • Developers who need stable in-China access to third-party models like DeepSeek and Kimi from one platform.

Use the credit for a one-time comparison. For everyday use, Zhipu's permanently free tier or Groq's rate-limited tier fit better.

If you're in mainland China

If a platform won't connect directly, it's of little use to you. These notes come from community reports in September 2026, and your network may differ, so test before you commit:

  • Usually reachable directly: Groq (the most China-friendly of the international platforms, and the cheapest place to move existing OpenAI code) and OpenRouter, free pool included.
  • Test it yourself: SambaNova, Cerebras and Puter. For a product aimed at users in China, confirm that signing up for a Puter account goes smoothly for them before you adopt the user-pays model.
  • Not reachable without a special network setup: Google Gemini, the hardest of the nine to reach. Without a stable connection, start with Zhipu's GLM-4-Flash instead. With one, Gemini has the strongest free tier here.
  • Direct within China: Zhipu BigModel (phone-number signup, no network cost) and Alibaba Model Studio (Alibaba Cloud account plus real-name verification).
  • Offline: LM Studio needs no network to run. Model downloads can go through a Hugging Face mirror to speed them up.

Who should pick what

Find the row that describes you.

You are First choice Backup Why
A student or learner Zhipu GLM-4-Flash Alibaba Model Studio new-user credit No barrier, free for good, no network cost; with a reliable international connection, Gemini is more capable
An indie developer building an MVP Groq SambaNova Fast and OpenAI-compatible, existing code runs unchanged; 1,000 requests a day covers development
Out of free quota and looking for a cheap way to continue A2Agent OpenRouter ($10 unlock) $5 at signup, another $10 on your first $10 top-up, model prices up to 50% off; same models and OpenAI-compatible format, so existing code needs no changes
A team selecting or evaluating models OpenRouter ($10 unlock) Puter One key to compare 20+ free models; $10 buys 1,000 requests a day for good
Processing long documents or batch text Cerebras SambaNova About 1M tokens a day; SambaNova allows about 10K tokens per request
Building a web app where users share the AI cost Puter OpenRouter Frontend integration, no developer bill, users pay from their own accounts
Wanting to run a 405B-class open model for free SambaNova Cerebras A rare flagship slot on a free tier, if 20 requests a day works for you
Working with sensitive data, offline or on an intranet LM Studio None Data stays local, no limits, works offline; the cost is your hardware
Wanting to try every Chinese model at once Alibaba Model Studio Zhipu 1M tokens for each of 70+ models; compare within 90 days, then decide who to pay

Six ways to stretch a free quota

Free limits are small, but with these habits even 50 requests a day can carry a learning project.

  1. Cache repeated answers. Asking the same question ten times burns ten requests, so store common Q&A locally and check the cache before calling the API. Of all six tips, this one saves the most quota.
  2. Trim your prompts. Limits count tokens as well as requests (200K a day on Groq, about 1M on Cerebras), and long context plus long prompts drains them fast. Send only what the model needs.
  3. Batch small tasks. Per-minute request caps are low (30 on Groq, 20 on SambaNova), so combine several small tasks into one call.
  4. Give easy tasks to Flash or other small models and save the big ones for hard problems. Small models in the free pool often have the same limits but use fewer tokens, so your quota goes further.
  5. Keep a backup provider. All nine options use the OpenAI format, so switching is a one-line base_url change. When one provider throttles you, move to another. Many people on free tiers do this, and it keeps you from depending on a single vendor.
  6. Watch the dashboard. Failed requests count against your quota (OpenRouter says so explicitly). Check each provider's usage panel before you ship, so a retry loop doesn't burn a day's allowance.

Free is where you start

Of the nine, there are two ways to use AI for free over the long run. One is a rate-limited free tier (Gemini, Groq, SambaNova, Cerebras, OpenRouter, Puter, Zhipu), with enough quota for learning and small tools. The other is local deployment with LM Studio, which has no limits and costs you hardware instead. If you write against the OpenAI-compatible format from day one, moving to paid when traffic grows means changing one line of base_url, so there's no cost to starting free and paying later.

Free tier tracker

We update this table whenever a provider changes its free quota or free model list. Last updated: October 8, 2026. Check back here before you plan a project around a number.

Platform Type of free What you get free To sign up Link
A2Agent Signup credit $5 credit at signup; first $10 top-up earns another $10 Sign up at a2agent.me a2agent.me/pricing
Google Gemini Rate-limited Flash models at roughly 10 to 15 requests/min; live limits in the console Google account, no card ai.google.dev
Groq Rate-limited 30 requests/min, 1,000 requests/day, 200K tokens/day Email, no card console.groq.com
SambaNova Rate-limited Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day per model; plus a one-time $5 credit No card cloud.sambanova.ai
Cerebras Rate-limited About 1M tokens/day No card cloud.cerebras.ai
OpenRouter Rate-limited + rotating pool 20+ :free models at 20 requests/min, 50 requests/day (1,000/day after a one-time $10 top-up) Google or GitHub account openrouter.ai
Puter User-pays Free for developers; end users get starter credit (about 100 requests/day per account in community tests) Each end user needs a Puter account puter.com
LM Studio Local Unlimited; limited only by your hardware No account lmstudio.ai
Zhipu BigModel Rate-limited, permanently free GLM-4-Flash and GLM-4.7-Flash free for good, plus new-user bonus tokens Phone number open.bigmodel.cn
Alibaba Model Studio Trial credit 1M tokens per model, over 70M in total, valid 90 days Alibaba Cloud account with real-name verification bailian.console.aliyun.com

Data note: quotas in this post were checked on September 30, 2026, mainly against each provider's official documentation and pricing pages; the sources are listed below. Free model pools and limits change at any time, and Google and Cerebras no longer publish static limit tables, so the live numbers in each console take precedence. This post is not a service commitment on behalf of any platform.

Sources