Free AI APIs in 2026: Quotas, Catches and Who Each One Fits
A2Agent Team · 2026-10-08T00:00:00Z
Almost every model provider now advertises a free tier. Sign up for a few and the differences show up fast. Some want a credit card. Some quotas are gone by lunchtime. Some terms change from month to month, and some free tiers train on whatever you send them.
This guide covers six international platforms, one local tool and two platforms based in China. For each one you'll find the quota, the trade-offs, how to make a first call and who it suits. We checked the quotas on September 30, 2026. Providers change them often, so trust each provider's console over any number here.
Five kinds of "free"
In this guide, "free" means you can call a model's API without paying. That covers five different arrangements:
- Rate-limited. Free with no end date, capped by requests per minute, requests per day or tokens per day. The quota never runs out as long as you stay under the cap. Gemini, Groq, SambaNova, Cerebras, the OpenRouter free pool and Zhipu work this way.
- Trial credit. A one-time balance billed per token, gone once it's spent. It suits a short burst of testing. Example: Alibaba Model Studio's 90-day credit.
- Rotating pool. The quota is permanent, but the set of free models keeps changing. You can stay free for a long time as long as your project isn't tied to one specific model. Example: OpenRouter's
:freemodels. - User-pays. Free for the developer, with no developer-side bill. AI usage is charged to each end user's own account, and new accounts usually come with starter credit. Example: Puter.
- Local. The model runs on your own machine, so there are no request or token limits and it works offline. You pay for it in hardware. Example: LM Studio.
What you pay instead of money
The terms for Gemini's free tier and OpenRouter's free models allow your inputs and outputs to be used for training, so keep sensitive data out of them. Some platforms want personal details instead: Zhipu asks for a phone number, and Puter needs each of your users to sign in to a Puter account. LM Studio costs nothing in API fees, but you supply the GPU and memory.
Free quotas are sized for learning and development. They won't carry production traffic, and once a project grows you will pay.
The map at a glance
| Platform | Category | Free quota (highlights) | Best for | Main catch |
|---|---|---|---|---|
| A2Agent | Aggregator | $5 credit at signup; first $10 top-up earns another $10; model prices up to 50% off | Everyday calls, and cheap continued use after free quotas run out | Apart from new-user offers, promotions for existing users change from time to time |
| Google Gemini | Official developer platform | No card; static limits no longer published, Flash models roughly 10 to 15 requests/min | A first AI project, multimodal work | Free-tier data used for training (except in the EEA, Switzerland and the UK) |
| Groq | Fast inference | 30 requests/min, 1,000 requests/day, 200K tokens/day, no card | Low-latency apps, existing OpenAI code | Few models, fast turnover; old tutorials break |
| SambaNova | Fast inference | Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day (per model), no card | Running 405B-class open models for free | Only 20 requests a day |
| Cerebras | Fast inference | About 1M tokens/day, no card; published request limits disagree | Long documents, batch text jobs | Free model lineup changes often |
| OpenRouter | Aggregator | Free models: 20 requests/min, 50 requests/day; a one-time $10 top-up raises that to 1,000/day for good | Comparing models side by side | Rotating pool; requires a training/logging opt-in |
| Puter | Aggregator | Free for developers; AI usage billed to end users' accounts, new accounts get starter credit | Web apps, frontend developers | Undisclosed daily cap; browser only |
| LM Studio | Local | No limits; capability depends on your hardware | Sensitive data, offline or intranet work | You need the GPU and RAM |
| Zhipu BigModel | Official platform (China) | GLM-4-Flash free permanently, plus new-user credit | Learners in China, high-volume agent workloads | Bonus-credit terms per Zhipu's official rules |
| Alibaba Model Studio | Official platform (China) | 1M tokens per model, over 70M in total, valid 90 days | Trying Chinese models side by side | Expires after 90 days; Beijing region only |
Quotas come from each provider's documentation and community verification pages from September 2026. Google and Cerebras in particular adjust limits and free models often, so check the console before you plan around a number.
How the platforms differ
Official developer platforms: Gemini, Zhipu, Alibaba Model Studio
These are the model makers' own developer sites, where you register and call their models directly. The free tiers are aimed at individual developers who are trying things out. Quotas are fairly generous (Gemini's multimodal access, Zhipu's permanently free model), and the terms usually say your inputs and outputs can be used to train their models.
Fast inference platforms: Groq, SambaNova, Cerebras
These companies serve open-weight models on their own inference chips at several hundred to several thousand tokens per second, which helps when you're developing something where latency matters. They meter differently. Groq counts requests, SambaNova allows few requests but plenty of tokens per request, and Cerebras counts total tokens.
Aggregators: OpenRouter and Puter
Third-party platforms that give you one entry point to models from many vendors. OpenRouter puts 500+ models behind one key, 20+ of them free. Puter targets browser apps: developers integrate at no cost, and the AI bill lands on the end user's account.
Local: LM Studio
The model and the compute both live on your computer, and nothing goes through a server. You don't need an account, there's no quota, and your data stays on your machine. How strong a model you can run depends on your GPU and memory.
One API format across all nine
All nine free options accept the OpenAI API format. Once you have working openai library code, you can move between them by changing the base URL, the key and the model name, which is why the "keep a backup provider" tip further down works. LM Studio additionally speaks the Anthropic format (/v1/messages), so tools like Claude Code can talk to a local model directly. If you're new to switching providers this way, our OpenAI SDK integration guide shows the two-line change.
The platforms, one by one
Google Gemini (AI Studio)
Official developer platform · rate-limited
Gemini's official API, offered through Google AI Studio. The free tier needs no credit card and handles text, images and code, which makes it the most well-rounded free option for everyday use.
Free quota
- No credit card; sign up and start
- Flash models at roughly 10 to 15 requests per minute
- Some accounts see per-model tokens-per-minute limits in the millions
- Since September 2026 Google no longer publishes a static limits table; the console shows your live limits
Strengths
- Handles text, images and code, with the strongest multimodal capability in the free pool
- Lowest barrier: no card and no phone verification
- Fairly generous limits, and Flash models handle volume well
- Mature SDKs and a smooth path to paid
Drawbacks
- Limits aren't transparent, and the figures quoted in the press contradict each other
- Free-tier data may be used for training (except in the EEA, Switzerland and the UK), so don't send sensitive data
- Pro models have mostly left the free tier, and enabling billing on a project replaces that project's free tier
- Hardest to reach from mainland China (see the China section below)
Getting started
- Go to ai.google.dev and sign in to AI Studio with a Google account.
- Click "Get API key". You don't need a card, and it takes about a minute.
- Use the OpenAI-compatible endpoint
https://generativelanguage.googleapis.com/v1beta/openai/or Google's own SDK. Example model:gemini-2.5-flash. - Check the rate-limit panel in AI Studio for your account's current limits before you start.
Best for
- Students and career changers building a first AI project. No card or verification, and one platform covers text, images and code while you learn.
- Indie developers who need mixed image and text processing. Nothing else in the free pool matches Gemini's vision capability.
- Developers with reliable international connectivity building content or automation tools. The quota is large enough to keep a small tool running long term, which most rate-limited tiers can't do.
Groq
Fast inference · rate-limited
Groq serves open-weight models (OpenAI's gpt-oss series, Qwen variants and others) on its own LPU chips. The free tier runs at several hundred tokens per second, and for chat or voice apps it feels close to a paid product.
Free quota (per organization, no card)
- 30 requests/min and 1,000 requests/day
- 8,000 tokens/min and 200,000 tokens/day
- Models:
gpt-oss-120b,gpt-oss-20b,qwen3.6-27b
Strengths
- The fastest of the free tiers here
- Fully OpenAI-compatible: change
base_urland the key in your existing code and it runs - 200K tokens a day is generous for a rate-limited tier, enough for a project in development
- Email signup, no card
Drawbacks
- Few models, and they rotate quickly. The Llama models were retired in August 2026, so code from older tutorials fails
- Open-weight models only, nothing at the level of Gemini's closed flagships
- The 200K daily token cap goes faster than you'd expect with long prompts and large contexts
- Requests may queue at peak times; less stable than paid
Getting started
- Sign up with email at console.groq.com. No card needed.
- Create a key on the API Keys page, and check the Limits page for your live limits.
- Point your OpenAI code at
https://api.groq.com/openai/v1and use the modelopenai/gpt-oss-120b. - The widely quoted "14,400 requests a day" is out of date. Plan for 1,000.
Best for
- Indie developers building an MVP. 1,000 requests a day covers development, and debugging feels the same as on a paid tier.
- Real-time chat and voice apps. Groq's free tier responds faster than any other here.
- Anyone with OpenAI-format code who wants to run it for free, since only two config lines change.
SambaNova
Fast inference · rate-limited
SambaNova Cloud runs inference on its own RDU chips at speeds close to Groq's. Its free list includes Llama 3.1 405B, a flagship open model, while most free tiers stop at small and mid-size models.
Free quota (no card, counted per model)
- Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day
- Llama 3.1 405B: about 10 requests/min
- DeepSeek-V3.1 and gpt-oss-120b are also free
- New users get an extra one-time $5 credit
Strengths
- Free access to 405B-class open models, which few free tiers offer
- Fast inference on RDU chips
- 200K tokens a day per model works out to about 10K tokens per request, which suits long tasks
- No card, OpenAI-compatible, and rate-limit status comes back in the response headers
Drawbacks
- Only 20 requests a day, which suits a few big jobs and rules out frequent chat
- The lineup is mostly Llama, so the choice is narrow
- Official documentation is vague about some limits, and sources disagree on the token cap
Getting started
- Register at cloud.sambanova.ai. No card needed.
- Create an API key in the console.
- Use the OpenAI-compatible endpoint
https://api.sambanova.ai/v1. - Try the model
Meta-Llama-3.3-70B-Instruct, and read the rate-limit fields in the response headers to track what's left.
Best for
- Anyone who wants to run a 405B-class open model for free. Hardly any other free tier has one.
- Occasional heavy jobs such as long-document analysis, where about 10K tokens per request across 20 requests a day is a good fit.
- A backup for Groq. Both are fast, but the models and quota structure differ, so each covers the other's gaps.
Cerebras
Fast inference · rate-limited
Cerebras runs inference on wafer-scale chips at several thousand tokens per second, so long outputs arrive almost instantly. Its free quota is measured in total tokens, about 1M a day, the largest token allowance of the nine.
Free quota
- About 1,000,000 tokens/day
- No card, free with no end date
- Request limits are reported inconsistently: older sources say 30/min, 2026 sources say 5/min. Check the console
- Current free models:
gpt-oss-120band GLM-4.7 (Llama has been phased out)
Strengths
- Very fast on long outputs
- The highest daily token allowance here, enough for real batch processing
- 131K context for long documents
- No card, OpenAI-compatible
Drawbacks
- Published request limits contradict each other, so it's hard to plan around them
- The free lineup changes quickly; Llama is gone, replaced by gpt-oss-120b and GLM-4.7
- Unused tokens don't roll over, and the count resets daily
Getting started
- Register at cloud.cerebras.ai. No card needed.
- Create an API key in the console.
- Use the OpenAI-compatible endpoint
https://api.cerebras.ai/v1. - Check the current free model list before choosing a model, because it changes every few months. Examples:
gpt-oss-120b,glm-4.7.
Best for
- Long-document and batch text jobs. About 1M tokens a day is more than any other free tier here gives you.
- Apps where users wait on long replies. At thousands of tokens per second, they barely wait at all.
OpenRouter
Aggregator · rate-limited + rotating pool
OpenRouter is a single gateway to 500+ models, so one account and one key reach almost every major model. More than 20 free models (IDs ending in :free) rotate through the catalog, including Llama 3.3 70B, Qwen3 Coder, GPT-OSS and Nemotron. It's most useful when you haven't decided which model to use yet.
Free quota
- 20 requests/min and 50 requests/day
- A one-time $10 top-up raises the daily cap to 1,000 requests permanently, even after the balance is spent
- Free models cost $0 per token; the limit is the request count
- The free router
openrouter/freepicks an available model for you
Strengths
- One key for 500+ models, so comparing them costs almost nothing
- A one-time $10 permanently raises the daily limit 20× (50 to 1,000 requests), the best-value one-off spend in this guide
- The free router fails over automatically, so if one model is down it moves to another
- Switching models means changing one parameter, so you aren't locked in
Drawbacks
- Free models require you to enable training/logging, which puts the data terms in the same category as Gemini's
- Failed requests count against the daily quota, so a retry loop can burn through 50 requests quickly
- Free endpoints may have smaller context windows than paid ones, and upstream providers throttle at peak times
- OpenRouter itself says free models aren't meant for production
Getting started
- Sign in at openrouter.ai with Google or GitHub.
- Create an API key on the Keys page.
- Use the endpoint
https://openrouter.ai/api/v1and add:freeto a model ID to use the free quota, for exampleqwen/qwen3-coder:free. - Without a top-up you get 50 requests a day. After one $10 top-up you get 1,000 a day for good.
Best for
- Anyone who wants to compare models before paying. Send one prompt to 20+ free models and you'll know which fits your task before you buy a subscription.
- Teams running model evaluations. With one interface and one meter, an evaluation script works across every model.
- Developers who don't want to depend on one vendor. Your code talks to OpenRouter, and the vendors behind it can change.
Puter
Aggregator · user-pays
Puter (puter.com) is an open-source "internet operating system". Its Puter.js library lets a web app add AI in a few lines of code, with access to 400+ models from 70+ providers, including OpenAI, Anthropic, Gemini, Qwen and Nemotron as well as image, video and speech models. Developers integrate at no cost, and AI usage is charged to each end user's own Puter account.
Free quota
- Developer side: free to integrate, no bill
- User side: new accounts get starter credit
- Community testing finds a hidden cap of about 100 requests per account per day (not officially published)
- 400+ models, plus access to OpenRouter's
:freepool
Strengths
- No backend and no keys: you add a few lines of frontend code and never see a bill
- The widest model coverage in this guide
- User-pays fits products that give AI features to users for free
- Open source and self-hostable
Drawbacks
- The daily cap isn't published. Community tests put it around 100 requests per account per day, after which calls fail with a
usage-limitederror - Browser only; no server-side or scheduled jobs
- Not a complete API surface: no Responses API, no structured outputs
- Every end user has to register a Puter account, so the sign-up hassle moves to them
Getting started
- Read the developer docs at puter.com and include Puter.js in your page.
- Call
puter.ai.chat()with a model name. - Your users sign in with their own Puter accounts, and AI usage is charged to them.
- For zero-cost backend-style work, route through the OpenRouter free models it proxies.
Best for
- Frontend developers building web apps and demos. You manage no backend, keys or bill, and users pay for their own AI usage.
- Teams that want every model behind one entry point. Coverage is the widest here, and changing models takes one line.
LM Studio
Local · no limits
A desktop app for Mac, Windows and Linux. You download open models (Llama, Qwen, DeepSeek and others) from its catalog and run them on your own computer, and it includes an OpenAI-compatible server at http://localhost:1234/v1. Because the model and the compute are yours, there are no limits and no account, and your data never leaves your machine.
Free quota
- Unlimited requests and tokens, offline
- Real capability depends on your RAM and GPU
- The app is free and the models are free to download
Strengths
- Free and unlimited, so you can run automated tests and agent debugging as much as you like
- Data stays on your machine, and it works on intranets and offline
- Supports the OpenAI format, Responses and Anthropic Messages, so tools like Claude Code can connect to local models directly
Drawbacks
- You need the hardware: 16GB of RAM is a sensible minimum, and large models need a GPU or lots of VRAM
- Local inference speed depends on your hardware and is usually slower than the cloud
- You're limited to the open models your machine can run (quantized versions help shrink them)
Getting started
- Download and install from lmstudio.ai.
- Search for and download a model in the app (GGUF format; quantized versions such as Q4 are available).
- Turn on the local server on the Developer page. The default address is
http://localhost:1234/v1. - In your code, set
base_urlto the local address, put anything as the key and use the name of the loaded model.
Best for
- Developers with sensitive data or intranet environments. When data must stay on the machine, local deployment is the only option.
- Anyone running automated tests or agent debugging at volume, since there's no quota to burn.
- Work without reliable connectivity. It runs on a plane, or in a server room with the network down.
To point Claude Code at a local model, set ANTHROPIC_BASE_URL to http://localhost:1234.
Zhipu BigModel
Official platform (China) · rate-limited, permanently free
The developer platform of Chinese model maker Zhipu, serving the GLM model family. GLM-4-Flash is permanently free with a 128K context, and GLM-4.7-Flash is also free with a 200K context. Registration takes a phone number. For readers in mainland China it's the easiest place to start, and it needs no special network setup.
Free quota
- GLM-4-Flash: free permanently
- GLM-4.7-Flash: free permanently
- Concurrency around 30 (sources differ)
- New users get bonus tokens (amount and validity per Zhipu's official terms)
Strengths
- Free for good, with no time limit and no dollar cap, as long as you stay within the rate limit
- No card, and no special network setup in China
- OpenAI-compatible, and the 128K and 200K contexts suit document work
- Direct connection within China gives the best latency and stability for users there
Drawbacks
- The free models are Flash-class, below the flagship open models in the international free pools
- Claims of "tens of millions of free tokens at signup" vary between sources (5 million, 20 million, monthly allowances), so don't plan a long-term project around the bonus
- Concurrency limits are also reported inconsistently; check the console before running volume
Getting started
- Register with a phone number at open.bigmodel.cn.
- Create an API key in the console. The bonus credit arrives automatically.
- Use the OpenAI-compatible endpoint
https://open.bigmodel.cn/api/paas/v4/with the modelglm-4-flash.
Best for
- Learners and students in China making their first API call. It costs nothing, needs no network workaround, and takes about ten minutes from signup to a working request.
- Agent workflows, coding assistance and other high-volume use. The model is permanently free with long context, so repeated agent calls cost nothing.
- Developers in China who want a stable direct connection without fiddling with network setup.
Flash models have limits. One way to work around them is to build and test the workflow on GLM-4-Flash, then use Alibaba Model Studio's new-user credit to try flagships such as Qwen and DeepSeek. Stronger Zhipu models are paid. An A2Agent key serves the same models with the same request format at a lower price, so you can keep going for less after you upgrade.
Alibaba Model Studio (Bailian)
Official platform (China) · trial credit
Alibaba Cloud's model platform brings together 70+ Chinese and third-party models, including the full Qwen family, DeepSeek, Kimi, MiniMax and GLM. New users get 1M tokens per model, over 70M in total. That makes it a trial-credit offer: a 90-day balance meant for trying Chinese models side by side. When the 90 days are up, usage moves to pay-as-you-go billing. If you already know how to call these models, an A2Agent key gives you the same models behind the same OpenAI-compatible interface at a lower price, and you can switch without changing your code.
Free quota
- 1M tokens per model
- Over 70M tokens in total
- Valid for 90 days, with no extension after expiry
- Beijing (China North 2) region only, and only for real-time inference
Strengths
- 70+ models, including third-party flagships like DeepSeek and Kimi, under one signup
- 70M tokens is enough for a serious comparison
- Direct connection in China, and companies already on Alibaba Cloud can move to paid easily
Drawbacks
- The 90-day expiry is firm, so it suits a burst of testing or a short project
- Requires an Alibaba Cloud account with real-name verification, a higher bar than Zhipu
- Per-model quotas aren't shared, and the same verified identity can't claim them twice
- Once the credit runs out, billing switches to pay-as-you-go by default, so turn on the "stop when free quota is used up" setting
Getting started
- Create an Alibaba Cloud account and complete real-name verification.
- Activate Model Studio at bailian.console.aliyun.com. The free credit is issued automatically.
- Check each model's remaining quota and expiry in the model marketplace, then create an API key.
- Use the OpenAI-compatible endpoint
https://dashscope.aliyuncs.com/compatible-mode/v1. Example models:qwen-max,deepseek-v3. - Turn on "stop when free quota is used up" in settings to avoid surprise charges.
Best for
- Anyone who wants to try every major Chinese model at once. With 1M tokens each for Qwen, DeepSeek, Kimi and GLM, 90 days is plenty of time to pick one.
- Teams with a project or competition inside a 90-day window. Trial credit works best when you spend it in one push.
- Developers who need stable in-China access to third-party models like DeepSeek and Kimi from one platform.
Use the credit for a one-time comparison. For everyday use, Zhipu's permanently free tier or Groq's rate-limited tier fit better.
If you're in mainland China
If a platform won't connect directly, it's of little use to you. These notes come from community reports in September 2026, and your network may differ, so test before you commit:
- Usually reachable directly: Groq (the most China-friendly of the international platforms, and the cheapest place to move existing OpenAI code) and OpenRouter, free pool included.
- Test it yourself: SambaNova, Cerebras and Puter. For a product aimed at users in China, confirm that signing up for a Puter account goes smoothly for them before you adopt the user-pays model.
- Not reachable without a special network setup: Google Gemini, the hardest of the nine to reach. Without a stable connection, start with Zhipu's GLM-4-Flash instead. With one, Gemini has the strongest free tier here.
- Direct within China: Zhipu BigModel (phone-number signup, no network cost) and Alibaba Model Studio (Alibaba Cloud account plus real-name verification).
- Offline: LM Studio needs no network to run. Model downloads can go through a Hugging Face mirror to speed them up.
Who should pick what
Find the row that describes you.
| You are | First choice | Backup | Why |
|---|---|---|---|
| A student or learner | Zhipu GLM-4-Flash | Alibaba Model Studio new-user credit | No barrier, free for good, no network cost; with a reliable international connection, Gemini is more capable |
| An indie developer building an MVP | Groq | SambaNova | Fast and OpenAI-compatible, existing code runs unchanged; 1,000 requests a day covers development |
| Out of free quota and looking for a cheap way to continue | A2Agent | OpenRouter ($10 unlock) | $5 at signup, another $10 on your first $10 top-up, model prices up to 50% off; same models and OpenAI-compatible format, so existing code needs no changes |
| A team selecting or evaluating models | OpenRouter ($10 unlock) | Puter | One key to compare 20+ free models; $10 buys 1,000 requests a day for good |
| Processing long documents or batch text | Cerebras | SambaNova | About 1M tokens a day; SambaNova allows about 10K tokens per request |
| Building a web app where users share the AI cost | Puter | OpenRouter | Frontend integration, no developer bill, users pay from their own accounts |
| Wanting to run a 405B-class open model for free | SambaNova | Cerebras | A rare flagship slot on a free tier, if 20 requests a day works for you |
| Working with sensitive data, offline or on an intranet | LM Studio | None | Data stays local, no limits, works offline; the cost is your hardware |
| Wanting to try every Chinese model at once | Alibaba Model Studio | Zhipu | 1M tokens for each of 70+ models; compare within 90 days, then decide who to pay |
Six ways to stretch a free quota
Free limits are small, but with these habits even 50 requests a day can carry a learning project.
- Cache repeated answers. Asking the same question ten times burns ten requests, so store common Q&A locally and check the cache before calling the API. Of all six tips, this one saves the most quota.
- Trim your prompts. Limits count tokens as well as requests (200K a day on Groq, about 1M on Cerebras), and long context plus long prompts drains them fast. Send only what the model needs.
- Batch small tasks. Per-minute request caps are low (30 on Groq, 20 on SambaNova), so combine several small tasks into one call.
- Give easy tasks to Flash or other small models and save the big ones for hard problems. Small models in the free pool often have the same limits but use fewer tokens, so your quota goes further.
- Keep a backup provider. All nine options use the OpenAI format, so switching is a one-line
base_urlchange. When one provider throttles you, move to another. Many people on free tiers do this, and it keeps you from depending on a single vendor. - Watch the dashboard. Failed requests count against your quota (OpenRouter says so explicitly). Check each provider's usage panel before you ship, so a retry loop doesn't burn a day's allowance.
Free is where you start
Of the nine, there are two ways to use AI for free over the long run. One is a rate-limited free tier (Gemini, Groq, SambaNova, Cerebras, OpenRouter, Puter, Zhipu), with enough quota for learning and small tools. The other is local deployment with LM Studio, which has no limits and costs you hardware instead. If you write against the OpenAI-compatible format from day one, moving to paid when traffic grows means changing one line of base_url, so there's no cost to starting free and paying later.
Free tier tracker
We update this table whenever a provider changes its free quota or free model list. Last updated: October 8, 2026. Check back here before you plan a project around a number.
| Platform | Type of free | What you get free | To sign up | Link |
|---|---|---|---|---|
| A2Agent | Signup credit | $5 credit at signup; first $10 top-up earns another $10 | Sign up at a2agent.me | a2agent.me/pricing |
| Google Gemini | Rate-limited | Flash models at roughly 10 to 15 requests/min; live limits in the console | Google account, no card | ai.google.dev |
| Groq | Rate-limited | 30 requests/min, 1,000 requests/day, 200K tokens/day | Email, no card | console.groq.com |
| SambaNova | Rate-limited | Llama 3.3 70B: 20 requests/min, 20 requests/day, 200K tokens/day per model; plus a one-time $5 credit | No card | cloud.sambanova.ai |
| Cerebras | Rate-limited | About 1M tokens/day | No card | cloud.cerebras.ai |
| OpenRouter | Rate-limited + rotating pool | 20+ :free models at 20 requests/min, 50 requests/day (1,000/day after a one-time $10 top-up) |
Google or GitHub account | openrouter.ai |
| Puter | User-pays | Free for developers; end users get starter credit (about 100 requests/day per account in community tests) | Each end user needs a Puter account | puter.com |
| LM Studio | Local | Unlimited; limited only by your hardware | No account | lmstudio.ai |
| Zhipu BigModel | Rate-limited, permanently free | GLM-4-Flash and GLM-4.7-Flash free for good, plus new-user bonus tokens | Phone number | open.bigmodel.cn |
| Alibaba Model Studio | Trial credit | 1M tokens per model, over 70M in total, valid 90 days | Alibaba Cloud account with real-name verification | bailian.console.aliyun.com |
Data note: quotas in this post were checked on September 30, 2026, mainly against each provider's official documentation and pricing pages; the sources are listed below. Free model pools and limits change at any time, and Google and Cerebras no longer publish static limit tables, so the live numbers in each console take precedence. This post is not a service commitment on behalf of any platform.
Sources
- A2Agent pricing
- Groq Free Tier 2026: 1,000 Requests a Day, Llama Is Gone (Klymentiev Blog)
- Groq's 14,400 requests a day is not for the chat models (DEV Community)
- Rate Limits Policy (SambaNova documentation)
- Cerebras Pricing 2026: Actual Rate Limits, $/MTok & Enterprise Tiers (Morph)
- Free, Unlimited AI API (Puter developer docs)
- Use your LM Studio Models in Claude Code (LM Studio blog)
- Spending $10 once on OpenRouter raises your free-tier cap 20x, permanently (DEV Community)
- Google raises free Gemini API quotas, some models reach 1M tokens per minute (Appinn, in Chinese)
- Free LLM API Tiers: Limits by Provider, Verified (BenchLM)
- Alibaba Cloud Model Studio guide: entry points, free token quota and FAQ (Alibaba Cloud Developer Community, in Chinese)