English · 简体中文
CI License: MIT providers sources checked
Permanent free tiers, no-card options, direct API key links, models, and verified limits — every claim points to an official source.
13 permanent provider free tiers · 13 require no credit card · 13 are OpenAI compatible · sources reviewed 2026年08月22日. No keys are distributed here. A probe describes one sampled request, not provider-wide uptime.
Browse the live directory · Pick by model · Set up a coding agent · Check your own key
Filterable LLM free-tier status page
Star this repository to bookmark the dataset and follow releases. A star changes nothing about any provider's keys, credits, or limits, and this project gives nothing in return for one.
| Goal | Pick | Why it appears here | Start |
|---|---|---|---|
| Highest published daily request limit | GroqCloud | 1,000 requests/day published | Get API key |
| Highest published requests per minute | SiliconFlow | 1,000 RPM published | Get API key |
| Works with the browser key checker | Google Gemini API | Browser CORS check supported | Get API key |
| Fast path for coding agents | GroqCloud | Documented OpenAI-compatible coding setup | Get API key |
These are rule-based shortcuts, not paid placements. Open the filterable directory for all 26 providers.
Free is not the same as cheap enough to keep running, and the number you need is what replaces the free tier once it runs out. Same maintainer, re-reading OpenRouter's whole catalog every day — 60 models on 2026年08月22日, cheapest five first:
| Model | Input / M tokens | Output / M tokens | Design Arena agents |
|---|---|---|---|
| Gemini 3.7 Flash (batch) | 0ドル.1875 | 0ドル.9375 | #3 androidnative |
| Gemini 3 Flash Preview (batch) | 0ドル.25 | 1ドル.50 | #9 agenticslides |
| MiniMax M3 | 0ドル.30 | 1ドル.20 | #10 python-pptxslides |
| Gemini 3.6 Flash (batch) | 0ドル.375 | 1ドル.875 | #6 agenticgamedev |
| GLM 4.7 | 0ドル.40 | 1ドル.75 | #27 androidnative |
A (batch) row is the queued price, not the interactive one — an agent waiting on the reply pays the other number. All 60 rows as JSON or CSV, re-read tomorrow.
A row price is not a request price. 12 of those 60 rows carry a second, higher rate card that switches on when the prompt gets long — and it applies to every token in the request, including the ones before the threshold, so a prompt one token over the line costs about double one token under it. The thresholds are 200k and 272k prompt tokens. The counter-intuitive part: the bigger the advertised context window, the smaller the share of it the advertised price covers — x-ai/grok-4.20 advertises 2 000k of context at 1ドル.25 per million input and steps to 2ドル.50 at 200k, so 90% of that window bills at a number that is not in its row. None of the five above is one of them, which is most of why they are the cheap ones. Both numbers ride on every row of the JSON and CSV as long_context_from and long_input_per_million.
A row price is not an hour price either. DeepSeek's V4 line runs two rate cards off the clock: peak is 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, and every other hour is off-peak at exactly half. Since 2026年08月23日 the entire weekend is off-peak too — 14 hours a week that most published summaries, and most of the billing code I have read, still charge at 2x. Per million: deepseek-v4-flash input 0ドル.22 off-peak against 0ドル.44 peak, deepseek-v4-pro output 1ドル.98 against 3ドル.96. A nightly batch moved a few hours earlier pays half, and no code change buys that. Which hours, and the two instants where a UTC weekday and a Beijing weekday disagree.
A list price is what a vendor publishes; what a month of it came to is a different number, and only somebody who paid it can tell you. The same repository keeps those too — every figure quoted in a write-up, with the full sentence it came from on every row:
| Price | Unit | The sentence it was published in |
|---|---|---|
0ドル.06 / 0ドル.2 |
per million | At the low end, MiniMax M3 runs 0ドル.06 to 0ドル.2 per million and draws 60% to 70% of its revenue from outside its home market. → |
0ドル.19 / 5ドル |
per million tokens | Chinese AI models provide a cost-effective alternative to their American counterparts, with input costs as low as 0ドル.19 per million tokens, compared to OpenAI's 5ドル-12. → |
1ドル |
per million tokens | Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near 1ドル per million tokens at inference gross margins of 10% to 20%. → |
1ドル.25 / 4ドル.25 |
per million | Meta priced Muse Spark 1.1 at 1ドル.25 per million input and 4ドル.25 per million output, roughly 75% and 83% below Anthropic's Opus, and the tradeoff is visible in the benchmarks, since it leads on MCP Atlas and JobBench while trailing on SWE-Bench Pro and DeepSWE 1.1. → |
3ドル |
per million input tokens | The 3ドル per million input tokens price point means developers should carefully evaluate whether the premium model's capabilities justify the increased costs for their specific use cases. → |
A 1ドル.43 is never left ambiguous between per million tokens, per month and per seat, because the sentence travels with it. Readable in code as JSON or CSV, or as prose: where the token bill actually goes.
Paying for a model whose bill surprised you? Name it in one line — one field, and it decides which price gets chased next.
These Provider Free Tiers are the main list: they do not expire like trial credits, and none currently require a credit card.
| Provider | Models | Published limits | Card | OpenAI compatible | Get API key |
|---|---|---|---|---|---|
| Google Gemini API | Free-tier eligibility varies by model gemini-2.5-flash gemini-2.5-flash-lite |
Dynamic / model-dependent | Not required | Yes | Open |
| GroqCloud | openai/gpt-oss-120b openai/gpt-oss-20b openai/gpt-oss-safeguard-20b |
30 RPM, 1,000 requests/day | Not required | Yes | Open |
| SambaNova Cloud | DeepSeek-V3.1 Meta-Llama-3.3-70B-Instruct gpt-oss-120b |
20 RPM, 20 requests/day | Not required | Yes | Open |
| Cohere | command-a-03-2025 command-r-plus embed-v4.0 |
20 RPM | Not required | Yes | Open |
| Cloudflare Workers AI | @cf/meta/llama-3.3-70b-instruct-fp8-fast @cf/openai/gpt-oss-120b @cf/qwen/qwen2.5-coder-32b-instruct |
Dynamic / model-dependent | Not required | Yes | Open |
| Hugging Face Inference Providers | deepseek-ai/DeepSeek-V3-0324 openai/gpt-oss-120b 200+ models routed across partner providers |
Dynamic / model-dependent | Not required | Yes | Open |
| SiliconFlow | Qwen/Qwen3-8B THUDM/GLM-4-9B-0414 deepseek-ai/DeepSeek-R1 |
1,000 RPM | Not required | Yes | Open |
| Fireworks AI | accounts/fireworks/models/llama-v3p3-70b-instruct accounts/fireworks/models/gpt-oss-120b |
10 RPM | Not required | Yes | Open |
| Z.AI Open Platform | GLM-4.7-Flash GLM-4.5-Flash GLM-4.6V-Flash |
Dynamic / model-dependent | Not required | Yes | Open |
| Mistral La Plateforme | mistral-small-latest open-mistral-nemo codestral-latest |
Dynamic / model-dependent | Not required | Yes | Open |
| Alibaba Cloud Model Studio | qwen-plus qwen-turbo qwen3-coder-plus |
Dynamic / model-dependent | Not required | Yes | Open |
| Pollinations.AI | openai mistral Community-hosted open models |
4 RPM | Not required | Yes | Open |
| Ollama Cloud | gpt-oss:120b-cloud gpt-oss:20b-cloud qwen3-coder:480b-cloud |
Dynamic / model-dependent | Not required | Yes | Open |
These entries can still be useful, but they are aggregators, trial credits, retiring tiers, or metered services — not permanent Provider Free Tiers.
| Provider | Access type | Models | Published limits | Card | OpenAI compatible | Get API key |
|---|---|---|---|---|---|---|
| Novita AI | Metered access | Llama 3.1 8B Instruct OpenAI: GPT OSS 20B BAAI:BGE-M3 |
Dynamic / model-dependent | Not required | Yes | Open |
| Moonshot AI (Kimi) | Metered access | kimi-k2-0905-preview moonshot-v1-8k moonshot-v1-128k |
Dynamic / model-dependent | Not required | Yes | Open |
| Cerebras Inference | Free trial credit | gpt-oss-120b gemma-4-31b |
5 RPM | Required | Yes | Open |
| Vercel AI Gateway | Free trial credit | Free Tier eligible model subset openai/gpt-oss-120b moonshotai/kimi-k2 |
Dynamic / model-dependent | Not required | Yes | Open |
| IBM watsonx.ai | Free trial credit | ibm/granite-3-8b-instruct meta-llama/llama-3-3-70b-instruct mistralai/mistral-large |
Dynamic / model-dependent | Not required | No | Open |
| OpenRouter | Free model aggregator | Model IDs ending in :free openrouter/free |
20 RPM, 50 requests/day | Not required | Yes | Open |
| GitHub Models | Retired free tier Retired 2026年07月30日 |
Retired — the model catalog is gone | Dynamic / model-dependent | Not required | Yes | Closed to new users |
| Together AI | Metered access | meta-llama/Llama-3.3-70B-Instruct-Turbo openai/gpt-oss-120b Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo |
Dynamic / model-dependent | Not required | Yes | Open |
| Nebius Token Factory | Metered access | deepseek-ai/DeepSeek-V3 meta-llama/Llama-3.3-70B-Instruct Qwen/Qwen3-235B-A22B |
Dynamic / model-dependent | Not required | Yes | Open |
| Perplexity API | Metered access | sonar sonar-pro sonar-reasoning |
50 RPM | Not required | Yes | Open |
| DeepInfra | Metered access | deepseek-ai/DeepSeek-V3 Qwen/Qwen3-Next-80B-A3B-Instruct meta-llama/Llama-4-Scout-17B-16E |
Dynamic / model-dependent | Not required | Yes | Open |
| Chutes | Metered access | zai-org/GLM-5 Qwen/Qwen3-32B unsloth/Mistral-Nemo-Instruct-2407 |
Dynamic / model-dependent | Not required | Yes | Open |
| Scaleway Generative APIs | Metered access | llama-3.3-70b-instruct gpt-oss-120b qwen3-coder-30b-a3b-instruct |
Dynamic / model-dependent | Required | Yes | Open |
Limits marked dynamic or model-dependent are intentionally not replaced with guessed numbers. Follow the linked official source for the current quota.
After creating your own Groq API key, the OpenAI SDK needs only a different Base URL and model id:
import os from openai import OpenAI client = OpenAI( api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1", ) response = client.chat.completions.create( model="llama-3.3-70b-versatile", messages=[{"role": "user", "content": "Say hello in one sentence."}], ) print(response.choices[0].message.content)
For coding agents, generate a client-specific configuration that reads the key from your environment:
npx free-llm-api setup claude-code
Client guides: Claude Code · Codex CLI · Cline · all clients.
Every row on this page is one file, published with no key and Access-Control-Allow-Origin: *:
curl -s https://xyzs996.github.io/free-llm-api/providers.json
A row carries base_url, openai_compatible, credit_card_required, browser_check, the published limits, and the official_sources each limit was read from — every one of them dated. That is enough for a program to pick a provider instead of a person re-reading the table above. The 23 that speak OpenAI and ask for no card:
curl -s https://xyzs996.github.io/free-llm-api/providers.json | jq -r '.[] | select(.openai_compatible and .credit_card_required == false) | "\(.id)\t\(.base_url)"'
For a build that should not reach a Pages host, the same bytes are on a CDN. @main follows the branch; put a tag there to freeze it.
Open the browser key checker. Nothing is installed or stored: the request goes from your browser straight to the chosen provider. Its Content Security Policy allows the 26 catalog origins and no analytics or project server. 21 providers answer cross-origin browser requests; blocked providers get an equivalent curl command.
These three answer a question this page only gets to in passing. Each opens with a direct answer, then shows the reviewed figures behind it and the date each was read:
- Which free LLM APIs work without a credit card — and what are the published limits? — every permanent free tier that asks for no card, with the limit each one publishes and the ones that publish none.
- What are the Gemini, Groq and OpenRouter free tier rate limits right now? — the per-minute and per-day figures each one documents, why one of the three publishes no single number, and what that means for quoting it.
- Which free LLM APIs are OpenAI compatible, and what base URL do I point my coding agent at? — the base URL for every compatible provider, the one that is not a drop-in, and what the compatible flag does not promise.
- Every limit and lifecycle claim links to an official source and carries a review date.
- Trial credit, metered access, aggregators, and retiring tiers are separated from permanent free tiers.
- No working API keys are stored or distributed. Use environment variables for your own credentials.
- A probe describes one sampled request, not provider-wide uptime. A
429does not reveal the key's remaining quota.
Run probes explicitly outside CI. Keys are read only from the provider environment variable:
GROQ_API_KEY=YOUR_API_KEY npm run probe -- --provider groq
The ignored data/probe-output.json contains only a bounded classification, status, latency, and timestamp — never the key, response body, or raw exception. Read the full methodology.
Week of 2026年08月22日. First re-check since publication. Every source was read again: GitHub Models is gone, Novita has no free rows left, and Groq dropped both Llama models from the free table. Two entries could not be re-read and keep their old check date.
- Lifecycle — GitHub Models: Retired on 2026年07月30日 as announced: playground, model catalog, inference API and BYOK endpoints are all gone. Kept as a tombstone entry.
- Lifecycle — Novita AI: inclusionai/Ling-3.0-flash and Mind Lab Macaron V1 Venti are no longer priced at zero; no model on the pricing table is free any more. Moved to metered access.
- Lifecycle — Moonshot AI (Kimi): Docs moved to platform.kimi.ai. The lowest tier still requires a 1ドル recharge before any call goes through, so the entry moved out of the permanent free tier list into metered access.
- Limits changed — GroqCloud: llama-3.3-70b-versatile and llama-3.1-8b-instant left the free plan table. openai/gpt-oss-safeguard-20b, qwen/qwen3.6-27b and groq/compound-mini are in it. The RPM/RPD columns now quote openai/gpt-oss-120b at 30 RPM / 1K RPD.
- Limits changed — Cerebras Inference: zai-glm-4.7 left the Free Trial table; only gpt-oss-120b and gemma-4-31b remain at 5 RPM / 30K TPM / 1M TPH / 1M TPD.
- Added — Z.AI Open Platform: GLM-4.6V-Flash, a vision model, is now listed at 0ドル across every pricing column alongside GLM-4.7-Flash and GLM-4.5-Flash.
- Added — SambaNova Cloud: DeepSeek-V3.2 and gemma-4-31B-it appear in the free table as preview models, on the same 20 RPM / 20 RPD / 200K TPD numbers.
- Corrected (4): Pollinations.AI, Nebius Token Factory, Vercel AI Gateway, Cloudflare Workers AI
Every entry above is dated and sourced in the catalog below. Full history: data/changelog.json.
A free tier closing turns into a bill, and this catalog stops where the free tier does. What the paid ones actually cost, per million tokens, each figure carrying the sentence it was published in: the field notes cost table.
Corrections are the contribution this project runs on: a limit that moved, a provider that closed signups, or a link that died. See CONTRIBUTING.md and the correction form.
Not sure enough to file one? Say it in one line — one field, no link, no screenshot, no source. Somebody else can go find the page.
data/providers.jsonis the reviewed source dataset.llms.txtis the whole catalog as one text file: every provider on one line with its limit, its base URL and the date that limit was read.data/changelog.jsonrecords weekly changes.README.md,docs/providers.json, and the static pages are generated deterministically.npm run validaterejects stated quotas without an official source.
Run the generated site locally:
npm run render && npm run serveOpen http://127.0.0.1:4173. Node.js 20+ is required; there are no runtime dependencies and no API keys are needed.
This repository contains no working credentials. Keep probe keys in environment variables and redact Authorization headers from reports. See SECURITY.md.
- Free Tier LLM Router combines your own provider keys behind one local endpoint with controlled failover.
- AI Coding Field Notes publishes every figure it has cited — anything carrying a unit — as JSON and CSV, each row paired with the sentence it came from, plus the write-ups behind them.
If rotating free-tier keys and handling different limits becomes the work, create a PekPik API account for one OpenAI-compatible hosted endpoint. The free directory above remains usable without it.