Skip to content

Navigation Menu

Sign in
Sign up

Repository files navigation

Free LLM APIs

English · 简体中文

CI License: MIT providers sources checked

Permanent free tiers, no-card options, direct API key links, models, and verified limits — every claim points to an official source.

13 permanent provider free tiers · 13 require no credit card · 13 are OpenAI compatible · sources reviewed 2026年08月22日. No keys are distributed here. A probe describes one sampled request, not provider-wide uptime.

Browse the live directory · Pick by model · Set up a coding agent · Check your own key

Filterable LLM free-tier status page

Star this repository to bookmark the dataset and follow releases. A star changes nothing about any provider's keys, credits, or limits, and this project gives nothing in return for one.

Pick a free API by goal

Goal Pick Why it appears here Start
Highest published daily request limit GroqCloud 1,000 requests/day published Get API key
Highest published requests per minute SiliconFlow 1,000 RPM published Get API key
Works with the browser key checker Google Gemini API Browser CORS check supported Get API key
Fast path for coding agents GroqCloud Documented OpenAI-compatible coding setup Get API key

These are rule-based shortcuts, not paid placements. Open the filterable directory for all 26 providers.

Free is not the same as cheap enough to keep running, and the number you need is what replaces the free tier once it runs out. Same maintainer, re-reading OpenRouter's whole catalog every day60 models on 2026年08月22日, cheapest five first:

Model Input / M tokens Output / M tokens Design Arena agents
Gemini 3.7 Flash (batch) 0ドル.1875 0ドル.9375 #3 androidnative
Gemini 3 Flash Preview (batch) 0ドル.25 1ドル.50 #9 agenticslides
MiniMax M3 0ドル.30 1ドル.20 #10 python-pptxslides
Gemini 3.6 Flash (batch) 0ドル.375 1ドル.875 #6 agenticgamedev
GLM 4.7 0ドル.40 1ドル.75 #27 androidnative

A (batch) row is the queued price, not the interactive one — an agent waiting on the reply pays the other number. All 60 rows as JSON or CSV, re-read tomorrow.

A row price is not a request price. 12 of those 60 rows carry a second, higher rate card that switches on when the prompt gets long — and it applies to every token in the request, including the ones before the threshold, so a prompt one token over the line costs about double one token under it. The thresholds are 200k and 272k prompt tokens. The counter-intuitive part: the bigger the advertised context window, the smaller the share of it the advertised price covers — x-ai/grok-4.20 advertises 2 000k of context at 1ドル.25 per million input and steps to 2ドル.50 at 200k, so 90% of that window bills at a number that is not in its row. None of the five above is one of them, which is most of why they are the cheap ones. Both numbers ride on every row of the JSON and CSV as long_context_from and long_input_per_million.

A row price is not an hour price either. DeepSeek's V4 line runs two rate cards off the clock: peak is 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, and every other hour is off-peak at exactly half. Since 2026年08月23日 the entire weekend is off-peak too — 14 hours a week that most published summaries, and most of the billing code I have read, still charge at 2x. Per million: deepseek-v4-flash input 0ドル.22 off-peak against 0ドル.44 peak, deepseek-v4-pro output 1ドル.98 against 3ドル.96. A nightly batch moved a few hours earlier pays half, and no code change buys that. Which hours, and the two instants where a UTC weekday and a Beijing weekday disagree.

A list price is what a vendor publishes; what a month of it came to is a different number, and only somebody who paid it can tell you. The same repository keeps those too — every figure quoted in a write-up, with the full sentence it came from on every row:

Price Unit The sentence it was published in
0ドル.06 / 0ドル.2 per million At the low end, MiniMax M3 runs 0ドル.06 to 0ドル.2 per million and draws 60% to 70% of its revenue from outside its home market.
0ドル.19 / 5ドル per million tokens Chinese AI models provide a cost-effective alternative to their American counterparts, with input costs as low as 0ドル.19 per million tokens, compared to OpenAI's 5ドル-12.
1ドル per million tokens Top-tier Chinese models such as GLM5.2 and DeepSeek V4 Pro sit near 1ドル per million tokens at inference gross margins of 10% to 20%.
1ドル.25 / 4ドル.25 per million Meta priced Muse Spark 1.1 at 1ドル.25 per million input and 4ドル.25 per million output, roughly 75% and 83% below Anthropic's Opus, and the tradeoff is visible in the benchmarks, since it leads on MCP Atlas and JobBench while trailing on SWE-Bench Pro and DeepSWE 1.1.
3ドル per million input tokens The 3ドル per million input tokens price point means developers should carefully evaluate whether the premium model's capabilities justify the increased costs for their specific use cases.

A 1ドル.43 is never left ambiguous between per million tokens, per month and per seat, because the sentence travels with it. Readable in code as JSON or CSV, or as prose: where the token bill actually goes.

Paying for a model whose bill surprised you? Name it in one line — one field, and it decides which price gets chased next.

Permanent free tiers

These Provider Free Tiers are the main list: they do not expire like trial credits, and none currently require a credit card.

Provider Models Published limits Card OpenAI compatible Get API key
Google Gemini API Free-tier eligibility varies by model
gemini-2.5-flash
gemini-2.5-flash-lite
Dynamic / model-dependent Not required Yes Open
GroqCloud openai/gpt-oss-120b
openai/gpt-oss-20b
openai/gpt-oss-safeguard-20b
30 RPM, 1,000 requests/day Not required Yes Open
SambaNova Cloud DeepSeek-V3.1
Meta-Llama-3.3-70B-Instruct
gpt-oss-120b
20 RPM, 20 requests/day Not required Yes Open
Cohere command-a-03-2025
command-r-plus
embed-v4.0
20 RPM Not required Yes Open
Cloudflare Workers AI @cf/meta/llama-3.3-70b-instruct-fp8-fast
@cf/openai/gpt-oss-120b
@cf/qwen/qwen2.5-coder-32b-instruct
Dynamic / model-dependent Not required Yes Open
Hugging Face Inference Providers deepseek-ai/DeepSeek-V3-0324
openai/gpt-oss-120b
200+ models routed across partner providers
Dynamic / model-dependent Not required Yes Open
SiliconFlow Qwen/Qwen3-8B
THUDM/GLM-4-9B-0414
deepseek-ai/DeepSeek-R1
1,000 RPM Not required Yes Open
Fireworks AI accounts/fireworks/models/llama-v3p3-70b-instruct
accounts/fireworks/models/gpt-oss-120b
10 RPM Not required Yes Open
Z.AI Open Platform GLM-4.7-Flash
GLM-4.5-Flash
GLM-4.6V-Flash
Dynamic / model-dependent Not required Yes Open
Mistral La Plateforme mistral-small-latest
open-mistral-nemo
codestral-latest
Dynamic / model-dependent Not required Yes Open
Alibaba Cloud Model Studio qwen-plus
qwen-turbo
qwen3-coder-plus
Dynamic / model-dependent Not required Yes Open
Pollinations.AI openai
mistral
Community-hosted open models
4 RPM Not required Yes Open
Ollama Cloud gpt-oss:120b-cloud
gpt-oss:20b-cloud
qwen3-coder:480b-cloud
Dynamic / model-dependent Not required Yes Open

Other access options

These entries can still be useful, but they are aggregators, trial credits, retiring tiers, or metered services — not permanent Provider Free Tiers.

Provider Access type Models Published limits Card OpenAI compatible Get API key
Novita AI Metered access Llama 3.1 8B Instruct
OpenAI: GPT OSS 20B
BAAI:BGE-M3
Dynamic / model-dependent Not required Yes Open
Moonshot AI (Kimi) Metered access kimi-k2-0905-preview
moonshot-v1-8k
moonshot-v1-128k
Dynamic / model-dependent Not required Yes Open
Cerebras Inference Free trial credit gpt-oss-120b
gemma-4-31b
5 RPM Required Yes Open
Vercel AI Gateway Free trial credit Free Tier eligible model subset
openai/gpt-oss-120b
moonshotai/kimi-k2
Dynamic / model-dependent Not required Yes Open
IBM watsonx.ai Free trial credit ibm/granite-3-8b-instruct
meta-llama/llama-3-3-70b-instruct
mistralai/mistral-large
Dynamic / model-dependent Not required No Open
OpenRouter Free model aggregator Model IDs ending in :free
openrouter/free
20 RPM, 50 requests/day Not required Yes Open
GitHub Models Retired free tier
Retired 2026年07月30日
Retired — the model catalog is gone Dynamic / model-dependent Not required Yes Closed to new users
Together AI Metered access meta-llama/Llama-3.3-70B-Instruct-Turbo
openai/gpt-oss-120b
Qwen/Qwen3-Coder-480B-A35B-Instruct-Turbo
Dynamic / model-dependent Not required Yes Open
Nebius Token Factory Metered access deepseek-ai/DeepSeek-V3
meta-llama/Llama-3.3-70B-Instruct
Qwen/Qwen3-235B-A22B
Dynamic / model-dependent Not required Yes Open
Perplexity API Metered access sonar
sonar-pro
sonar-reasoning
50 RPM Not required Yes Open
DeepInfra Metered access deepseek-ai/DeepSeek-V3
Qwen/Qwen3-Next-80B-A3B-Instruct
meta-llama/Llama-4-Scout-17B-16E
Dynamic / model-dependent Not required Yes Open
Chutes Metered access zai-org/GLM-5
Qwen/Qwen3-32B
unsloth/Mistral-Nemo-Instruct-2407
Dynamic / model-dependent Not required Yes Open
Scaleway Generative APIs Metered access llama-3.3-70b-instruct
gpt-oss-120b
qwen3-coder-30b-a3b-instruct
Dynamic / model-dependent Required Yes Open

Limits marked dynamic or model-dependent are intentionally not replaced with guessed numbers. Follow the linked official source for the current quota.

Quick start

After creating your own Groq API key, the OpenAI SDK needs only a different Base URL and model id:

import os
from openai import OpenAI
client = OpenAI(
 api_key=os.environ["GROQ_API_KEY"],
 base_url="https://api.groq.com/openai/v1",
)
response = client.chat.completions.create(
 model="llama-3.3-70b-versatile",
 messages=[{"role": "user", "content": "Say hello in one sentence."}],
)
print(response.choices[0].message.content)

For coding agents, generate a client-specific configuration that reads the key from your environment:

npx free-llm-api setup claude-code

Client guides: Claude Code · Codex CLI · Cline · all clients.

Pull the catalog into your own code

Every row on this page is one file, published with no key and Access-Control-Allow-Origin: *:

curl -s https://xyzs996.github.io/free-llm-api/providers.json

A row carries base_url, openai_compatible, credit_card_required, browser_check, the published limits, and the official_sources each limit was read from — every one of them dated. That is enough for a program to pick a provider instead of a person re-reading the table above. The 23 that speak OpenAI and ask for no card:

curl -s https://xyzs996.github.io/free-llm-api/providers.json | jq -r '.[] | select(.openai_compatible and .credit_card_required == false) | "\(.id)\t\(.base_url)"'

For a build that should not reach a Pages host, the same bytes are on a CDN. @main follows the branch; put a tag there to freeze it.

Check a key you already have

Open the browser key checker. Nothing is installed or stored: the request goes from your browser straight to the chosen provider. Its Content Security Policy allows the 26 catalog origins and no analytics or project server. 21 providers answer cross-origin browser requests; blocked providers get an equivalent curl command.

Answered in full, with the sources

These three answer a question this page only gets to in passing. Each opens with a direct answer, then shows the reviewed figures behind it and the date each was read:

Why trust this list

  • Every limit and lifecycle claim links to an official source and carries a review date.
  • Trial credit, metered access, aggregators, and retiring tiers are separated from permanent free tiers.
  • No working API keys are stored or distributed. Use environment variables for your own credentials.
  • A probe describes one sampled request, not provider-wide uptime. A 429 does not reveal the key's remaining quota.

Run probes explicitly outside CI. Keys are read only from the provider environment variable:

GROQ_API_KEY=YOUR_API_KEY npm run probe -- --provider groq

The ignored data/probe-output.json contains only a bounded classification, status, latency, and timestamp — never the key, response body, or raw exception. Read the full methodology.

Changed this week

Week of 2026年08月22日. First re-check since publication. Every source was read again: GitHub Models is gone, Novita has no free rows left, and Groq dropped both Llama models from the free table. Two entries could not be re-read and keep their old check date.

  • Lifecycle — GitHub Models: Retired on 2026年07月30日 as announced: playground, model catalog, inference API and BYOK endpoints are all gone. Kept as a tombstone entry.
  • Lifecycle — Novita AI: inclusionai/Ling-3.0-flash and Mind Lab Macaron V1 Venti are no longer priced at zero; no model on the pricing table is free any more. Moved to metered access.
  • Lifecycle — Moonshot AI (Kimi): Docs moved to platform.kimi.ai. The lowest tier still requires a 1ドル recharge before any call goes through, so the entry moved out of the permanent free tier list into metered access.
  • Limits changed — GroqCloud: llama-3.3-70b-versatile and llama-3.1-8b-instant left the free plan table. openai/gpt-oss-safeguard-20b, qwen/qwen3.6-27b and groq/compound-mini are in it. The RPM/RPD columns now quote openai/gpt-oss-120b at 30 RPM / 1K RPD.
  • Limits changed — Cerebras Inference: zai-glm-4.7 left the Free Trial table; only gpt-oss-120b and gemma-4-31b remain at 5 RPM / 30K TPM / 1M TPH / 1M TPD.
  • Added — Z.AI Open Platform: GLM-4.6V-Flash, a vision model, is now listed at 0ドル across every pricing column alongside GLM-4.7-Flash and GLM-4.5-Flash.
  • Added — SambaNova Cloud: DeepSeek-V3.2 and gemma-4-31B-it appear in the free table as preview models, on the same 20 RPM / 20 RPD / 200K TPD numbers.
  • Corrected (4): Pollinations.AI, Nebius Token Factory, Vercel AI Gateway, Cloudflare Workers AI

Every entry above is dated and sourced in the catalog below. Full history: data/changelog.json.

A free tier closing turns into a bill, and this catalog stops where the free tier does. What the paid ones actually cost, per million tokens, each figure carrying the sentence it was published in: the field notes cost table.

Contributing

Corrections are the contribution this project runs on: a limit that moved, a provider that closed signups, or a link that died. See CONTRIBUTING.md and the correction form.

Not sure enough to file one? Say it in one line — one field, no link, no screenshot, no source. Somebody else can go find the page.

Data and local development

  • data/providers.json is the reviewed source dataset.
  • llms.txt is the whole catalog as one text file: every provider on one line with its limit, its base URL and the date that limit was read.
  • data/changelog.json records weekly changes.
  • README.md, docs/providers.json, and the static pages are generated deterministically.
  • npm run validate rejects stated quotas without an official source.

Run the generated site locally:

npm run render && npm run serve

Open http://127.0.0.1:4173. Node.js 20+ is required; there are no runtime dependencies and no API keys are needed.

Security

This repository contains no working credentials. Keep probe keys in environment variables and redact Authorization headers from reports. See SECURITY.md.

Star history

Star History Chart

Related projects

  • Free Tier LLM Router combines your own provider keys behind one local endpoint with controlled failover.
  • AI Coding Field Notes publishes every figure it has cited — anything carrying a unit — as JSON and CSV, each row paired with the sentence it came from, plus the write-ups behind them.

Need one stable endpoint?

If rotating free-tier keys and handling different limits becomes the work, create a PekPik API account for one OpenAI-compatible hosted endpoint. The free directory above remains usable without it.

About

Free LLM API providers list: verified free tier limits, API keys with no credit card, and OpenAI-compatible endpoints for developers.

Topics

Resources

Contributing

Security policy

Stars

8 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages

AltStyle によって変換されたページ (->オリジナル) /