# Free Inference > Every provider that gives developers free API access to LLM inference — usable from a harness, agent, or CLI. Web-chat-only free tiers (ChatGPT, Claude.ai, Gemini app, deepseek.com, Grok, Copilot) are excluded on purpose. > Last verified: 2026-08-12 · 13 providers · 70 free models. > Machine data: https://freeinference.dev/data/providers.json (validate against https://freeinference.dev/data/schema.json). ## Providers - [OpenCode Zen](https://opencode.ai/docs/zen): Promo free (limited time) · models: Big Pickle, DeepSeek V4 Flash, MiMo-V2.5, Laguna S 2.1, Ling-3.0-tiny, LongCat-2.0, North Mini Code, Nemotron 3 Ultra - Big Pickle: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - DeepSeek V4 Flash: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - MiMo-V2.5: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - Laguna S 2.1: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - Ling-3.0-tiny: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - LongCat-2.0: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - North Mini Code: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - Nemotron 3 Ultra: cost $0 (promo) · context 256K - 1M · RPM Not published · TPM Not published · RPD Not published · day tokens Not published · verified 2026-08-12 - [OpenRouter](https://openrouter.ai/models?max_price=0): Rate-limited free · models: cohere/north-mini-code:free, google/gemma-4-26b-a4b-it:free, google/gemma-4-31b-it:free, inclusionai/ling-3.0-tiny:free, liquid/lfm-2.5-2.6b:free, nvidia/nemotron-3-nano-30b-a3b:free, nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free, nvidia/nemotron-3-super-120b-a12b:free, nvidia/nemotron-3-ultra-550b-a55b:free, nvidia/nemotron-3.5-content-safety:free, nvidia/nemotron-3.5-lightning:free, nvidia/nemotron-nano-12b-v2-vl:free, nvidia/nemotron-nano-9b-v2:free, openai/gpt-oss-20b:free, poolside/laguna-s-2.1:free, poolside/laguna-xs-2.1:free - cohere/north-mini-code:free: cost $0 · context 250K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - google/gemma-4-26b-a4b-it:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - google/gemma-4-31b-it:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - inclusionai/ling-3.0-tiny:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - liquid/lfm-2.5-2.6b:free: cost $0 · context 125K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3-nano-30b-a3b:free: cost $0 · context 250K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free: cost $0 · context 250K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3-super-120b-a12b:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3-ultra-550b-a55b:free: cost $0 · context 1000000 · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3.5-content-safety:free: cost $0 · context 125K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-3.5-lightning:free: cost $0 · context 1000000 · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-nano-12b-v2-vl:free: cost $0 · context 125K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - nvidia/nemotron-nano-9b-v2:free: cost $0 · context 125K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - openai/gpt-oss-20b:free: cost $0 · context 128K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - poolside/laguna-s-2.1:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - poolside/laguna-xs-2.1:free: cost $0 · context 256K · RPM 20 · TPM Provider-dependent · RPD 50 (below $10 credits) / 1,000 ($10+ credits) · day tokens Not published · verified 2026-08-12 - [Google AI Studio (Gemini API)](https://ai.google.dev/gemini-api/docs/rate-limits): Rate-limited free · models: gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3-flash-preview, gemini-3.1-flash-lite, gemini-3.1-flash-lite-preview, gemini-3.1-pro-preview, gemini-3.1-pro-preview-customtools, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-flash-latest, gemini-flash-lite-latest, gemini-omni-flash-preview, gemini-pro-latest, gemma-4-26b-a4b-it, gemma-4-31b-it - gemini-2.5-flash: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-2.5-flash-lite: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-2.5-pro: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3-flash-preview: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.1-flash-lite: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.1-flash-lite-preview: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.1-pro-preview: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.1-pro-preview-customtools: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.5-flash: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.5-flash-lite: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-3.6-flash: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-flash-latest: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-flash-lite-latest: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-omni-flash-preview: cost $0 · context 128K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemini-pro-latest: cost $0 · context 1024K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemma-4-26b-a4b-it: cost $0 · context 256K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - gemma-4-31b-it: cost $0 · context 256K · RPM Varies by account · TPM Varies by account · RPD Varies by account · day tokens Not published · verified 2026-08-12 - [Groq](https://console.groq.com/docs/rate-limits): Rate-limited free · models: Llama 3.1 / 3.3, Llama 4 Scout, Qwen 3, Kimi K2, GPT-OSS 120B, Gemma 2 9B - Llama 3.1 / 3.3: cost $0 · context 128K · RPM 30 (15 for some) · TPM 6K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - Llama 4 Scout: cost $0 · context 128K · RPM 30 (15 for some) · TPM 6K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - Qwen 3: cost $0 · context 128K · RPM 30 (15 for some) · TPM 6K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - Kimi K2: cost $0 · context 128K · RPM 30 (15 for some) · TPM 6K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - GPT-OSS 120B: cost $0 · context 128K · RPM 30 (15 for some) · TPM 6K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - Gemma 2 9B: cost $0 · context 128K · RPM 30 (15 for some) · TPM 15K · RPD 14,400 (org-wide) · day tokens Not published · verified 2026-08-12 - [Cerebras](https://inference-docs.cerebras.ai/support/rate-limits): Rate-limited free (Free Trial) · models: gpt-oss-120b, zai-glm-4.7, gemma-4-31b - gpt-oss-120b: cost $0 · context 8K cap on free tier · RPM 5 · TPM 30K · RPD Not published · day tokens 1M tokens/day (per model) · verified 2026-08-12 - zai-glm-4.7: cost $0 · context 8K cap on free tier · RPM 5 · TPM 30K · RPD Not published · day tokens 1M tokens/day (per model) · verified 2026-08-12 - gemma-4-31b: cost $0 · context 8K cap on free tier · RPM 5 · TPM 30K · RPD Not published · day tokens 1M tokens/day (per model) · verified 2026-08-12 - [Cloudflare Workers AI](https://developers.cloudflare.com/workers-ai/platform/limits): Rate-limited free · models: ~80 free models (Llama 3.x/4, Qwen, Gemma, DeepSeek-R1 distills, FLUX, Whisper, BGE embeddings) - ~80 free models (Llama 3.x/4, Qwen, Gemma, DeepSeek-R1 distills, FLUX, Whisper, BGE embeddings): cost $0 · context 8K-128K+ (model-dependent) · RPM 300 text gen (720-1,500 other tasks) · TPM Model-dependent · RPD Not published · day tokens 10,000 neurons/day · verified 2026-08-12 - [Hugging Face](https://huggingface.co/docs/inference-providers): Rate-limited free · models: Llama 3.2 8B (Serverless), Qwen 2.5 7B (Serverless), Mistral 7B (Serverless), Inference Providers gateway (15+ partners) - Llama 3.2 8B (Serverless): cost $0 · context Model-dependent · RPM ~300 req/hour (Serverless) · TPM Model max context · RPD Not published · day tokens Small monthly credit pool (PRO: 2M credits/mo) · verified 2026-08-12 - Qwen 2.5 7B (Serverless): cost $0 · context Model-dependent · RPM ~300 req/hour (Serverless) · TPM Model max context · RPD Not published · day tokens Small monthly credit pool (PRO: 2M credits/mo) · verified 2026-08-12 - Mistral 7B (Serverless): cost $0 · context Model-dependent · RPM ~300 req/hour (Serverless) · TPM Model max context · RPD Not published · day tokens Small monthly credit pool (PRO: 2M credits/mo) · verified 2026-08-12 - Inference Providers gateway (15+ partners): cost $0 · context Model-dependent · RPM ~300 req/hour (Serverless) · TPM Model max context · RPD Not published · day tokens Small monthly credit pool (PRO: 2M credits/mo) · verified 2026-08-12 - [Mistral La Plateforme](https://docs.mistral.ai): Rate-limited free (Experiment) · models: Mistral Large, Codestral, All API models (Experiment tier) - Mistral Large: cost $0 (Experiment) · context Model-dependent · RPM ~1 req/sec · TPM Not published · RPD Not published · day tokens ~1B tokens/month · verified 2026-08-12 - Codestral: cost $0 (Experiment) · context Model-dependent · RPM ~1 req/sec · TPM Not published · RPD Not published · day tokens ~1B tokens/month · verified 2026-08-12 - All API models (Experiment tier): cost $0 (Experiment) · context Model-dependent · RPM ~1 req/sec · TPM Not published · RPD Not published · day tokens ~1B tokens/month · verified 2026-08-12 - [SambaNova Cloud](https://cloud.sambanova.ai/apis): Rate-limited free · models: DeepSeek-V3.1, MiniMax-M2.7, Gemma 4 31B preview - DeepSeek-V3.1: cost $0 · context 128K · RPM 20 · TPM Not published · RPD 20 · day tokens 200K tokens/day · verified 2026-08-12 - MiniMax-M2.7: cost $0 · context 128K · RPM 20 · TPM Not published · RPD 20 · day tokens 200K tokens/day · verified 2026-08-12 - Gemma 4 31B preview: cost $0 · context 128K · RPM 20 · TPM Not published · RPD 20 · day tokens 200K tokens/day · verified 2026-08-12 - [NVIDIA NIM (build.nvidia.com)](https://build.nvidia.com): Trial credits · models: GLM-5, Kimi-2.5, NIM-packaged open models - GLM-5: cost $0 (trial credits) · context Model-dependent · RPM ~40 · TPM Not published · RPD Not published · day tokens ~1,000 credits at signup (up to ~5,000) · verified 2026-08-12 - Kimi-2.5: cost $0 (trial credits) · context Model-dependent · RPM ~40 · TPM Not published · RPD Not published · day tokens ~1,000 credits at signup (up to ~5,000) · verified 2026-08-12 - NIM-packaged open models: cost $0 (trial credits) · context Model-dependent · RPM ~40 · TPM Not published · RPD Not published · day tokens ~1,000 credits at signup (up to ~5,000) · verified 2026-08-12 - [Z.AI (Zhipu)](https://z.ai): Rate-limited free · models: GLM-5.1, GLM-4.5-Flash, GLM-4.7-Flash, GLM-4.6V-Flash (vision) - GLM-5.1: cost $0 · context Up to 203K · RPM 3 (burst) · TPM Not published · RPD 1,000 (1K RPD tier) · day tokens Not published · verified 2026-08-12 - GLM-4.5-Flash: cost $0 · context Up to 203K · RPM 3 (burst) · TPM Not published · RPD 1,000 · day tokens Not published · verified 2026-08-12 - GLM-4.7-Flash: cost $0 · context Up to 203K · RPM 3 (burst) · TPM Not published · RPD 1,000 · day tokens Not published · verified 2026-08-12 - GLM-4.6V-Flash (vision): cost $0 · context Up to 203K · RPM 3 (burst) · TPM Not published · RPD 1,000 · day tokens Not published · verified 2026-08-12 - [Together AI](https://www.together.ai/pricing): Trial credits · models: 200+ open models (Llama, Qwen, DeepSeek) - 200+ open models (Llama, Qwen, DeepSeek): cost $0 (trial credits) · context Model-dependent · RPM ~60 (Build tier) · TPM ~100K (Build tier) · RPD Not published · day tokens $25 one-time credit · verified 2026-08-12 - [DeepInfra](https://docs.deepinfra.com/account/rate-limits): Trial credits · models: 40+ open models (Llama, Qwen, DeepSeek, Mistral) - 40+ open models (Llama, Qwen, DeepSeek, Mistral): cost $0 (trial credits) · context Model-dependent · RPM ~60 (varies by model) · TPM Not published · RPD Not published · day tokens $5 signup credits · verified 2026-08-12 ## Definitions - RPM: Requests per minute — API calls in any 60-second window. - TPM: Tokens per minute — input + output tokens across all calls in a minute. - RPD: Requests per day — hard daily cap, resets on the provider's clock. - Day tokens: Daily or periodic token/compute quota (some providers meter a token pool per day, month, or one-time credits instead). - Timeout: Max request duration before the provider kills the connection. - Context: Maximum context window available on the free tier. ## Caveats Limits change without notice and vary per model, account age, region, and peak hours. Treat this data as a map, not a contract — verify before you architect on it.