| OpenCode Zen | Promo free (limited time) | Big Pickle, DeepSeek V4 Flash, MiMo-V2.5, Laguna S 2.1, Ling-3.0-tiny, LongCat-2.0, North Mini Code, Nemotron 3 Ultra | — | — | — | — | — | Not published | Free models are limited-time promos; some may use data for training. No credit card. OpenAI-compatible base URL opencode.ai/zen/v1. List free models: opencode models | grep -i free |
| OpenRouter | Rate-limited free | cohere/north-mini-code:free, google/gemma-4-26b-a4b-it:free, google/gemma-4-31b-it:free, inclusionai/ling-3.0-tiny:free, liquid/lfm-2.5-2.6b:free, nvidia/nemotron-3-nano-30b-a3b:free, nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free, nvidia/nemotron-3-super-120b-a12b:free, nvidia/nemotron-3-ultra-550b-a55b:free, nvidia/nemotron-3.5-content-safety:free, nvidia/nemotron-3.5-lightning:free, nvidia/nemotron-nano-12b-v2-vl:free, nvidia/nemotron-nano-9b-v2:free, openai/gpt-oss-20b:free, poolside/laguna-s-2.1:free, poolside/laguna-xs-2.1:free | — | — | — | — | — | Not published | Gateway, not a model owner. Shared best-effort capacity, no SLA. BYOK program: 1M free routing requests/month with your own provider keys. |
| Google AI Studio (Gemini API) | Rate-limited free | gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.5-pro, gemini-3-flash-preview, gemini-3.1-flash-lite, gemini-3.1-flash-lite-preview, gemini-3.1-pro-preview, gemini-3.1-pro-preview-customtools, gemini-3.5-flash, gemini-3.5-flash-lite, gemini-3.6-flash, gemini-flash-latest, gemini-flash-lite-latest, gemini-omni-flash-preview, gemini-pro-latest, gemma-4-26b-a4b-it, gemma-4-31b-it | — | — | — | — | — | Not published | Free forever, no card, most generous major-provider tier. Live-verified against the v1beta API with an API key (2026-08-12) — model list and context windows are real; per-account quota varies by account age (see AI Studio rate-limit dashboard). Free-tier data may be used for training. |
| Groq | Rate-limited free | Llama 3.1 / 3.3, Llama 4 Scout, Qwen 3, Kimi K2, GPT-OSS 120B, Gemma 2 9B | — | — | — | — | — | Not published | LPU hardware, fastest time-to-first-token. Limits per API key per model; some models get half allowance. No credit card. Limits visible in x-ratelimit-* response headers. |
| Cerebras | Rate-limited free (Free Trial) | gpt-oss-120b, zai-glm-4.7, gemma-4-31b | — | — | — | — | — | Not published | Wafer-scale hardware, 2,000+ tok/s. Free tier caps context at 8K (temporary). Per-model limits on a rotating shortlist. No credit card. |
| Cloudflare Workers AI | Rate-limited free | ~80 free models (Llama 3.x/4, Qwen, Gemma, DeepSeek-R1 distills, FLUX, Whisper, BGE embeddings) | — | — | — | — | — | Not published | Neurons are normalized GPU-compute units, shared pool across all models. Resets 00:00 UTC. Hard stop when pool exhausted, errors not overage. No credit card. |
| Hugging Face | Rate-limited free | Llama 3.2 8B (Serverless), Qwen 2.5 7B (Serverless), Mistral 7B (Serverless), Inference Providers gateway (15+ partners) | — | — | — | — | — | Not published | Two products: Serverless API (free, rate-limited) and Inference Providers gateway. Cold starts 10-30s on unpopular models. Limits not published as fixed numbers. |
| Mistral La Plateforme | Rate-limited free (Experiment) | Mistral Large, Codestral, All API models (Experiment tier) | — | — | — | — | — | Not published | Phone (SMS) verification required, no card. Exact limits no longer published - see Admin Console Limits per workspace. Evaluation tier, not production. |
| SambaNova Cloud | Rate-limited free | DeepSeek-V3.1, MiniMax-M2.7, Gemma 4 31B preview | — | — | — | — | — | Not published | OpenAI-compatible base URL api.sambanova.ai/v1. Preview models can be pulled at any time. No credit card, no phone verification. |
| NVIDIA NIM (build.nvidia.com) | Trial credits | GLM-5, Kimi-2.5, NIM-packaged open models | — | — | — | — | — | Not published | Credit-based, not a rate-limited-free tier. Larger grants need corporate email, tie to ~90-day evaluation windows. |
| Z.AI (Zhipu) | Rate-limited free | GLM-5.1, GLM-4.5-Flash, GLM-4.7-Flash, GLM-4.6V-Flash (vision) | — | — | — | — | — | Not published | Free-tier limits revised twice in the past year - verify. Peak-hour throttling. Flash models are free regardless of tier. OpenAI-compatible. |
| Together AI | Trial credits | 200+ open models (Llama, Qwen, DeepSeek) | — | — | — | — | — | Not published | One-time credit, not forever-free. Card required once credits run out. Startup programs can grant far more. |
| DeepInfra | Trial credits | 40+ open models (Llama, Qwen, DeepSeek, Mistral) | — | — | — | — | — | Not published | 200 concurrent requests per model - high ceiling. Credit-based, no daily cap while credits last. OpenAI-compatible. |