

Groq
Groq pricing bills per million tokens, in arrears, with no prepaid balance and no credit system.
Updated on:
Groq pricing: the speed vendor charges the going rate
Groq pricing bills per million tokens, in arrears, with no prepaid balance and no credit system. Four chat models carry a published rate. Everything else, including both Llama models and MiniMax M2.7, says Contact Sales. Speed is the pitch, but the rates sit mid-pack, and the marketing site no longer hosts a pricing page.
Key takeaways
GPT-OSS 120B costs $0.15 per million input tokens and $0.60 output on Groq, exactly the rate Together AI, Amazon Bedrock and Nebius charge for the same model.
Of the 13 models Groq lists publicly, 10 carry a price and 3 say Contact Sales, including both Llama models that made Groq's name.
The free tier grants no credit. It's the paid API at tight limits: 8K tokens per minute and 200K per day on GPT-OSS 120B, against 250K per minute and no daily cap on Developer.
Batch cuts 50% and cached input cuts 50%, and the docs say the two discounts don't stack.
Groq pricing in 2026
Plan | Price | What's included | Metered limit |
|---|---|---|---|
Free | $0 | Same per-token rates, no card required | 30 RPM, 1K RPD, 8K TPM on GPT-OSS 120B |
Developer | $0 + usage | Batch, Flex tier, spend limits, chat support, 100 MB audio files | 1K RPM, 500K RPD, 250K TPM, no daily token cap |
Enterprise | Custom | Performance tier, Llama and MiniMax access, 99.9% availability SLA | Contracted |
Model | Input / 1M | Cached input | Output / 1M |
openai/gpt-oss-120b | $0.15 | $0.075 | $0.60 |
openai/gpt-oss-20b | $0.075 | $0.037 | $0.30 |
openai/gpt-oss-safeguard-20b | $0.075 | $0.037 | $0.30 |
qwen/qwen3.8-27b (preview) | $0.80 | Not supported | $4.00 |
whisper-large-v3 / turbo | $0.111 / $0.04 per audio hour | n/a | n/a |
Orpheus English / Arabic | $22 / $40 per 1M characters | n/a | n/a |
What Groq actually meters
Groq meters three different units and keeps them apart. Text generation bills input and output tokens at separate rates, with a third rate for cached input on the GPT-OSS family. Speech to text bills by the hour of audio processed, not by the token. Text to speech bills per million characters, so Arabic at $40 runs 1.8 times the English rate of $22.
The published catalog is narrower than the pitch suggests. Thirteen models appear on the models page. Llama 3.1 8B, Llama 3.3 70B and MiniMax M2.7 are tagged Enterprise and show Contact Sales in both the price and the rate limit column. That leaves four general chat models with a public per-token rate, two of which are the same GPT-OSS 20B weights under different names.
Cached input is automatic, needs no code change and carries no fee. It applies to the three GPT-OSS models only, expires after two hours unused, and cached tokens don't count against rate limits. The headline discount is 50%: $0.075 against $0.15 on the 120B. On the 20B, Groq publishes $0.037 against $0.075, rounding the half-cent down in your favour.
Groq's own pricing page is gone. groq.com/pricing returns a 308 to the homepage and the plan comparison needs a login, so rates survive only in the docs.
What happens when you hit the limit
Groq blocks rather than throttles, and it bills you before the month ends. New Developer accounts run progressive billing: the card gets charged the moment lifetime usage crosses $1, $10, $100, $500 and $1,000. Past $1,000 the account settles monthly. Indian billing addresses get a different ladder, $1, then $10, then every $100 for good.
Spend limits are the real guardrail. Set a monthly cap and every key in the organisation starts returning a 400 with code blocked_api_access once you reach it. Spend tracking lags 10 to 15 minutes, so Groq says plainly you may overshoot. The limit resets on the first.
Rate limits bite before spend does: you get a 429 with a retry-after header. Flex tier raises limits tenfold at the same price and fails fast with a 498 capacity_exceeded when capacity runs out.
How Groq pricing has changed
Date | Milestone | Source |
|---|---|---|
18 Apr 2026 | MiniMax M2.5 and Qwen3-VL 32B ship Enterprise-only, with no published rate | Vendor |
30 Jan 2026 | PlayAI text to speech retires platform-wide; Orpheus replaces it at $22 and $40 per 1M characters | Vendor |
29 Oct 2025 | GPT-OSS-Safeguard 20B launches at $0.075 input, $0.30 output, caching on from day one | Vendor |
21 Oct 2025 | Prompt caching reaches GPT-OSS 120B, cached input $0.075 against $0.15 | Vendor |
25 Sep 2025 | Prompt caching reaches GPT-OSS 20B, cached input $0.037 against $0.075 | Vendor |
5 Sep 2025 | Kimi K2-0905 lands at $1.00 input and $3.00 output | Vendor |
20 Aug 2025 | Prompt caching introduced, Kimi K2 first, 50% off cached tokens | Vendor |
Customer
Sentiment Highlights
"It's common for me to run workloads in Groq that cost less than $100, while the same workload can approach $1,000 on Bedrock or Gemini"
Groq user comparing provider bills, Hacker News, December 2025
"Cheap enough for now, but of all the companies selling inference at a loss, Cerebras and Groq are probably losing the most per token"
Hacker News commenter on inference economics, November 2025
Explore other providers

Together AI
Infrastructure Platform
Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning.

OpenRouter
Infrastructure Platform
OpenRouter pricing works differently from everything else in this index.

Replicate
Infrastructure Platform
Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on.
How much does Groq cost?
Is Groq more expensive than other providers?
Does Groq have a free tier?
Does Groq sell dedicated capacity?
























