

Together AI
Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning.
Updated on:
Together AI pricing: three ways to buy the same model
Together AI pricing charges three meters for the same model: per token on serverless, per GPU-minute on dedicated endpoints, and per token trained on fine-tuning. A fourth option, provisioned throughput, bills $0.05 per unit per minute on a monthly commitment. Nothing is bundled, so what you pay depends less on the model you pick than on the meter you're standing on.
Key takeaways
Serverless runs $0.14 to $3.00 per million input tokens, and several models price cached input separately at 42% to 81% off.
Dedicated inference bills per GPU-hour by the minute, per ready replica, and drops to zero when a deployment scales to zero.
Fine-tuning bills every token processed across training and validation, from $0.34 to $40.00 per million on a LoRA supervised run, and every model carries its own minimum charge.
There's no free tier. Together is fully prepaid, requires a $5 minimum credit purchase, and suspends API access when the balance hits zero.
Together AI pricing in 2026
Mode | Unit | Rate | Notes
|
|---|---|---|---|
Serverless | Per 1M tokens | Kimi K3 $3.00 in / $15.00 out; DeepSeek V4 Flash $0.14 / $0.28 | Cached input priced separately on most models |
Batch | Per 1M tokens | Up to 50% off serverless | Discount applies to selected models only |
Dedicated (DMI) | Per GPU-hour, billed per minute | H100 80GB $3.99 promotional to 30 Sep 2026, $5.49 list; B200 180GB $8.99; H200, B300, GB300 Custom | Per ready replica |
Provisioned throughput | Per PTU-minute | $0.05 | One month minimum, sales only |
Fine-tuning | Per 1M tokens trained | LoRA supervised $0.34 to $40.00, DPO $0.84 to $100.00 | Minimum charge per model, $4.00 to $60.00 |
GPU clusters | Per GPU-hour | H100 $3.99 on-demand, $1.99 preemptible | Reserved terms cut to $3.19 at 91 to 180 days |
What Together AI actually meters
Together meters three distinct things and never mixes them into one bill line. On serverless it's the token, split into input, cached input and output, and cached input carries its own published rate on most models rather than a blanket multiplier. MiniMax M3 charges $0.30 input against $0.06 cached. GLM-5.2 charges $1.40 against $0.26, an 81% discount, while Qwen3.5-397B-A17B charges $0.60 against $0.35 and saves only 42%. The discount isn't uniform, so caching pays back very differently model to model.
On dedicated model inference the meter switches to hardware. Together bills per GPU-hour, measured by the minute, per replica, and only while a replica is ready to serve. Provisioning, cold starts and DEGRADED replicas don't bill. Token volume is irrelevant here: the model affects cost only through the GPU count it needs.
Fine-tuning meters tokens processed, defined as (n_epochs × training tokens) + (n_evals × validation tokens). Disable packing and it recalculates as dataset length multiplied by max_seq_length, which moves the number a long way. Cancelled jobs pay for completed steps only, and failed jobs get fully refunded.
How credits work
Together is fully prepaid, and credits are the only currency. You buy a balance, minimum $5, and every service draws from it: API calls, dedicated deployments, fine-tuning and evaluation jobs. There's no free trial and no signup grant. Credits don't expire, and Together commits to advance notice if that changes, which is more than most prepaid vendors say.
Auto-recharge tops the balance back to a target when it falls below a threshold you set, as a single transaction on your default payment method. It works only when that default is a card: setting a US bank account as default turns auto-recharge off automatically. One sharp edge, credits bought after an invoice is generated can't clear that invoice or any past due balance.
What happens when you hit the limit
The balance hitting zero is the limit, and Together suspends API access until you add credits. That's a hard stop rather than an overage charge, which is what prepaid should mean and frequently doesn't. Creating a dedicated endpoint or a fine-tuning job needs enough balance up front to cover the cost.
Rate limits behave differently. They're dynamic, applied per model rather than per account, and grow with sustained reliable traffic. The old Build Tier 1 to 5, Scale and Enterprise labels are retired, so there's no tier to buy into. Every serverless response returns headers carrying the current limit.
How Together AI pricing has changed
Date | Milestone | Source
|
|---|---|---|
10 Sep 2026 | Preemptible GPU cluster compute enters public preview at a flat discount to on-demand, metered every one to two minutes | Vendor |
1 Sep 2026 | H100 80GB dedicated endpoint hardware drops to $3.99 per hour, down from $5.49 | Vendor |
25 Jun 2026 | Seedance 2.0 adds a 4K tier at $0.836 per second, against $0.40 at 1080p | Vendor |
9 Jun 2026 | Cached input pricing added for GLM-5.1 and Qwen3.5-397B-A17B. DeepSeek V4 Pro cut from $2.10 to $1.74 input and $4.40 to $3.48 output | Vendor |
29 May 2026 | Three models raised. Llama 3.3 70B goes $0.88 to $1.04, Qwen3.5 9B goes $0.10 to $0.17 input | Vendor |
10 Mar 2026 | First cached input rate published: MiniMax M2.5 at $0.06 per 1M, 80% off standard input | Vendor |
Customer
Sentiment Highlights
"They delivered a 2x reduction in latency and cut our costs by approximately a third"
Customer quoted on together.ai's own customer stories page (vendor-published)
Explore other providers

Synthesia
Image / Video Generation
Synthesia pricing sells a credit balance, and generated video draws 2 credits for every second, so one finished minute costs 120 credits.

Replit
Developer Tool
Replit pricing gives you a subscription that converts into a dollar allowance, then spends that allowance on agent work priced by effort.

Suno
AI Voice
Suno pricing sells credits, and one generation costs 10 of them and returns two songs, which works out at 5 credits a song.
How much does Together AI cost?
Does Together AI have a free tier?
How does Together AI bill dedicated endpoints?
Do Together AI credits expire?
























