

Replicate
Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on.
Updated on:
Replicate pricing: paying by the second, until the model is official
Replicate pricing charges per second of compute, and the per-second rate depends on which GPU your model runs on. An A100 costs $0.001400 a second, an H100 $0.001525. But a subset of models, the ones Replicate calls official, ignore the clock entirely and bill per output image, per token or per second of generated video. Which meter applies isn't a setting you choose. It's a property of the model you picked.
Key takeaways
Replicate's single-GPU rates run from $0.000225 a second on a T4 to $0.001525 a second on an H100, with an 8x H100 node at $0.012200, all published on the pricing page with no login.
Official models bill per output unit instead of per second: $0.04 per image on FLUX 1.1 Pro, $0.09 per second of output video on Wan 2.1 i2v 480p.
On public models you pay only for active processing time. Setup and idle time are free, so cold boots cost you latency but not money.
On private models and deployments you pay for setup, idle and active time, because the hardware is dedicated to you.
Replicate pricing in 2026
Hardware | Per second | Per hour | Spec
|
|---|---|---|---|
CPU (small) | $0.000025 | $0.09 | 1x CPU, 2GB RAM |
Nvidia T4 | $0.000225 | $0.81 | 16GB GPU RAM |
Nvidia L40S | $0.000975 | $3.51 | 48GB GPU RAM |
Nvidia A100 80GB | $0.001400 | $5.04 | 80GB GPU RAM |
Nvidia H100 | $0.001525 | $5.49 | 80GB GPU RAM |
Nvidia H200 | $0.001525 | $5.49 | Committed spend contract only |
8x Nvidia H100 | $0.012200 | $43.92 | Committed spend contract only |
What Replicate actually meters
Replicate runs two meters, and the model decides which one you get.
Most community and private models meter wall-clock seconds of compute at the hardware's published rate. Every run creates a prediction, and Replicate charges the seconds that prediction spends actually executing.
Official models replace that with output units. FLUX 1.1 Pro charges $0.04 per output image, FLUX Schnell $3.00 per thousand output images, DeepSeek R1 $3.75 per million input tokens and $0.01 per thousand output tokens, and Wan 2.1 i2v 480p $0.09 per second of output video. Replicate named this category on 29 January 2025 and described the switch plainly: instead of being charged for the time a model runs, you're charged by output.
Two wrinkles matter. Models that call other models bill you for the root model's compute plus every downstream model it invokes, and the downstream list appears in the model's pricing section. And when Nano Banana Pro falls back to Seedream 5.0 lite under rate limiting, you pay the fallback model's price, not Nano Banana Pro's.
How credits work
Since 16 July 2025, every new Replicate account bills through prepaid credit. You buy a balance, and usage deducts from it.
Purchased credit is valid for one year from the purchase date and isn't refundable. Replicate documents no rollover concept, because the balance is simply money you've already spent.
Auto reload tops the balance back up when it drops past a threshold you set. The minimum threshold is $5 and the minimum reload balance is $15, and Replicate's own example is a $10 threshold with a $50 reload, which adds $40 when you hit $10. If your balance already sits at or below the threshold when you save, the reload fires immediately.
Accounts created before 16 July 2025 can stay on monthly arrears billing, where Replicate charges the previous month's usage at the start of the next one. Replicate says it intends to migrate most accounts to prepaid eventually. Credit is account-scoped, and organizations get a shared balance rather than pooling individual ones.
What happens when you hit the limit
Replicate throttles you before it stops you. As your credit balance approaches zero, it applies progressively stronger rate limits so you have time to top up rather than getting cut off without warning. The docs recommend keeping the balance above $20 via auto reload.
At zero, Replicate prevents new work from starting and shuts down any infrastructure running for you. A prediction can occasionally run past the balance, and Replicate charges the outstanding amount to your default payment method at month end.
Normal rate limits sit at 600 prediction creations a minute and 3,000 requests a minute elsewhere, with short bursts allowed. Accounts holding granted credit with no payment method on file get 1 request per second and 6 a minute. Predictions time out at 30 minutes.
How Replicate's pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
2 Mar 2026 | Nano Banana Pro fallback bills at the fallback model's rate | Vendor |
21 Nov 2025 | Approximate per-run cost shown on prediction and training pages | Vendor |
8 Oct 2025 | Invoice PDFs available from billing settings via Stripe | Vendor |
26 Sep 2025 | Low-balance throttling documented as a spend guardrail | Vendor |
16 Jul 2025 | All new accounts moved to prepaid credit. Existing accounts keep monthly arrears | Vendor |
29 Jan 2025 | Official models named and switched from time-based to per-output billing | Vendor |
22 Nov 2024 | Per-video pricing support added, alongside L40S GPUs and preview hardware pricing | Vendor |
Customer
Sentiment Highlights
"On replicate.com a single image takes 1.5s at a price of 1000 images per $1."
aleyan, Hacker News, December 2025
"Neither fal nor replicate return accurate pricing in the response body."
seblavoie, Hacker News, February 2026
Explore other providers

Perplexity API
API
Perplexity API pricing charges you twice for the same call.

Pinecone
Data Platform
Pinecone pricing meters four things on a serverless index: read units, write units, gigabytes stored and gigabytes returned.

OpenRouter
Infrastructure Platform
OpenRouter pricing works differently from everything else in this index.
How much does Replicate cost per hour?
Does Replicate charge for cold boots?
Do Replicate credits expire?
What happens if you run out of credit on Replicate?
























