

Perplexity API
Perplexity API pricing charges you twice for the same call.
Updated on:
Perplexity Sonar API pricing: two meters running on every call
Perplexity API pricing charges you twice for the same call. You pay per million tokens like any model API, and you pay a separate request fee for the web search that grounds the answer. That fee isn't flat: it climbs with the search_context_size you ask for. On a short Sonar query the search fee is the bill and the tokens round to nothing. Perplexity documents all of it.
Key takeaways
Sonar bills $1 per million tokens in and out, then adds $5, $8 or $12 per 1,000 requests depending on search context size.
Perplexity's own Sonar worked example totals $0.005420 for one query, and $0.005 of that is the request fee rather than tokens.
Sonar Deep Research carries no request fee. It bills citation tokens at $2 per million, reasoning tokens at $3 per million, and search queries at $5 per 1,000.
Sonar Chat Completions is deprecated. Perplexity supports it until 27 September 2026 and points new builds at the Agent API, which prices tools per invocation.
Perplexity API pricing in 2026
Model | Input | Output
|
|---|---|---|
Sonar | $1 | $1 |
Sonar Pro | $3 | $15 |
Sonar Reasoning Pro | $2 | $8 |
Sonar Deep Research | $2 | $8 |
Model | Low context | Medium context |
--- | --- | --- |
Sonar | $5 | $8 |
Sonar Pro | $6 | $10 |
Sonar Reasoning Pro | $6 | $10 |
Sonar Pro with Pro Search (search_type: pro) | $14 | $18 |
Source: docs.perplexity.ai/docs/getting-started/pricing, read 22 Sep 2026.
What Perplexity actually meters
Perplexity meters tokens and requests separately, and the docs state the formula outright: total cost per query equals token costs plus a request fee. That fee applies to Sonar, Sonar Pro and Sonar Reasoning Pro, and not to Sonar Deep Research.
The fee tracks search_context_size, set inside web_search_options. Low is the default and the cheapest, high buys maximum search depth. Moving Sonar from low to high more than doubles the fee, from $5 to $12 per 1,000, without touching the token rate.
That split matters. Perplexity's published Sonar example uses 9 input and 411 output tokens at low context: tokens cost $0.000009 and $0.000411, the request fee costs $0.005. So 92% of that line is search, not the model. Budget from token counts alone and you'll be out by an order of magnitude.
Sonar Deep Research inverts the shape, dropping the request fee and adding three meters instead. Perplexity's worked example for one call totals $0.816123, of which $0.581841 is reasoning tokens. Every response carries a usage.cost object splitting request_cost from token costs, so you reconcile per call.
How credits work
Perplexity runs prepaid credits, not postpaid invoicing. You buy credits in the API console, Stripe takes the payment, and calls draw the balance down. Auto reload adds credits when the balance falls below a threshold you set, and the docs recommend enabling it.
Cumulative purchases also drive your usage tier, which sets your rate limits. Tier 0 starts at $0 and Tier 5 at $5,000 of lifetime purchases, with $50, $250, $500 and $1,000 in between. Tiers count lifetime spend rather than current balance and never downgrade: on the Agent API that's 1 query per second at Tier 0 against 33 at Tier 5, and on Sonar 50 requests a minute against 4,000.
Credit expiry, rollover, pooling across projects and deduction order are all not documented, and neither is any refund path for unused balance. Enterprises can buy credits through AWS Marketplace.
What happens when you hit the limit
You get blocked, not billed. Perplexity's docs are blunt about it: run out of credits and your API keys are blocked until you top the balance up. Calls then return a 401, which the FAQ lists alongside invalid and deleted keys as a credential failure rather than a billing one. There's no grace allowance and no arrears billing, which makes auto reload worth checking before any launch.
Rate limiting behaves differently. Exceed your tier's QPS or per-minute limit and you get a 429 with a Retry-After header, and Router requests rejected with a 429 aren't billed.
How Perplexity's pricing has changed across all these years
Date | Milestone | Source
|
|---|---|---|
27 Sep 2026 | Sonar support ends | docs.perplexity.ai/docs/resources/changelog |
Jul 2026 | Sonar folded into the Agent API. The new Router API bills per token with no request fee | docs.perplexity.ai/docs/resources/changelog |
Apr 2026 | API credits become purchasable through AWS Marketplace | docs.perplexity.ai/docs/resources/changelog |
Nov 2025 | Pro Search reaches GA on Sonar Pro, carrying a higher request-fee band | docs.perplexity.ai/docs/resources/changelog |
Jul 2025 | Responses return a cost object with request_cost split from token costs | docs.perplexity.ai/docs/resources/changelog |
Mar 2025 | Search context modes introduced and citation tokens stop billing except on Deep Research. Default from 18 April 2025 | docs.perplexity.ai/docs/resources/changelog |
Jan 2025 | Sonar and Sonar Pro launch, replacing the llama-3.1-sonar family | docs.perplexity.ai/docs/resources/changelog |
Source: docs.perplexity.ai/docs/resources/changelog, read 22 September 2026. Wayback's CDX index
Customer
Sentiment Highlights
Explore other providers

OpenRouter
Infrastructure Platform
OpenRouter pricing works differently from everything else in this index.

ChatGPT
Chatbot
ChatGPT pricing sells seats from $0 to $100 a month and meters the ambitious work on top in credits.

Microsoft 365 Copilot
Workspace Platform
Microsoft 365 Copilot pricing runs two meters at once.
How much does the Perplexity Sonar API cost?
What is the Perplexity request fee?
Does the Perplexity API have a free tier?
What happens if my Perplexity credits run out?
























