Skip to main content
Perplexity

Search documentation

Type to search this documentation.

On this pageOverview

Pricing

Router API usage is billed per token at each model's published rates — there are no per-request fees. Every model has its own input and output rate, with cache reads billed at the model's discounted cache-read rate and reasoning tokens billed at the output rate. You are always billed at the requested model's rates, regardless of how the request is served.

See Router Models & Pricing for the full per-model rate card, including cache-write rates and long-context pricing.

The Agent API provides access to third-party models from providers including OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, and NVIDIA with transparent, token-based pricing at each model's published rates.

Agent API pricing varies by provider and model, with each provider offering multiple models at different price points.

View Complete Third-Party Model Pricing

See the full pricing breakdown for all available models from OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, and NVIDIA, including cache rates and provider documentation links on the Agent API Models page.

When using tools with the Agent API:

Tool Price Description
web_search (standard) $0.0025 per invocation Standard web search with search_type: "web". $2.50 per 1,000 invocations
web_search (Fast Search) $0.001 per invocation The same tool with search_type: "fast". $1.00 per 1,000 invocations. See Fast Search
fetch_url $0.0005 per invocation Fetches and extracts content from specific URLs
people_search $0.005 per invocation Looks up professionals, employees, and people. $5 per 1,000 tool invocations
finance_search $0.005 per invocation Retrieves financial data and market information. $5 per 1,000 tool invocations
sandbox $0.03 per session Isolated container for executing code during an Agent API request. A session covers up to 20 minutes of active use for billing purposes — this is the billing window, not a runtime cap. SDK search queries made from inside the sandbox are billed at $0.0025 per request (same as web_search).
API Price per 1K requests Description
Search API $5.00 Raw web search results with advanced filtering
Search API with Fast Search $1.00 Lower-latency web search results. Set search_type: "fast"

Generate high-quality text embeddings for semantic search, retrieval-augmented generation (RAG), and other machine learning applications.

Model Dimensions Price ($/1M tokens)
pplx-embed-v1-0.6b 1024 $0.004
pplx-embed-v1-4b 2560 $0.03
Model Dimensions Price ($/1M tokens)
pplx-embed-context-v1-0.6b 1024 $0.008
pplx-embed-context-v1-4b 2560 $0.05

View Embeddings API Documentation

Learn how to use the Embeddings API for semantic search, RAG, and more.

Token and Cost Glossary

Input Tokens

The number of tokens in your prompt or message to the API. This includes:

  • Your question or instruction
  • Any context or examples you provide
  • System messages and formatting

Example: "What is the weather in New York?" = ~8 input tokens

Output Tokens

The number of tokens in the API's response. This includes:

  • The generated answer or content
  • Any explanations or additional context
  • Search results and references

Example: "The weather in New York is currently sunny with a temperature of 72°F." = ~15 output tokens

Search Context Size vs Context Window

Search context size is not the same as the context window.

  • Search context size: How much web information is retrieved during search
  • Context window: Maximum tokens the model can process in one request (affects token limits)

Agent API Web Search

openai/gpt-5.6-terra • 500 input + 200 output tokens • 1 web search

ComponentCost
Input tokens$0.001
Output tokens$0.0024
web_search$0.0025
Total$0.0059

Agent API Research Preset

low preset representative run • 2,000 input + 1,000 output tokens • 1 web search + 1 fetch

ComponentCost
Model input tokens$0.001
Model output tokens$0.003
web_search$0.0025
fetch_url$0.0005
Total$0.007

Actual preset costs vary with the selected model, token usage, and tool invocations. When present on a completed response, usage.cost.total_cost reports the calculated request cost.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu