Skip to main content
This page shows pricing information to help you understand API costs.For billing setup, payment methods, and usage monitoring, visit the Admin section. For rate limits, see the Rate Limits & Usage Tiers page.

Estimate your cost

Router API Pricing

Router API usage is billed per token at each model’s published rates — there are no per-request fees. Every model has its own input and output rate, with cache reads billed at the model’s discounted cache-read rate and reasoning tokens billed at the output rate. You are always billed at the requested model’s rates, regardless of how the request is served. See Router Models & Pricing for the full per-model rate card, including cache-write rates and long-context pricing.

Agent API Pricing

The Agent API provides access to third-party models from providers including OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, and NVIDIA with transparent, token-based pricing at each model’s published rates.

Model Pricing

Agent API pricing varies by provider and model, with each provider offering multiple models at different price points.

View Complete Third-Party Model Pricing

See the full pricing breakdown for all available models from OpenAI, Anthropic, Google, xAI, Z.AI, Moonshot AI, and NVIDIA, including cache rates and provider documentation links on the Agent API Models page.

Tool Pricing

When using tools with the Agent API:
Most tool costs are per invocation. sandbox is billed per container session — a 20-minute billing window per container, not a runtime cap — plus per SDK search query made from inside it. Tool costs are separate from model token costs.

Search API Pricing

Billing unit: Search API charges for each successful POST /search request, not for each query in the request. A successful request containing an array of up to five queries is one billing unit. Invalid requests, rate-limited requests, and upstream failures are not billed. A successful response is billed even when it returns no results. There are no additional token-based charges.

Embeddings API Pricing

Generate high-quality text embeddings for semantic search, retrieval-augmented generation (RAG), and other machine learning applications.

Standard Embeddings

Contextualized Embeddings

View Embeddings API Documentation

Learn how to use the Embeddings API for semantic search, RAG, and more.

Decisions API Pricing

The Decisions API answers yes/no, multiple-choice, and scored questions about text and images with calibrated probabilities.
Billing unit: you pay for input tokens only. The state, every question, and any images count as input; the response reports the exact number as usage.input_tokens, so you can compute the cost of a request from the response you already receive. A request with 600 input tokens costs $0.000024.

View Decisions API Documentation

Learn how to ask structured questions and read the probabilities the Decisions API returns.

Input Tokens

The number of tokens in your prompt or message to the API. This includes:
  • Your question or instruction
  • Any context or examples you provide
  • System messages and formatting
Example: “What is the weather in New York?” = ~8 input tokens

Output Tokens

The number of tokens in the API’s response. This includes:
  • The generated answer or content
  • Any explanations or additional context
  • Search results and references
Example: “The weather in New York is currently sunny with a temperature of 72°F.” = ~15 output tokens

Search Context Size vs Context Window

Search context size is not the same as the context window.
  • Search context size: How much web information is retrieved during search
  • Context window: Maximum tokens the model can process in one request (affects token limits)
Token Calculation: 1 token ≈ 4 characters in English text. The exact count may vary based on language and content complexity.

Cost Examples

Agent API Web Search

openai/gpt-5.6-terra • 500 input + 200 output tokens • 1 web search

Agent API Research Preset

low preset representative run • 2,000 input + 1,000 output tokens • 1 web search + 1 fetch
Actual preset costs vary with the selected model, token usage, and tool invocations. When present on a completed response, usage.cost.total_cost reports the calculated request cost.

Purchase Options

Perplexity API Platform on AWS Marketplace

Purchase API credits through AWS Marketplace with consolidated billing and enterprise procurement.

Contact Sales Team

Fill out our enterprise inquiry form to discuss custom pricing, dedicated support, and enterprise features for teams and organizations.