---
name: perplexity
description: Use when building AI agents that need real-time web search,
  multi-provider model access, structured reasoning, or tool orchestration.
  Reach for this skill when a user asks to integrate web-grounded AI, build
  agentic workflows, access multiple LLM providers through one API, or implement
  search-powered applications.
metadata:
  mintlify-proj: perplexity
  version: "1.0"
---

# Perplexity API Skill

## Product summary

Perplexity API is a multi-provider LLM platform with built-in web search, tool orchestration, and real-time data access. Use it to build agents that answer questions with live citations, run code in a sandbox, call custom functions, or access models from OpenAI, Anthropic, Google, and xAI through one unified interface. The primary endpoint is `POST https://api.perplexity.ai/v1/agent`. Key files: set `PERPLEXITY_API_KEY` environment variable. SDKs available for Python (`perplexityai`) and TypeScript (`@perplexity-ai/perplexity_ai`). See [Agent API Quickstart](https://docs.perplexity.ai/docs/agent-api/quickstart) for full documentation.

## When to use

Reach for this skill when:
- A user wants to build an AI agent that answers questions with current web data and inline citations
- You need to access multiple LLM providers (OpenAI, Anthropic, Google, xAI) through one API without managing separate keys
- A task requires running code in an isolated sandbox, searching the web, fetching URLs, or calling custom business logic
- You're building a research assistant, financial analysis tool, talent sourcer, or any agentic workflow that combines reasoning with tool use
- A user asks about configuring models, presets, tools, reasoning effort, or multi-turn conversations
- You need to enforce structured JSON output or control response format
- The application requires real-time financial data, people search, or domain-filtered web results

## Quick reference

### Authentication
```bash
export PERPLEXITY_API_KEY="your_key_here"
```

### Core request structure
```python
from perplexity import Perplexity

client = Perplexity()

response = client.responses.create(
    preset="low",  # or model="openai/gpt-5.6-sol"
    input="Your question here",
    tools=[{"type": "web_search"}],
    max_steps=5,
    instructions="System prompt rules here"
)

print(response.output_text)
```

### Presets (trade depth vs latency)

| Preset | Best for | Max steps | Use when |
|--------|----------|-----------|----------|
| `fast` | Single facts, quick lookups | 1 | Speed matters, minimal research needed |
| `low` | Everyday research, light multi-step | 3-5 | Current info + light tool use |
| `medium` | Multi-hop browsing, wide aggregation | 8-10 | Chaining evidence across sources |
| `high` | Expert reasoning, exhaustive coverage | 12-15 | Broadest coverage, completeness matters |
| `xhigh` | Open-ended agentic work, code execution | 20+ | Long tool loops, sandbox code, complex orchestration |

### Available models (provider/model format)

**OpenAI:** `openai/gpt-5.6-sol` (reasoning), `openai/gpt-5.6-terra` (general), `openai/gpt-5.6-luna` (fast)  
**Anthropic:** `anthropic/claude-opus-4-6` (best), `anthropic/claude-sonnet-4-6` (balanced), `anthropic/claude-haiku-4-5` (fast)  
**Google:** `google/gemini-3.1-pro-preview` (long-context), `google/gemini-3.1-flash-lite` (speed)  
**xAI:** `xai/grok-4.5` (conversational)  
**Perplexity:** `perplexity/sonar` (search-grounded)

### Built-in tools

| Tool | Type | Use for |
|------|------|---------|
| `web_search` | Search live web with filters | Current events, facts, research |
| `fetch_url` | Extract content from specific URLs | Deep dives into single sources |
| `sandbox` | Run Python code in isolated container | Computation, data processing, file generation |
| `people_search` | Find professionals by role/company/skill | Recruiting, sourcing, org mapping |
| `finance_search` | Structured financial & market data | 10-K filings, stock data, earnings |
| `image_search` | Find images with domain/format filters | Visual research, asset discovery |

### Response structure
```python
response.output_text  # Convenience: all text content aggregated
response.output       # Full array: messages, search_results, function_calls, etc.
response.usage        # Token counts and costs
response.status       # "completed", "incomplete", or error state
```

## Decision guidance

### When to use preset vs explicit model

| Scenario | Use preset | Use explicit model |
|----------|-----------|-------------------|
| Want automatic improvements as Perplexity optimizes | ✓ | |
| Need exact reproducibility, no future changes | | ✓ |
| Prototyping or exploring capabilities | ✓ | |
| Production with pinned behavior | | ✓ |
| Cost/latency sensitive, want Perplexity's tuning | ✓ | |
| Specific model required by spec | | ✓ |

### When to use web_search vs fetch_url

| Scenario | web_search | fetch_url |
|----------|-----------|-----------|
| Find relevant sources on a topic | ✓ | |
| Deep dive into one known URL | | ✓ |
| Current events, trending topics | ✓ | |
| Extract full content from specific page | | ✓ |
| Multi-source aggregation | ✓ | |
| Bypass search, go straight to source | | ✓ |

### When to use custom functions vs built-in tools

| Scenario | Custom function | Built-in tool |
|----------|-----------------|---------------|
| Call your internal API or database | ✓ | |
| Search the live web | | ✓ |
| Run arbitrary Python code | | ✓ (sandbox) |
| Integrate with your business logic | ✓ | |
| No external dependencies needed | | ✓ |

### When to use instructions vs input

| Content | instructions | input |
|---------|--------------|-------|
| Role, tone, citation rules | ✓ | |
| Do/don't constraints (apply every turn) | ✓ | |
| This specific question or task | | ✓ |
| Output format rules | ✓ | |
| Search query or research focus | | ✓ |

## Workflow

### 1. Start with a preset or model
Choose a preset for automatic optimization or specify a model for reproducibility. Presets bundle model, tools, system prompt, and step budget.

```python
response = client.responses.create(
    preset="low",  # Start here for most tasks
    input="Your question"
)
```

### 2. Add tools if needed
Enable built-in tools (web_search, sandbox, fetch_url, etc.) or declare custom functions. The model decides when to call them.

```python
response = client.responses.create(
    preset="low",
    input="Find and analyze recent AI funding trends",
    tools=[
        {"type": "web_search"},
        {"type": "fetch_url"}
    ],
    max_steps=8  # Allow multiple search rounds
)
```

### 3. Set system rules in instructions
Define role, tone, citation style, and constraints that apply on every turn. Keep it lean—every token re-processes on each step.

```python
response = client.responses.create(
    preset="low",
    input="...",
    instructions="You are a financial analyst. Cite every claim by source domain. Never speculate beyond retrieved evidence."
)
```

### 4. Handle tool results
For built-in tools, results attach to `response.output` as `search_results`, `sandbox_results`, etc. For custom functions, the run pauses and returns a `function_call` item—you execute it and send the result back.

```python
# Built-in tool results are in response.output
for item in response.output:
    if item.type == "search_results":
        for result in item.results:
            print(f"{result.title}: {result.url}")

# Custom function: run it yourself
function_call = next(i for i in response.output if i.type == "function_call")
result = my_function(**json.loads(function_call.arguments))

# Continue the run with the result
followup = client.responses.create(
    input=[
        {"role": "user", "content": "original question"},
        {"type": "function_call", "call_id": function_call.call_id, ...},
        {"type": "function_call_output", "call_id": function_call.call_id, "output": json.dumps(result)}
    ],
    tools=tools
)
```

### 5. Verify and iterate
Check `response.status`, token usage, and cost. For multi-turn conversations, use `previous_response_id` to continue or replay the conversation manually for precise control.

```python
if response.status == "completed":
    print(f"Cost: ${response.usage.cost.total_cost}")
    print(f"Tokens: {response.usage.total_tokens}")
else:
    print(f"Error: {response.error}")

# Continue conversation
followup = client.responses.create(
    input="Follow-up question",
    previous_response_id=response.id  # Automatic context carry
)
```

## Common gotchas

- **Anthropic models require `max_output_tokens`**: Always set it when using `anthropic/*` models, or the request fails with HTTP 400.
- **Presets replace instructions, not append**: Setting `instructions` with a preset overwrites the preset's system prompt. Omit `instructions` to keep the preset's tuned prompt.
- **`max_steps: 1` prevents tool loops**: At step 1, direct tools may run but the agent cannot reason over results. Tools like `finance_search` need `max_steps: 3+` to initialize and run.
- **Custom functions don't execute server-side**: The API pauses and returns a `function_call` item. You must run the function and send the result back with the same `call_id`.
- **Search results are not citations**: `web_search` returns results; the model generates citations inline. Pull URLs from `search_results` items, not from the text.
- **Prompt tokens re-process on every step**: Keep `instructions` lean. Use request parameters (tool filters, `response_format`) instead of prose rules—they're enforced directly.
- **`previous_response_id` only works for completed responses**: If a response errored or is incomplete, you must replay the conversation manually in the `input` array.
- **Structured output (JSON schema) has first-request delay**: New schemas take 10–30 seconds to prepare on the first call. Subsequent requests are fast.
- **Avoid URLs in structured JSON output**: Models can hallucinate or malform URLs inside JSON. Extract URLs from `search_results` or `fetch_url` results instead.
- **Rate limits use leaky-bucket algorithm**: Short bursts are allowed; sustained high throughput requires tier upgrade or custom limit request.

## Verification checklist

Before submitting work with Perplexity API:

- [ ] API key is set as `PERPLEXITY_API_KEY` environment variable or passed explicitly
- [ ] If using Anthropic models, `max_output_tokens` is set
- [ ] `max_steps` is high enough for tool loops (≥3 for `finance_search`, ≥8 for multi-source research)
- [ ] `instructions` is lean and role/tone focused, not a list of do/don'ts
- [ ] Tool results are read from `response.output` array, not assumed to be in text
- [ ] Custom function calls are handled: run the function, return result with matching `call_id`
- [ ] Multi-turn conversations use `previous_response_id` or manually replay conversation
- [ ] Response status is checked: `response.status == "completed"` before using output
- [ ] Token usage and cost are logged for monitoring
- [ ] Error handling catches `APIStatusError`, `RateLimitError`, `AuthenticationError`
- [ ] Structured output schemas are tested (first request may timeout)

## Resources

Comprehensive page-by-page navigation: [https://docs.perplexity.ai/llms.txt](https://docs.perplexity.ai/llms.txt)

Critical documentation:
- [Agent API Quickstart](https://docs.perplexity.ai/docs/agent-api/quickstart) — basic usage, presets, tools
- [Building Agents Guide](https://docs.perplexity.ai/docs/agent-api/building-agents/define-the-run) — define runs, prompts, tools, output control
- [Models & Pricing](https://docs.perplexity.ai/docs/agent-api/models) — available models, token costs, service tiers

---

> For additional documentation and navigation, see: https://docs.perplexity.ai/llms.txt