Skip to main content
The Decisions API is billed at $0.04 per million input tokens. Output tokens are free. See Pricing.
The Decisions API answers questions with probabilities. You send the content as state, which can be text, JSON, or images, attach one or more named questions, and get one answer per question:

What a decision model is

A decision model is a class of model built to make fast, structured decisions that software can use directly. It reads natural-language text and images the way a multimodal language model does, but instead of writing text it returns typed answers with probabilities: yes or no, one of your options, or a level on your rubric. It does not write replies, generate code, or explain its reasoning; your code does the reasoning with the numbers it returns. decider-27b is a decision model, and the Decisions API is how you call it. Use it where you would otherwise ask a chat model for a label and parse the reply: classifying, routing, grading against a rubric, or any decision you want to threshold. You get numbers you can compare against a cutoff, no output parsing, and as many questions as you need about the same content in one request.

Quickstart

1

Set your API key

Any Perplexity API key works; create one at console.perplexity.ai if you need one.
2

Install an HTTP client

The Decisions API is a single JSON endpoint, so any HTTP client works. The Python example uses httpx. The TypeScript example uses the fetch built into Node.js 18 and later and has no dependencies; save it as decisions.mts and run npx tsx decisions.mts (npx downloads tsx on first use), or save it as decisions.mjs and run node decisions.mjs.
3

Ask three questions about one review

Send a POST to https://api.perplexity.ai/v1/decisions with the content as state, the model decider-27b, and your questions.
4

Read the answers

The response has one answer per question, under the name you gave it. This is the response the request above returned, with the values as the API sent them:
Read it like this:
  • defect: a 94% probability that the review reports a defect. Compare noul against a threshold you choose; values near 0.5 mean the model is unsure.
  • sentiment: mixed is the option with the highest probability (95%). probabilities covers every option you defined and sums to about 1, so you can see how close the runner-up came.
  • severity: score is the probability-weighted average of the level indices, so 1.78 sits between Inconvenient (1) and Product unusable (2), closer to 2. legend maps each index back to your rubric, and probabilities shows the full distribution over levels.
confidence on choice and score answers is the model’s own certainty estimate, from 0 to 1. It is not the top probability: in the example above, sentiment has a top probability of 0.95 and a confidence of 0.93. It drops when the runner-up is close.Identical requests usually return identical numbers. Occasionally they differ in the second decimal place, so set thresholds with some margin. model echoes the model name you sent.

Question types

Every question in a request refers to the same state, which can be a string, an object, or an array. You name each question, and the response uses the same names. Each question has a type, instructions (what to decide), and, depending on the type, criteria.

noul: yes or no

Ask a question or state something to check. Give instructions, criteria, or both; criteria defines what counts as yes and what counts as no. A noul with neither returns 400.
The answer is noul, the probability of yes or true, from 0 to 1.

choice: one of your options

criteria maps each option name to a description of when it applies. Use null as the description to let the name speak for itself. A question accepts 1 to 255 options.
The answer has choice (the option with the highest probability), probabilities (one value per option, summing to about 1), and confidence.

score: a level on an ordered rubric

criteria is an ordered array of level descriptions. The index in the array is the level’s score, starting at 0. A question accepts up to 10 levels. Use at least two: with a single level there is nothing to decide, so the answer is always a score of 0 with probability 1.
The answer has score (the probability-weighted average of the level indices, which can fall between two levels), legend (each index, as a string, mapped back to your rubric entry), probabilities (one value per level, keyed like legend), and confidence.

Images in state

state can carry images next to text. Pass state as an array and put each image in an OpenAI-style image part with a base64 data URL:
  • PNG, JPEG, and WebP data URLs are accepted. The API never fetches a URL: an http or https image URL returns 400.
  • An image can also be the whole state, with no text.
  • The API reads images in 32 × 32 pixel tiles. Keep each image at or under 2,048 tiles: round the width and the height to the nearest multiple of 32 and keep (width / 32) × (height / 32) at or under 2,048. 1440 × 1440 and 2048 × 1024 fit; 1600 × 1310 does not. A larger image does not return 400: the request waits about a minute and then returns 504. Resize before you send.
  • Image tokens count toward usage.input_tokens and the input limit like text. In our tests an image cost about 1,000 input tokens per megapixel.

Request limits

state must be a string, an object, or an array; null returns 400. score levels follow the same rule. choice descriptions can also be null. An unknown top-level field returns 400.

Model

One model serves the Decisions API: decider-27b. Set it on every request. The response model field echoes the name you sent. Both names serve the same model today. A missing or unknown model returns 400:

Endpoint and authentication

Send the key as Authorization: Bearer <PERPLEXITY_API_KEY> and the body as JSON. A key in an x-api-key header is not read, so the request returns 401. Another method on the endpoint returns 405 with Allow: POST, and any other path, including a trailing slash, returns 404.

Request ids

Responses carry an x-request-id header with a UUID, on success and on most errors. Log it with your results and quote it in support requests. A 401, a 404, and a 504 carry no request id.

Rate limits

Every organization can send 10 requests per second to the Decisions API, on every plan. A token limit also applies to large bursts. Successful responses carry x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-used, and x-ratelimit-reset (Unix seconds). A request over a limit returns 429 with a Retry-After header in seconds. Wait that long, then retry.

Errors

Most errors return a JSON body with an error object. Read error.message for the reason and error.type for the category; don’t branch on error.code, which is a string, a number, or null depending on the error. A 404 or 405 has an empty body, and a 504 can return an HTML page, so check the status before you parse the body.

Timeouts

Response time grows with input size. In our tests on September 30, 2026, a request with a few hundred input tokens answered in under 2 seconds, about 90,000 tokens took 5 seconds, about 190,000 tokens took 14 seconds, and just under the input limit took 23 seconds. If the model does not answer in time, the request returns 504; in our tests that took about a minute. Set your client timeout to fit the input. The examples on this page use 30 seconds, which covers any request under the input limit; for small inputs, 10 seconds is plenty.

Pricing

The Decisions API costs $0.04 per million input tokens. Output tokens are free, and there is no per-request fee. Input tokens are the usage.input_tokens value in each response, so you can compute the cost of a request from the response you already receive. Usage is billed to the organization that owns the API key, like every other Perplexity API.

Next steps

Answer questions reference

Full request and response schema for POST /v1/decisions.

Cookbook: triage support tickets

Route twelve tickets with three questions per ticket, then send only the escalations to the Agent API.

Agent API

Web-grounded answers with citations, tools, and structured output.

Router API

Direct access to open-weight models through OpenAI- and Anthropic-compatible endpoints.
Need help? Check out our community for support.