Skip to main content
Ask a chat model “is this ticket urgent?” and you get a paragraph. Somewhere in it is a “yes”, a “probably”, or a “this may warrant attention”. Now write the if statement that reads that paragraph. You can’t. So you ask for JSON instead, parse it, retry when the JSON is malformed, and still have no idea how sure the model was. And you paid a frontier model to read every ticket, including the spam and the password resets, to get there. Triage is a pile of small decisions, and each one needs a number you can put a threshold on. That is what the Decisions API returns. You send one piece of content and a set of named questions. You get back a probability for every yes/no question, a probability for every option of every choice question, and a probability for every level of every score question. No prose to parse. Thresholds become one-line comparisons, and tuning them means changing a constant, not rewriting a prompt. A decision model is a class of model built to make fast, structured decisions that software can use directly. It reads text and images like a multimodal language model, but instead of writing text it returns typed answers with probabilities: yes or no, one of your options, or a level on your rubric. It does not write replies or explain its reasoning; your code does that part. decider-27b is a decision model, and the Decisions API is how you call it. This recipe sends text only; the same request accepts images in state when your tickets carry screenshots. This recipe triages twelve inbound support tickets for a fictional SaaS product. Each ticket gets one Decisions API request that answers three questions at once: does a person need to act, what is the customer asking for, and how severe is the impact. A small routing function turns those numbers into escalate, review, queue, auto_reply, or close. Only the escalated tickets go on to the investigation step, an Agent API request that drafts an investigation plan. In the run below, 3 of 12 tickets reached it. The other 9 never touched a frontier model.

What you need

  • Python 3.10 or newer on macOS or Linux. The shell commands use Unix syntax.
  • A Perplexity API key. Every key can call the Decisions API; there is no separate signup. The same key runs the optional Agent API step.
The Decisions API is one JSON endpoint, POST https://api.perplexity.ai/v1/decisions, so the script calls it with httpx, a plain HTTP client. The Perplexity Python library is used only for the Agent API step. One key, two APIs.

Set up

Everything lives in one directory. You create the first three files from this page; the script writes the fourth on every successful run.
Everything you need is on this page. Start with the dependencies, pinned to the versions this recipe was tested with.
requirements.txt
Create the directory and save requirements.txt in it. Then create a virtual environment, install the two packages, and export your key.

The tickets

Twelve tickets, one JSON object per line. They cover an outage, a leaked credential, accidental data loss, a partial SSO failure, two billing requests, a how-to question, a feature request, a password reset, a rate-limit question, an angry repeat complaint, and one piece of spam. Save the file as tickets.jsonl. Each line is sent whole as the request’s state. state accepts a string, an object, or an array, so you do not flatten the ticket into a prompt. The model sees plan, subject, and body as labeled fields. In production, put the full thread, prior tickets from the same account, and any account metadata in the same object. See the Decisions API docs for limits.
tickets.jsonl

The script, stage by stage

The rest of this page walks through triage_tickets.py one piece at a time. Every block below is an exact excerpt of the full file, which is at the end of the page.

Imports and constants

The script needs httpx for the Decisions API and the Perplexity Python library for the Agent API. The endpoint and the model live in one place. decider-27b is the model that serves the Decisions API; use decider-27b-v0 instead if you want to pin the current version. The thresholds are the whole routing policy. Tune them against your own labeled tickets and nothing else in the file changes.

Three questions, one request

The three questions cover the three question types, written as the JSON the API takes. A noul is a yes/no question and returns one probability. A choice returns a probability for every option plus the top option and a confidence. A score takes an ordered rubric, where the position in the list is the level, and returns the expected level, a probability for every level, and a confidence. The names you give the questions (needs_human, intent, severity) are the keys you read the answers from. confidence is the API’s own certainty estimate for the pick. It is not max(probabilities), and the two can differ by a lot: on T-1007 below the top option had 0.58 of the mass while confidence came back at 0.51. Confidence drops when the runner-up is close or when no option fits the content well. The expected level a score question returns is the probability-weighted average of the levels, so a ticket split evenly between level 2 and level 3 scores about 2.5. That is why the script also reads the probability on the top level separately: an average can look calm while the tail is not. Write criteria the way you would brief a new support hire. The option descriptions in intent and the level descriptions in severity do most of the work. Vague criteria give you flat probabilities.

Loading tickets and recording decisions

DecisionsError is the error the script raises when the Decisions API answers with anything but 200. load_tickets reads one JSON object per line. Decision is what the script records for each ticket: the route, a one-line reason, every probability the route used, and the request id from the x-request-id response header. You will want all of it later when you tune thresholds or ask why a ticket went where it did, and the request id is what support asks for.

From probabilities to a route

route is a pure function of five numbers and a label, taken as keyword arguments so a call site cannot silently swap two floats. Order matters. Danger first: if the probability mass on the top severity level is at or above ESCALATE_P_CRITICAL, or the expected severity clears ESCALATE_MIN_SCORE, the ticket escalates no matter what else the model said. Uncertainty second: if the model is not confident about the intent, a person decides what the ticket is. Only then do the automated exits run, and each one requires a confident, low-risk signal. Escalation checks the tail, not just the average. A ticket with severity probabilities {1: 0.5, 3: 0.5} has an expected score of 2.0, which looks moderate. It also has a coin-flip chance of being an outage. p_critical catches that; the expected score alone does not.

One request per ticket

classify posts one request per ticket: the model, the ticket as state, and all three questions. Anything but 200 raises DecisionsError with the status, the error body, and the request id, so a failure tells you what to fix and what to quote. The response has one entry per question under answers, keyed by the name you gave it. answers["needs_human"]["noul"] is the yes probability. answers["intent"] has choice, confidence, and probabilities. answers["severity"] has score, confidence, legend, and probabilities. The level keys in legend and probabilities are strings ("0" to "3"), because JSON object keys always are, so the script builds the top level’s key from the length of the rubric and reads the mass on it.

The investigation step

Only escalated tickets reach draft_escalation. It uses the Perplexity Python library to make one Agent API request with the low preset and asks for a short investigation plan: what to check first, what to tell the customer, who to page. This is the call you would not want to make for all twelve tickets, and the triage step is what keeps it to three. The plan text changes from run to run and may carry citation markers such as [web:1] where the Agent API used a web source.

The run loop

main reads the key from PERPLEXITY_API_KEY, opens one httpx.Client that sends the key as a bearer token on every request, and classifies every ticket in order. The client timeout is 30 seconds, far more than a ticket needs, so only a real problem trips it. It catches DecisionsError and httpx.HTTPError, the base class for timeouts and connection failures, so every failure prints the same one-line message and stops the run before any results are written. Then it prints a table and writes one JSON line per ticket. With --draft-escalations it builds one Perplexity client and sends the escalated tickets to the Agent API; a failure there prints a line to stderr and the loop continues, because the triage results are already on disk and worth keeping. The API allows 10 requests per second per organization. More than 10 requests in the same second get 429 with a Retry-After header. This loop is sequential, so it never gets close.

The full file

Save this as triage_tickets.py. It is the seven excerpts above in order, nothing else.
triage_tickets.py

Run it

From the directory with the three files, with the virtual environment active and the key exported:
Output observed on 2026-09-30 at 19:53 UTC against the Decisions API. The escalation plan below is from a --draft-escalations run a few minutes earlier, which printed the same table.
The twelve classification requests take about 2.5 seconds in total. In our tests most requests took 150 to 250 ms for 590 to 659 input tokens per ticket. The first request of a process is slower because it opens the connection, and an occasional request took several seconds. Plan for hundreds of milliseconds per ticket, not tens, and measure your own. Add --draft-escalations to run the Agent API step for the escalated tickets. That run took 17.2 seconds end to end, and 15.6 of those seconds were the three Agent API requests.
The P(crit) column is the probability on the top severity level, the one the escalation rule reads. triage_results.jsonl has one line per ticket with the route, the reason, every probability the route used at full precision, and the request id. Keep these files. They are your labeled dataset for tuning thresholds later.

Reading the numbers

A few rows in that table show why probabilities beat free text for this job. T-1007, the deleted workspace. Intent confidence was 0.51. The probabilities were spread across outage_or_bug (0.58), feature_request (0.20), and how_to (0.13), which is fair; none of the seven options describes “I deleted my own data”. A confidence-first policy would have parked this in review. But P(crit) was 0.85: whatever this ticket is, it is probably data loss. Danger-first ordering sends it to escalate. In our run the Agent API draft told the on-call engineer to check the audit logs and whether the reports can be recovered from a soft-delete window or a backup, and to tell the customer not to recreate the workspace while recovery is investigated. T-1011, the SSO loop. Intent confidence 0.55, with the mass split between account_access (0.61) and outage_or_bug (0.37). Expected severity 2.19 and P(crit) 0.24, under both escalation thresholds. This lands in review, which is right. Thirty of two hundred users cannot log in, and a person should decide in the next few minutes whether that is an incident. A free-text classifier would have picked one label and hidden the disagreement. T-1004, the dark mode request. This is the one ticket in the set that taught us something about thresholds. Across five runs on 2026-09-29, needs_human came back at 0.29, 0.29, 0.31, 0.31, and 0.32. On 2026-09-30 it was 0.31 on every run. The first version of this script used an AUTO_REPLY_MAX_HUMAN of 0.30, so the same ticket flipped between auto_reply and queue with nothing about it changing. A cutoff that sits on top of where the model settles turns small run-to-run movement into a different decision. That is why the constant is 0.35: clear of the cluster, so T-1004 auto-replies on every run. Whether a feature request should auto-reply at all is a team choice. If you want a person to see them, drop feature_request from AUTO_REPLY_INTENTS rather than pulling the threshold down onto the cluster. Replay triage_results.jsonl to see what else a change moves. T-1005, the spam. needs_human was 0.42, which looks odd for spam until you read the criteria: the model was asked whether “a canned answer” resolves it, and spam does not fit either description. The route does not use needs_human for spam; it closes on intent == "spam" with a severity under 0.5. Pick the signal that answers the question you are actually asking. Identical requests usually return identical numbers, but not always. On 2026-09-30, 11 of 13 full runs printed the same table byte for byte. One of the other two moved T-1009 and T-1012 by at most 0.06, and the other put T-1011’s intent confidence at 0.53 instead of 0.55. On 2026-09-29, needs_human on T-1004 ranged from 0.29 to 0.32. With the cutoffs clear of those ranges, every route held on every run. Set thresholds with margin, do not read the third decimal place, and check where your borderline tickets cluster before you trust a cutoff that sits on top of them.

Adapt it

  • Your tickets. Replace tickets.jsonl with an export from your help desk. Keep one JSON object per line with an id and a subject; other fields are up to you. Pass --tickets path/to/file.jsonl.
  • Your categories. Edit the criteria in QUESTIONS. Describe every option; undescribed options are interpreted by name alone.
  • Your thresholds. Label a few hundred tickets, run the script, and compare the routes against the labels. Adjust the constants. Escalation misses are expensive, so start with ESCALATE_P_CRITICAL low and raise it only when the on-call rota complains.
  • More context. Put the conversation history, account tier, and open incidents in the state object. The questions do not change.
  • Throughput. Write an async classify with httpx.AsyncClient and classify tickets concurrently. Stay under 10 requests per second per organization, and on 429 wait Retry-After seconds before you retry.

Troubleshooting

Every Decisions API error prints as <ticket>: Decisions API failed: <status> <error body> (x-request-id <id>). Read error.message in the body. A 504 returns an HTML page instead of JSON; retry it.
  • 400 with Invalid model means DECISIONS_MODEL is not decider-27b or decider-27b-v0.
  • 400 with any other message is a request problem, for example an empty question name, state set to null, a null score level, or more than 255 options. The message names it.
  • 401 is a missing or invalid key. The Decisions API reads Authorization: Bearer only, which httpx.Client sends for you here.
  • 429 means more than 10 requests landed in one second, or the requests were very large (a token limit also applies). A sequential loop like this one stays well under the request limit, so if you see it, something else in your organization is probably sharing the budget. Wait Retry-After seconds and rerun.
  • A timeout or connection error prints the httpx error message. Check your network, or raise timeout on the client if you send much larger tickets.
  • triage_results.jsonl is written only after every ticket is classified, so a failed run leaves the previous results in place.