Skip to main content
Use the Decisions API to give a confidence score to the web_search tool answers from the Agent API. Your application asks the Agent API a question, then sends each cited sentence and its source snippets from the Agent API response to the Decisions API. Your application gets back a probability that the sources support the sentence. Your code compares the probabilities to cutoffs and labels each sentence supported, weak, unsupported, contradicted, or uncited. Low-scoring sentences go back to the Agent API to search again, and your code scores the rewrite the same way. A web-grounded answer puts a citation after most sentences. A citation marker points to a search result. It does not prove that the sentence is on that page. The Decisions API answers β€œdoes passage 2 state or imply this claim” with a probability. Your code compares the probabilities to cutoffs, keeps what the sources support, and renders the report. Each probability is the model’s estimate of how well the snippets you sent support the sentence. It is not proof that the sentence is true, and a supported rewrite can leave out part of the original answer.

A live run of this recipe on the K2-18b question. The step captions and tool panels were added for the recording, and request IDs are shortened. The script prints the terminal lines and writes out/report.html and out/run.json.

Why the Decisions API

  • Several probabilities choose an action. Checkable, support, and contradiction together tell your code which sentences to keep and which need another search. The same policy then scores the rewrites.
  • Probabilities, not prose. Each yes/no (noul) question in this tutorial returns a probability from 0 to 1. Your cutoffs are plain comparisons, with no prose response to parse.
  • Many questions, one request. One request per sentence asks support and contradiction for each passage, plus whether the sentence is checkable. Two passages return five probabilities.
  • Graded results. A 0.682 and a 0.986 lead to different actions. Your code keeps the 0.986 and leaves the 0.682 unconfirmed.
  • Works without a source. A sentence with no citation still gets the checkable question, so your code can tell an uncited claim from a transition.
  • Cheap enough to run on every sentence. The scoring step costs a small fraction of the Agent API calls around it. See the pricing page.
  • Text and images. This tutorial is text only, but the same endpoint scores screenshots. See Drive a Browser with the Decisions API.

What you will build

A command-line tool. You give it a question. It gets a live answer from the Agent API, checks every sentence, investigates the weak ones, and writes a report. The project is five files in one folder:
Every run takes seven steps:
  1. agent_client.py asks the Agent API your question with web_search enabled. The answer has a marker like [3] after each factual sentence, plus numbered search results with snippets.
  2. answer_gate.py splits the answer into sentences and reads each sentence’s markers.
  3. decisions_client.py sends each sentence and its cited snippets to the Decisions API with three kinds of yes/no question: does each snippet support it, does each snippet contradict it, and is it checkable.
  4. answer_gate.py compares the probabilities to cutoffs and labels the sentence.
  5. For the first four weak, unsupported, contradicted, or uncited sentences, agent_client.py asks the Agent API to search again and rewrite the sentence. Any beyond four are marked β€œnot reviewed”.
  6. answer_gate.py scores the rewrite the same way. Supported rewrite sentences replace the original; the rest are dropped. If nothing in the rewrite is supported, the original stays, marked β€œnot confirmed”.
  7. answer_gate.py writes out/report.html and out/run.json.
The Agent API searches and writes. The Decisions API returns probabilities. Only answer_gate.py labels sentences and chooses what to keep. One run makes one Agent API call, up to four more for investigations, and one Decisions API request for every original and rewrite sentence. Retries can add attempts.

What you need

  • Python 3.10 or newer on macOS or Linux. Check with python3 --version; on macOS the system python3 can be older.
  • A Perplexity API key from the API Console. One key works for both APIs.

Set up

requirements.txt pins the two packages to the tested versions. httpx sends the requests. pytest runs the tests.
requirements.txt
Run these commands. Save requirements.txt in the new folder when the comment says to.
In a new terminal, run the source and export lines again.

The Agent API client: agent_client.py

agent_client.py makes two kinds of Agent API call: the first answer, and the follow-up search for one sentence.

Settings and types

agent_client.py (part 1 of 3)
INSTRUCTIONS asks for plain prose, a marker after every factual sentence, and sentences that name their subject instead of starting with β€œit”. Each sentence is scored on its own, so β€œIt reaches end of life in 2029” is hard to check without the sentence before it. The model does not always follow this; you will still see some β€œIts” sentences.

Asking and investigating

agent_client.py (part 2 of 3)
ask sends your question. investigate sends the question, the full first answer for context, and the one sentence that scored low. It asks for the fewest replacement sentences, one to three, that state only what the sources support and do not repeat the rest of the answer. Both calls enable web_search. The prompt asks for a fresh search, and parse requires search results in the response.

Reading the response

agent_client.py (part 3 of 3)
The response has a message with the text and search_results with numbered sources. The marker [3] means result id 3. parse keeps each result’s url, title, and snippet. The snippet is the passage the Decisions API scores against. If the model answered without searching, parse raises an error.

The Decisions API client: decisions_client.py

decisions_client.py scores one sentence at a time.

Settings and the result type

decisions_client.py (part 1 of 3)
Verdict holds the probabilities and usage the API returned. It holds no labels.

The request

decisions_client.py (part 2 of 3)
The state is a JSON object with the claim and numbered passages. Every question is a noul, which returns the probability that the answer is yes.
  • checkable: is the sentence a factual statement? Questions and opinions are not.
  • supports_N: does passage N state or directly imply the claim? Being on the same topic does not count.
  • contradicts_N: does passage N say something that cannot be true if the claim is true?
A sentence with no citation gets only checkable. Separate support and contradiction questions, instead of one choice, let your code see a sentence that one source supports and another contradicts.

Sending it

decisions_client.py (part 3 of 3)
score sends the request and retries twice on 429, 500, 502, or 503, waiting Retry-After seconds when given. It keeps the x-request-id header for support requests. The tests use the client argument to pass a fake transport.

The gate: answer_gate.py

answer_gate.py turns probabilities into labels and runs the review. It is shown in six parts.

Cutoffs and the sentence record

answer_gate.py (part 1 of 6)
The four cutoffs are starting points, not tested production values. Tune them on labeled examples of your own answers. NEEDS_REVIEW lists the labels that trigger a second search, and MAX_INVESTIGATIONS caps how many follow-up Agent API calls one answer can make. Sentence records the claim, the passages it was scored against, every probability, the request ID, and its rewrite.

Sentences

answer_gate.py (part 2 of 6)
The splitter cuts at ., !, or ? plus any closing quote, and collects the [n] markers after it. It also moves markers written before the period (fact [1].) to after it, keeps β€œJohn J. Hopfield”, β€œe.g.”, and β€œi.e.” whole, and treats text with no final punctuation as one sentence. If any text is left over, it raises an error instead of dropping it. It is built for the plain prose the instructions ask for; abbreviations like β€œDr.” can still split a sentence.

The policy

answer_gate.py (part 3 of 6)
First matching rule wins: Contradiction comes before support, so a sentence that one passage supports and another contradicts is contradicted. A marker that points to a search result that does not exist is noted in the rule. Probabilities print with three decimals, so a 0.696 near the 0.7 cutoff does not show as 0.70. The comparison uses the full value, which run.json keeps: a 0.6999 prints as 0.700 and is still weak.

The loop

answer_gate.py (part 4 of 6)
score_answer sends one Decisions API request per sentence and prints each label as it arrives. review sends the first four low-scoring sentences to investigate, scores each rewrite with the same score_answer, and keeps the rewrite if any of its sentences are supported. Low-scoring sentences past the limit are marked β€œnot reviewed”. final_sentences builds the answer you keep: supported sentences stay, replaced sentences become the supported part of their rewrite, and β€œnot confirmed” sentences stay with their low label.

Output

answer_gate.py (part 5 of 6)
write_report shows the final answer colored by label, with replaced sentences underlined and β€œnot confirmed” or β€œnot reviewed” next to originals that stayed. Rewrites come from separate searches, so citations are renumbered by URL. Below the answer, the evidence table lists every scored sentence, including dropped rewrite sentences, with each passage’s probabilities, the request ID, and the passage exactly as sent. summarize counts sentences, investigations, Decisions API requests, input tokens, and labels before and after review.

Command line

answer_gate.py (part 6 of 6)
The script takes one question and an optional --out folder. It writes report.html and run.json, which holds every sentence, passage, probability, and request ID. If a review request fails, it still writes both files with the first-pass scores and every investigation that completed. The investigation that failed is not saved.

Test it

Save all five files in your folder, then run the tests before your first live run. They need no API key or network. They cover sentence splitting, the label rules, the request body, retries, and the review step with fake API responses.
test_answer_gate.py

Run it

Each run asks the Agent API live, so the answer, the sources, and the scores change from run to run. Use a separate --out folder for each question.
Recorded on 2026-10-07 at 21:11 UTC with pplx-decider-v1.1-27b. Your output will differ. The recorded run wrote to a different --out folder, so its last line shows that path instead of out_k2.

Reading the output

First pass. One sentence scored supported. Sentence 2 scored weak at 0.643. Sentences 1 and 3 came back with no citation marker, so each got only the checkable question and was labeled uncited. Investigation. All three went back to the Agent API.
  • Sentence 1’s rewrite scored 0.856 and replaced it. The rewrite also adds that later analyses questioned the carbon dioxide evidence.
  • Sentence 2’s rewrite scored 0.986 and replaced it. The original said low ammonia is consistent with a water ocean. The rewrite adds that a magma ocean could also explain it.
  • Sentence 3’s rewrite scored 0.682, just under the 0.7 cutoff. Your code left the original in place, marked β€œnot confirmed”.
Cost and time. The run made 4 Agent API calls and 7 Decisions API requests for 20,625 input tokens, in 14.3 seconds. Other live runs on the same code: What the scores mean. Each probability is the model’s estimate that one passage supports one sentence. It does not measure real-world truth or source quality, and the code uses the best single passage rather than combining them.
  • Some derived claims scored low in our runs, even when a reader could infer them from the passages.
  • A supported sentence can carry a source’s extra precision. In one test run, a rewrite took an exact end-of-life day from a version tracker, while the official Python schedule gives only the month.
  • Dropping unsupported rewrite sentences can remove true details. In one run, a sentence listing Python 3.13 features was replaced by a single supported sentence about colored tracebacks.
  • Green means supported, not well edited. Rewrites can repeat facts from elsewhere in the answer.

Adapt it

  • Raise the bar for specific claims. Add a score question with levels like β€œgeneral”, β€œspecific”, and β€œcontains a number or date”, and require more support for specific sentences.
  • Change what happens to unconfirmed sentences. Drop them, flag them in your UI, or send them to a human reviewer instead of keeping them.
  • Check that a rewrite still answers the question. Add a noul question asking whether the replacement keeps the essential information of the original. Like the cutoffs, test it on your own answers first.
  • Run it before an answer reaches a user. Call score_answer and review in your application, and alert when the share of unconfirmed sentences changes.

Troubleshooting

  • Set PERPLEXITY_API_KEY first: run the export line in this terminal.
  • 401 from either API: the key is wrong or inactive.
  • The Agent API response has no search results: the model answered without searching. Ask again or rephrase.
  • Could not split paragraph: the answer has text the splitter does not recognize. The error shows the text; adjust SENTENCE for it.
  • Many sentences are uncited: the model skipped citation markers. The review step investigates them, up to MAX_INVESTIGATIONS.
  • 429: you hit the rate limit. The client retries twice; wait and rerun if it still fails.

The complete files

Each file in full, for copying.
requirements.txt
agent_client.py
decisions_client.py
answer_gate.py
test_answer_gate.py