web_search tool answers from the Agent API. Your application asks the Agent API a question, then sends each cited sentence and its source snippets from the Agent API response to the Decisions API. Your application gets back a probability that the sources support the sentence. Your code compares the probabilities to cutoffs and labels each sentence supported, weak, unsupported, contradicted, or uncited. Low-scoring sentences go back to the Agent API to search again, and your code scores the rewrite the same way.
A web-grounded answer puts a citation after most sentences. A citation marker points to a search result. It does not prove that the sentence is on that page. The Decisions API answers βdoes passage 2 state or imply this claimβ with a probability. Your code compares the probabilities to cutoffs, keeps what the sources support, and renders the report.
Each probability is the modelβs estimate of how well the snippets you sent support the sentence. It is not proof that the sentence is true, and a supported rewrite can leave out part of the original answer.
A live run of this recipe on the K2-18b question. The step captions and tool panels were added for the recording, and request IDs are shortened. The script prints the terminal lines and writes out/report.html and out/run.json.
Why the Decisions API
- Several probabilities choose an action. Checkable, support, and contradiction together tell your code which sentences to keep and which need another search. The same policy then scores the rewrites.
- Probabilities, not prose. Each yes/no (
noul) question in this tutorial returns a probability from 0 to 1. Your cutoffs are plain comparisons, with no prose response to parse. - Many questions, one request. One request per sentence asks support and contradiction for each passage, plus whether the sentence is checkable. Two passages return five probabilities.
- Graded results. A 0.682 and a 0.986 lead to different actions. Your code keeps the 0.986 and leaves the 0.682 unconfirmed.
- Works without a source. A sentence with no citation still gets the checkable question, so your code can tell an uncited claim from a transition.
- Cheap enough to run on every sentence. The scoring step costs a small fraction of the Agent API calls around it. See the pricing page.
- Text and images. This tutorial is text only, but the same endpoint scores screenshots. See Drive a Browser with the Decisions API.
What you will build
A command-line tool. You give it a question. It gets a live answer from the Agent API, checks every sentence, investigates the weak ones, and writes a report. The project is five files in one folder:agent_client.pyasks the Agent API your question withweb_searchenabled. The answer has a marker like[3]after each factual sentence, plus numbered search results with snippets.answer_gate.pysplits the answer into sentences and reads each sentenceβs markers.decisions_client.pysends each sentence and its cited snippets to the Decisions API with three kinds of yes/no question: does each snippet support it, does each snippet contradict it, and is it checkable.answer_gate.pycompares the probabilities to cutoffs and labels the sentence.- For the first four weak, unsupported, contradicted, or uncited sentences,
agent_client.pyasks the Agent API to search again and rewrite the sentence. Any beyond four are marked βnot reviewedβ. answer_gate.pyscores the rewrite the same way. Supported rewrite sentences replace the original; the rest are dropped. If nothing in the rewrite is supported, the original stays, marked βnot confirmedβ.answer_gate.pywritesout/report.htmlandout/run.json.
answer_gate.py labels sentences and chooses what to keep.
One run makes one Agent API call, up to four more for investigations, and one Decisions API request for every original and rewrite sentence. Retries can add attempts.
What you need
- Python 3.10 or newer on macOS or Linux. Check with
python3 --version; on macOS the systempython3can be older. - A Perplexity API key from the API Console. One key works for both APIs.
Set up
requirements.txt pins the two packages to the tested versions. httpx sends the requests. pytest runs the tests.
requirements.txt
requirements.txt
requirements.txt in the new folder when the comment says to.
Install and set your key
Install and set your key
source and export lines again.
The Agent API client: agent_client.py
agent_client.py makes two kinds of Agent API call: the first answer, and the follow-up search for one sentence.
Settings and types
Settings and types
Settings and types
INSTRUCTIONS asks for plain prose, a marker after every factual sentence, and sentences that name their subject instead of starting with βitβ. Each sentence is scored on its own, so βIt reaches end of life in 2029β is hard to check without the sentence before it. The model does not always follow this; you will still see some βItsβ sentences.
Asking and investigating
Asking and investigating
Asking and investigating
ask sends your question. investigate sends the question, the full first answer for context, and the one sentence that scored low. It asks for the fewest replacement sentences, one to three, that state only what the sources support and do not repeat the rest of the answer. Both calls enable web_search. The prompt asks for a fresh search, and parse requires search results in the response.
Reading the response
Reading the response
Reading the response
message with the text and search_results with numbered sources. The marker [3] means result id 3. parse keeps each resultβs url, title, and snippet. The snippet is the passage the Decisions API scores against. If the model answered without searching, parse raises an error.
The Decisions API client: decisions_client.py
decisions_client.py scores one sentence at a time.
Settings and the result type
Settings and the result type
Settings and the result type
Verdict holds the probabilities and usage the API returned. It holds no labels.
The request
The request
The request
noul, which returns the probability that the answer is yes.
checkable: is the sentence a factual statement? Questions and opinions are not.supports_N: does passage N state or directly imply the claim? Being on the same topic does not count.contradicts_N: does passage N say something that cannot be true if the claim is true?
checkable. Separate support and contradiction questions, instead of one choice, let your code see a sentence that one source supports and another contradicts.
Sending it
Sending it
Sending it
score sends the request and retries twice on 429, 500, 502, or 503, waiting Retry-After seconds when given. It keeps the x-request-id header for support requests. The tests use the client argument to pass a fake transport.
The gate: answer_gate.py
answer_gate.py turns probabilities into labels and runs the review. It is shown in six parts.
Cutoffs and the sentence record
Cutoffs and the sentence record
Cutoffs and the sentence record
NEEDS_REVIEW lists the labels that trigger a second search, and MAX_INVESTIGATIONS caps how many follow-up Agent API calls one answer can make. Sentence records the claim, the passages it was scored against, every probability, the request ID, and its rewrite.
Sentences
Sentences
Sentences
., !, or ? plus any closing quote, and collects the [n] markers after it. It also moves markers written before the period (fact [1].) to after it, keeps βJohn J. Hopfieldβ, βe.g.β, and βi.e.β whole, and treats text with no final punctuation as one sentence. If any text is left over, it raises an error instead of dropping it. It is built for the plain prose the instructions ask for; abbreviations like βDr.β can still split a sentence.
The policy
The policy
The policy
contradicted. A marker that points to a search result that does not exist is noted in the rule. Probabilities print with three decimals, so a 0.696 near the 0.7 cutoff does not show as 0.70. The comparison uses the full value, which run.json keeps: a 0.6999 prints as 0.700 and is still weak.
The loop
The loop
The loop
score_answer sends one Decisions API request per sentence and prints each label as it arrives. review sends the first four low-scoring sentences to investigate, scores each rewrite with the same score_answer, and keeps the rewrite if any of its sentences are supported. Low-scoring sentences past the limit are marked βnot reviewedβ. final_sentences builds the answer you keep: supported sentences stay, replaced sentences become the supported part of their rewrite, and βnot confirmedβ sentences stay with their low label.
Output
Output
Output
write_report shows the final answer colored by label, with replaced sentences underlined and βnot confirmedβ or βnot reviewedβ next to originals that stayed. Rewrites come from separate searches, so citations are renumbered by URL. Below the answer, the evidence table lists every scored sentence, including dropped rewrite sentences, with each passageβs probabilities, the request ID, and the passage exactly as sent. summarize counts sentences, investigations, Decisions API requests, input tokens, and labels before and after review.
Command line
Command line
Command line
--out folder. It writes report.html and run.json, which holds every sentence, passage, probability, and request ID. If a review request fails, it still writes both files with the first-pass scores and every investigation that completed. The investigation that failed is not saved.
Test it
Save all five files in your folder, then run the tests before your first live run. They need no API key or network. They cover sentence splitting, the label rules, the request body, retries, and the review step with fake API responses.test_answer_gate.py
test_answer_gate.py
Run the tests
Run the tests
Run it
Each run asks the Agent API live, so the answer, the sources, and the scores change from run to run. Use a separate--out folder for each question.
Run the gate
Run the gate
pplx-decider-v1.1-27b. Your output will differ. The recorded run wrote to a different --out folder, so its last line shows that path instead of out_k2.
Observed output
Observed output
Reading the output
First pass. One sentence scored supported. Sentence 2 scored weak at 0.643. Sentences 1 and 3 came back with no citation marker, so each got only thecheckable question and was labeled uncited.
Investigation. All three went back to the Agent API.
- Sentence 1βs rewrite scored 0.856 and replaced it. The rewrite also adds that later analyses questioned the carbon dioxide evidence.
- Sentence 2βs rewrite scored 0.986 and replaced it. The original said low ammonia is consistent with a water ocean. The rewrite adds that a magma ocean could also explain it.
- Sentence 3βs rewrite scored 0.682, just under the 0.7 cutoff. Your code left the original in place, marked βnot confirmedβ.
- Some derived claims scored low in our runs, even when a reader could infer them from the passages.
- A supported sentence can carry a sourceβs extra precision. In one test run, a rewrite took an exact end-of-life day from a version tracker, while the official Python schedule gives only the month.
- Dropping unsupported rewrite sentences can remove true details. In one run, a sentence listing Python 3.13 features was replaced by a single supported sentence about colored tracebacks.
- Green means supported, not well edited. Rewrites can repeat facts from elsewhere in the answer.
Adapt it
- Raise the bar for specific claims. Add a
scorequestion with levels like βgeneralβ, βspecificβ, and βcontains a number or dateβ, and require more support for specific sentences. - Change what happens to unconfirmed sentences. Drop them, flag them in your UI, or send them to a human reviewer instead of keeping them.
- Check that a rewrite still answers the question. Add a
noulquestion asking whether the replacement keeps the essential information of the original. Like the cutoffs, test it on your own answers first. - Run it before an answer reaches a user. Call
score_answerandreviewin your application, and alert when the share of unconfirmed sentences changes.
Troubleshooting
Set PERPLEXITY_API_KEY first: run theexportline in this terminal.401from either API: the key is wrong or inactive.The Agent API response has no search results: the model answered without searching. Ask again or rephrase.Could not split paragraph: the answer has text the splitter does not recognize. The error shows the text; adjustSENTENCEfor it.- Many sentences are
uncited: the model skipped citation markers. The review step investigates them, up toMAX_INVESTIGATIONS. 429: you hit the rate limit. The client retries twice; wait and rerun if it still fails.
The complete files
Each file in full, for copying.requirements.txt
requirements.txt
agent_client.py
agent_client.py
decisions_client.py
decisions_client.py
answer_gate.py
answer_gate.py
test_answer_gate.py
test_answer_gate.py