Skip to main content
Most browser agents ask a chat model to describe the next click in text, then parse an element out of the reply with no measure of how certain the model was. When the click is wrong, nothing tells you it was a guess. The Decisions API returns probabilities instead of text. Your script numbers everything clickable on the page, screenshots it, and sends the picture to pplx-decider-v1.1-27b with the goal and six questions: which badge advances the goal, whether the goal is already reached, whether a banner is in the way, and three more. The model never clicks anything. A six-rule policy in your code compares the probabilities to cutoffs and does the clicking, scrolling, and stopping. The recipe is two Python files: one that calls the API and one that drives the browser with Playwright. The default goal navigates the public Perplexity docs to the Decisions API quickstart and finds the section about images, in five steps and about 37,000 input tokens.

A live run of this recipe. The side panel shows the probabilities from each Decisions API response and the policy rule they trigger. The panel was added for the recording. The script prints the top candidates to the console and writes every probability to out/run.jsonl.

What you need

  • Python 3.10 or newer on macOS or Linux. The commands below use Unix shell syntax.
  • A Perplexity API key from the API Console.
  • About 700 MB of free disk space for the Chromium browser that Playwright downloads.

Set up

The project is six files in one folder, all copied from this page. The script creates out/ when it runs.
Every step runs the same loop:
  1. browser_agent.py asks Playwright to run page_script.js, which draws a numbered badge on everything clickable.
  2. Playwright takes a screenshot.
  3. decisions_client.py sends the goal, the screenshot, and six questions to the Decisions API, which returns a probability for every answer.
  4. browser_agent.py compares the probabilities to its cutoffs and picks one action.
  5. Playwright clicks or scrolls. Repeat until the goal is reached or the policy stops.
Only browser_agent.py decides. Playwright never sees the goal, and the API never touches the browser. requirements.txt pins the three packages to the versions this tutorial was tested with.
requirements.txt
httpx sends the API requests. playwright controls the browser. pytest runs the tests. pyproject.toml holds settings for ruff (code style), mypy (type checks), and pytest. You do not need ruff or mypy to run the agent.
pyproject.toml
Create the folder, save requirements.txt and pyproject.toml in it, then run the remaining commands. The virtual environment (.venv) keeps this project’s packages separate from the rest of your system.
The Chromium download is the slowest step. In a new terminal, run the source and export lines again.

The page script

page_script.js runs inside the web page. Playwright passes it to the browser with page.evaluate. It is one function that takes a single command: "number" or "scroll".
page_script.js
Save this as page_script.js in the project folder. SCROLLERS collects every scrollable container: the page itself and any tall element with its own scrollbar, such as a docs sidebar. number does five things:
  1. Removes the badges, numbers, and outline from the previous step. Without this, an element numbered 20 last step would keep that number, and the script could click it instead of the element with the highest probability.
  2. Finds every link, button, input, select, and element with a button, link, or tab role that is on screen, at least 8 pixels wide and tall, visible, and enabled.
  3. Skips elements with no readable text and elements covered by another element, such as a link under a sticky header, by checking what the browser reports at the element’s center point. Elements inside iframes or shadow DOM are not found.
  4. Numbers each remaining element, stores the number on it as data-dm, and draws a red badge beside the element’s first line of text. Placement matters: in a sidebar with indented items, a badge at the corner of the element’s box sat 50 pixels from its label, and the probability on the correct element dropped from 94% to 65%.
  5. Returns a legend (number, kind, label) and canScroll. Short labels such as “Quickstart” include the nearest section heading, so the log can distinguish several “Quickstart” links. The legend is for the log only; it is never sent to the API.
scroll moves the page and every scrollable container down by three quarters of the window.

The API client: decisions_client.py

decisions_client.py calls the Decisions API and has no knowledge of the browser. Copy the complete file from The complete files, then read it here in three parts.

Endpoint, model, and retry settings

  • API_URL and MODEL set the endpoint and the model.
  • REQUEST_TIMEOUT_S, RETRY_STATUSES, MAX_RETRIES, and MAX_RETRY_DELAY_S control the timeout and retries. A 429 (rate limit), a 5xx (server error), or a dropped connection is retried twice. A 400 or 401 is a problem with the request and is not retried. If the server asks for a wait longer than 30 seconds, the script stops with an error rather than retry early.
  • RISK_LEVELS is a four-level scale from “Read only: browsing or navigating” to “Payment, purchase, or deletion,” used by the risk question and by the policy.
  • DecisionsError is raised when the API returns an error, or a 200 that is missing a field this script reads. A dedicated exception lets the agent catch this case and print one line instead of a stack trace.
  • make_client builds the single HTTP client for the run, with the API key and timeout set in one place.

The questions

build_questions returns the six questions sent on every step. The Decisions API has three question types, and this file uses all of them:
  • noul is a yes-or-no question. The answer is the probability of yes. blocked, target_visible, and goal_reached are noul.
  • choice is a select-one question. You supply the options; the answer is a probability for each. dismiss and next_action are choice. Their options are the badge numbers 1 to N with no description (None, sent as null), so the model must read each number from the screenshot and identify the element beside it.
  • score is a rating question with ordered levels. risk is score, using the four RISK_LEVELS.
No question names a website or page. Each refers to “the goal,” and the goal text is sent in the request.

The request

decide sends one request per step, plus retries if needed. The state is a list of four items: the goal, the previous action, the screenshot (as base64 text), and one sentence stating that the badges are numbered 1 to N. The goal is sent every time because the model retains nothing between requests. decide takes the httpx.Client as an argument rather than creating its own, so the tests can pass a fake client that returns prepared responses. The loop makes up to three attempts. A 200 returns the parsed JSON, the latency of that attempt in milliseconds, and the request id from the x-request-id header. A 429, 5xx, or dropped connection waits and retries; _retry_delay uses the Retry-After header when the server sends a number, otherwise 1 second, then 2. If the header asks for more than MAX_RETRY_DELAY_S, the loop stops with an error rather than retry early. Any other status raises DecisionsError with the status code, the API’s message, and the request id. Include that id in any support request. _parse checks a 200 for the fields this script reads: all six answers, probabilities for every choice and score, a choice that is one of the scored options, a number for every noul, and a token count. Anything else becomes a DecisionsError instead of a KeyError later. It is a sanity check, not full schema validation. The check iterates over build_questions(count), so the question names are defined in one place.

The agent: browser_agent.py

browser_agent.py drives the browser and makes the decisions. Copy the complete file from The complete files, then read it here in six parts.

Imports and cutoffs

  • The imports from decisions_client require both files to be in the same folder.
  • VIEWPORT is the browser window, 1280 by 800 pixels. The API reads images in 32 by 32 pixel tiles and accepts up to 2,048 per image; 1280 by 800 is 1,000 tiles. The screenshot accounts for about 6,000 of the roughly 7,300 input tokens in each request.
  • PAGE_SCRIPT reads page_script.js from the same folder once, at startup.
  • The six values from ACT_MIN to RISK_STOP_MIN are the cutoffs, and they are the entire policy. ACT_MIN = 0.6 means “click only if the model puts at least 0.6 on its top choice.” A probability is the model’s score, not a measured success rate. The cutoffs were chosen by observing where answers landed on these pages, and most sit far from 50% so that small run-to-run differences do not change the agent’s behavior.

Actions

Action is what the policy returns: a kind, a detail sentence for the console, the badge to click when there is one, and a stop_reason when the run is ending. Kind and StopReason are Literal types, so mypy catches a misspelled value before you run. Action is a frozen dataclass: nothing can modify it after the policy creates it.

The policy

choose is where the decision is made: it takes the model’s probabilities and returns one Action. It touches neither the browser nor the network, so the tests can check each rule in under a second. The order of the checks is the policy, read top to bottom:
  1. If goal_reached is at least 80%, stop. Completion takes priority over everything else.
  2. If the next click is probably a payment, purchase, or deletion (0.7 or more on level 3), stop for human review. This runs before any click, including a banner dismissal. It is a model’s estimate from a screenshot, not a safeguard, so use this agent on pages where a wrong click is harmless.
  3. If a banner is probably in the way (80%) and one element has at least 60% as the one that closes it, dismiss it. The loop remembers the last banner button it clicked so the same one is never clicked twice.
  4. If the target is probably not on screen (under 50%) and the page can still scroll, scroll.
  5. If the best click candidate is at least 60%, click it.
  6. Otherwise stop rather than guess.

Acting on the page

act is the only function that changes the page. If the action has a badge, it outlines that element in green, waits pause_s seconds (1.5 by default) so the decision is visible, and saves out/step-NN-decision.png. A scroll then runs the page script’s scroll command. A click finds the element by its data-dm number, removes any target attribute so links open in the same tab, clicks, and waits for the page to load; some clicks change the page without a load event, so that wait may time out silently. A stop changes nothing. Finally, it removes the badges.

The loop

run loads the start page, opens out/run.jsonl, and for each step:
  1. Numbers the page and takes a screenshot, saved as out/step-NN-input.png. This is exactly what the model receives. If nothing is clickable, it waits two seconds and checks again, then stops with no_candidates.
  2. Calls decide for the probabilities, then choose for the action.
  3. Prints the action, latency, tokens, and request id, plus the top three click candidates with their labels. Writes everything to run.jsonl, including every probability at full precision.
  4. Calls act.
  5. Stops if the action was a stop. Otherwise records the action, in words, for the next request’s state.
It returns a summary: the goal, the number of steps, why it stopped, and the final URL.

Command line and entry point

parse_args defines the options. Each has a default, so python browser_agent.py runs the tutorial flow. _step_count rejects a --max-steps value outside 1 to 100 before the browser opens. main reads the API key from the environment and returns an exit code: 2 if the key is missing, 1 if the API, network, or browser fails, 0 when the loop ends on its own for any stop reason. It closes the browser even when run raises, and prints API, network, and browser errors to standard error as a single line. It deletes any previous out/summary.json before starting and writes a new one when the loop ends. On an error, out/ keeps the screenshots and run.jsonl written so far, but no summary.json. The last two lines run main only when the file is executed directly, not when the tests import it.

Run it

From the project folder, with the virtual environment active and your key exported:
--headed shows the browser window. The defaults start on the docs overview page with the goal “Open the Decisions API quickstart and find the section about images in state.” Pass --start-url and --goal to use your own. Output from a headless run on 2026-10-06 at 18:17 UTC. The labels in the console come from the local legend; the model received only the badge numbers.
Five requests and 36,944 input tokens (tokens are how the API measures and bills a request): about 7,300 to 7,500 per step, of which about 6,000 are the screenshot. Our runs took 16 to 22 seconds, most of it page loads and the 1.5 second pause per decision. The first request took about a second; the others took a quarter to half a second. Open out/step-03-input.png to see the frame the model received when it put 99.8% on badge 20, and out/step-03-decision.png for the same frame with the chosen link outlined. out/run.jsonl has every probability for every step.

Reading the numbers

The model produced the probabilities; the policy produced the action. Steps 1 and 2: the policy scrolls. target_visible is 7% and then 8%: the Decisions API section is below the visible part of the sidebar. The click candidates are spread thin, the best at about a third. That is the expected result when the right answer is not on the page. The policy checks target_visible before the click, so it scrolls instead of clicking a weak candidate. Step 3: badge 20 receives 99.8%. “Quickstart” under the “Decisions API” heading is now on screen next to badge 20. target_visible rises to 99%, badge 20 receives 99.8%, the runner-up 0.1%, and the policy clicks. Three links on screen say “Quickstart”; scoring the right one requires reading the heading above it in the image. Step 4: a new page, a new badge 31. On the quickstart page, “Images in state” in the “On this page” list is beside badge 31. The model puts 99% on it and the policy clicks. Badge numbers are reassigned every step; nothing carries over. Step 5: goal_reached crosses the cutoff. It stayed far below 80% for four steps. With the “Images in state” heading at the top of the page it exceeds 99%, and the policy stops. Nothing in the code judges whether the goal is met; it asks the model and compares the answer to a cutoff. Identical requests return similar probabilities, not identical ones. Across our runs, badge 20 on step 3 ranged from 93% to 99.8%, badge 31 on step 4 from 92% to 99%, and the step 1 and 2 candidates varied below 40%. Every action was the same because the cutoffs sit far from those values. Leave margin when you set cutoffs.

Adapt it

  • Your goal and your site. Pass --start-url and --goal. Write the goal as a plain instruction: “Open the pricing page and find the enterprise tier.”
  • Add the labels. For pages with many similar elements, put each element’s label in its option description instead of None: {str(x["n"]): f"{x['kind']} '{x['label']}'" for x in legend}. In our test this raised the step 3 probability from 93% to 96% and added about 3,000 input tokens per step.
  • A more conservative stop. Lower RISK_STOP_LEVEL to 2 to stop before sign-ins and plan changes as well; choose sums every level at or above it. To ask for confirmation instead of stopping, replace that branch’s Action("stop", ...) with an input() prompt and click only on yes.
  • Form filling. Add a choice question for which input to fill and a source for the value. The model scores the inputs from the screenshot; your code selects the top one and types into it.
  • Verify each action. Add a noul question, “Did the last action described in the state produce the change it was meant to produce?”, and log it.
  • Larger pages. Keep the screenshot at or under 2,048 tiles of 32 by 32 pixels (1440 by 1440 or 2048 by 1024 fit). choice accepts up to 255 options; the page script caps at 250.

Troubleshooting

  • Set PERPLEXITY_API_KEY means the variable is not set in this terminal. Run the export line again.
  • Decisions API failed: 401 {"error":{"message":"Invalid API key provided. Ensure your API key is correct and active.","type":"invalid_api_key","code":401}} means the key is wrong or inactive.
  • Decisions API failed: 400 is a problem with the request; error.message gives the reason. If you raise the page script’s 250-badge cap, keep it at or under 255, the choice option limit. An image over 2,048 tiles returns 504 after about a minute, not 400; if you change VIEWPORT, keep width ÷ 32 × height ÷ 32 at or under 2,048.
  • Decisions API failed: 429 means your organization exceeded the request limit (10 per second) or the token limit, and two retries did not clear it. Another process is likely sharing the limit; wait and rerun.
  • FileNotFoundError: ... page_script.js or ModuleNotFoundError: No module named 'decisions_client' means a file is not next to browser_agent.py.
  • No candidate above 60%. Open out/step-NN-input.png. If two badges sit side by side next to the expected element, the image is ambiguous. If the right element is the top choice at around 50%, the goal sentence may be too vague.
  • The agent scrolls to the bottom and reaches --max-steps. Nothing on the page advances the goal as written, or the target is inside a scrollable container the page script does not detect.
  • The page shows “Just a moment…” or a verification challenge. Run with --headed; some sites challenge headless browsers.
  • A click opens a new tab. target="_blank" links are handled; buttons that open windows with JavaScript are not. Add a context.on("page") handler if you need them.

Test it

The tests use no API key, network, or browser. Save test_browser_agent.py next to the other files.
test_browser_agent.py
  • answers builds a fake set of answers in the same shape the API returns, so each test can set only the probabilities it cares about.
  • The policy. Six tests, one per rule, in policy order: stop when the goal is reached, stop on payment risk before any click, dismiss a banner only once, scroll when the target is off screen, click a confident candidate, and stop rather than guess.
  • The request. fake_client builds an httpx.Client on httpx.MockTransport, which returns prepared responses instead of calling the network. One test checks that the request carries the model, the screenshot, and the six questions. The other checks that a 429 is retried.
You should see:
Optional: the checks we ran before publishing. You should see exactly three lines: All checks passed!, 3 files already formatted, and Success: no issues found in 3 source files.

The complete files

The exact files used for the run above.
requirements.txt
pyproject.toml
page_script.js
decisions_client.py
browser_agent.py
test_browser_agent.py