What we measured
We evaluated Agent API presets and Sonar models on the same workloads, across three benchmarks:- BrowseComp measures agentic browsing on hard questions that require chaining many searches.
- DSQA (DeepSearchQA) measures answer quality on deep-search questions.
- WideSearch measures how completely a run gathers and fills structured results, scored by row-level F1.
For a benchmark built specifically for research agents that must search both wide and deep, see WANDR.
Which preset should you use?
Results vary by use case, but every preset lands higher on the quality curve than its Sonar counterpart — and in the deep-research tier, for less. Use the mapping below as a starting point, then confirm against your own traffic.Explore the presets
See what each preset configures and how to override individual settings.
Migrate your integration
Follow the field-by-field procedure to move a Sonar integration to the Agent API.