Skip to main content
Perplexity

Search documentation

Type to search this documentation.

On this pageOverview

Agent API vs Sonar benchmarks

The lead is widest on hard, multi-step tasks like BrowseComp and WideSearch, where the strongest preset more than doubles the best Sonar score.

We evaluated Agent API presets and Sonar models on the same workloads, across three benchmarks:

  • BrowseComp measures agentic browsing on hard questions that require chaining many searches.
  • DSQA (DeepSearchQA) measures answer quality on deep-search questions.
  • WideSearch measures how completely a run gathers and fills structured results, scored by row-level F1.

Agent API presets can even match Sonar's best model, Deep Research, at a fraction of the cost.

Results vary by use case, but every preset lands higher on the quality curve than its Sonar counterpart — and in the deep-research tier, for less. Use the mapping below as a starting point, then confirm against your own traffic.

Sonar Chat Completions Agent API preset Best for What improves
Sonar fast Single-fact lookups, definitions, quick summaries. More accurate than Sonar, with the clearest gains on research-style, multi-step questions.
Sonar Pro fast Everyday research questions, light multi-step lookups with current information. Much more accurate, and unlocks the real multi-step, agentic research that Sonar Pro can't reach.
Sonar Reasoning Pro low Multi-hop browsing and wide aggregation across many sources. Far stronger on step-by-step reasoning and on chaining evidence across several rounds of search.
Sonar Deep Research high Expert-level reasoning and exhaustive source coverage. The most accurate tier on demanding expert-level and deep-research work, and often at a lower per-request cost than Sonar Deep Research.
Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu