Perplexity

Ranked on the overall table for the run, at per query. It returns .

Overall score
ensemble median, 0-10
Cost / query
for the whole run
p50 latency
measured, not vendor-reported
Won outright
of queries; ended in a tie
Scored cleanly
%
of responses

Where Perplexity holds up, and where it does not

Each bar is how much of the category-leading score Perplexity reaches.

Share of the best score

by category
Bars start at 70%. The accent marks categories below 90% of the best available score.

A vendor that trails overall can still be the right call for one kind of query. That is why scores are reported per category. See the full table.

Perplexity, category by category

A cell below the coverage floor is shown as missing rather than as a thin sample.

Score is the mean of the ensemble medians in that category. "Of best" is the share of the category-leading score.
Category Score Of best Behind leader Category leader Scored

Against the rest of the set

The same table as the results page, with Perplexity marked.

Vendor Score Cost / query p50 latency Won outright Among best Returns

Known limits

of data. Scores carry roughly points of judge-to-judge disagreement. The judges are unaudited against humans. Known limits.

One configuration is measured. If it misrepresents Perplexity (wrong tier, depth or parameters), corrections get a changelog entry, not a silent edit. Report it, or recompute it from the export.