One run. Every cell.

queries against APIs, scored by judges. Every chart has a table view. Raw scores are on the data page.

how that was measured

queries
categories
complete ensembles
%
of responses
vendor spend
this run, all vendors
published weeks
in the public export

Quality by vendor and category

Each cell is the mean of the three-judge median. Queries without a complete ensemble are dropped.

Score matrix

0–10

Where each vendor's queries land

share of scored queries
The score column below is a mean; this is the spread behind it. Almost everything lands in the top three bands, which is why the query set is listed as a known limitation.

Vendor summary

this run
Overall score is the mean of a vendor's category scores. "Won outright" counts queries where one vendor alone took the highest ensemble median; "among best" counts queries where it was level with the highest. Those differ a great deal here, because of queries ended in a tie, which is itself the most useful thing this table says.
Vendor Score Cost / query Run cost p50 latency Won outright Among best Scored Returns
Routing headroom

Per query it is thinner still: of queries end with two or more vendors tied for the best median, so on most of them there is no pick to make.

Cost against quality

Cost is on a log scale. Volume discounts are not modelled.

Cost per query vs score

list price per query
Cost is what the vendor reported for this run where the API returns it, and list price otherwise. Volume discounts are not modelled.

Share of the best score, by category

How much of the category-leading score this vendor reaches. Bars start at 70%.

How long each API takes to answer

Median wall-clock time from request to response, measured by the harness.

Median response time

p50, measured
Measured over a residential connection during a single run. Treat the ordering as the signal and the absolute values as indicative.

How much the judges disagree

A single judge family would produce different absolute scores for the same responses.

How far apart the three judges were

per response
The mean disagreement quoted everywhere else on this site is a summary of this shape. The tail past three points is where the ensemble is doing real work.

Mean score by judge family

same responses
Same responses, same rubric, same prompt. The gap is the judge, not the vendor.
family spread
points, highest to lowest family mean
mean disagreement
points, max minus min per response
differ by > 3 points
of responses

Drop any one judge family and the top of places hold. Scored by a single family alone, of rank the vendors differently. How that was measured.

Every query, every score

Hover or focus a cell for the three individual judge scores.

Blank cells are responses that did not receive all three judge scores, usually a rate limit. They are excluded from every aggregate on this page. Breaking-news queries carry no gold answer by design, and are graded on source recency.