One run. Every cell.
queries against APIs, scored by judges. Every chart has a table view. Raw scores are on the data page.
Quality by vendor and category
Each cell is the mean of the three-judge median. Queries without a complete ensemble are dropped.
Score matrix
0–10Where each vendor's queries land
share of scored queriesVendor summary
this run| Vendor | Score | Cost / query | Run cost | p50 latency | Won outright | Among best | Scored | Returns |
|---|
Per query it is thinner still: of queries end with two or more vendors tied for the best median, so on most of them there is no pick to make.
Cost against quality
Cost is on a log scale. Volume discounts are not modelled.
Cost per query vs score
list price per queryShare of the best score, by category
How long each API takes to answer
Median wall-clock time from request to response, measured by the harness.
Median response time
p50, measuredHow much the judges disagree
A single judge family would produce different absolute scores for the same responses.
How far apart the three judges were
per responseMean score by judge family
same responsesDrop any one judge family and the top of places hold. Scored by a single family alone, of rank the vendors differently. How that was measured.
Every query, every score
Hover or focus a cell for the three individual judge scores.
No query matches that filter.
Blank cells are responses that did not receive all three judge scores, usually a rate limit. They are excluded from every aggregate on this page. Breaking-news queries carry no gold answer by design, and are graded on source recency.