The right provider depends on what you are asking it to do. ProviderGrade measures assessed quality — not popularity — on four separable axes, nightly, from runs you can re-check. Your agent asks at selection time and gets reasoning back.
Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, scored by a fixed LLM judge.
The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.
Axes are kept separable on purpose: agents weight them by their own constraints at query time.
Judge score for every (task, provider) pair. Hover a cell for the judge's reasoning; click to jump to the full text below.
Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.