ProviderGrade
Benchmarked ratings for the services agents consume

The right provider depends on what you are asking it to do. ProviderGrade measures assessed quality — not popularity — on four separable axes, nightly, from runs you can re-check. Your agent asks at selection time and gets reasoning back.

Code-execution providers, benchmarked for agents

Assessed quality, latency, cost, and reliability — every number computed from persisted benchmark runs, scored by a fixed LLM judge.

Preview ranking — not published methodology.
  • Five tasks only: provider gaps are well inside run-to-run noise. Treat the ordering as indicative, not decided.
  • Latency covers the full lifecycle an agent pays for — create, execute, teardown — measured from the benchmark host. Warm-pool or snapshot reuse is each provider's own optimization and shows up as their number.
  • Cost approximates published per-resource rates at each provider's default sandbox size, billed per second of measured lifetime; sizes are recorded per run in the pricing snapshot but are not identical across providers.

Ranking — weighted by your constraints

The composite is recomputed live from the four axis scores. Drag the sliders (or pick a preset) and watch the ranking change — this is the question a free popularity leaderboard can't answer.

Presets:

The four axes, side by side

Axes are kept separable on purpose: agents weight them by their own constraints at query time.

Quality per task

Judge score for every (task, provider) pair. Hover a cell for the judge's reasoning; click to jump to the full text below.

Judge reasoning — the audit trail

Every quality score links back to a stored judgment with reasoning, and every judgment links back to a raw run (request, response, latency, pricing snapshot, timestamp) preserved in that batch's archive.