A zero-dependency Python package for statistically paired A/B evaluation of LLM models, agents, and harnesses, combining programmatic gates with rubric scoring and exact paired tests.
As parsed from alloevil.github.io/paired-eval/ on 2026-09-08.
Explore the market
Prices and medians update for the tier you select.
Compare at
paired-eval · Pro
No Pro plan
—
Pro-provider median
—
insufficient comparable pricing
vs Pro providers
No comparison
paired-eval doesn't sell Pro
ⓘ Comparisons are same provider type (provider) and same buyer tier (Pro). Never across tiers.
Market density
13th pctile
7 competing vendors
Comparable listings · 30d
—
Price spectrum · Pro plans · 1 of 3 providers priced · log scale · $40 → $40/mo
APIs◻ shaded = middle 50% · line = median (all types)paired-eval has no comparable Pro price — not plotted
Products in this market
Ranked by how closely each one matches paired-eval's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.