
An open-source evaluation harness and dataset for benchmarking how well LLM agent sessions generate accessible HTML, using Playwright, axe-core, and custom per-test assertions.
Prices and medians update for the tier you select.
Ranked by how closely each one matches A11y Llm Eval's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.