Open-source AI evaluation framework with 50+ production-grade graders for assessing LLM agents, multimodal models, code generation, and math reasoning. Integrates with LangSmith, Langfuse, and RL frameworks.
Prices and medians update for the tier you select.
Ranked by how closely each one matches openjudge's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.