
Open-source LLM evaluation framework for testing AI agents, RAG pipelines, chatbots, and other LLM applications with metrics such as task completion, tool correctness, answer relevancy, hallucination, and G-Eval.
Prices and medians update for the tier you select.
Ranked by how closely each one matches deepeval's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.