Open-source benchmark for evaluating LLM tool-calling in agentic workflows across OpenAI-compatible serving endpoints, with deterministic scenarios, conversation traces, structured-output checks, and optional performance tests.
Prices and medians update for the tier you select.
Ranked by how closely each one matches tool-eval-bench's job. Prices show each provider's Pro state; entry prices are labelled as such. Unpriced products still belong to the market.
Market = the products most similar to this one by capability; prices are median / quartiles over its priced members, separated by provider type and buyer tier.