
Agentbench
Regression testing framework for AI agents. Provides replay, evaluation, assertions, and CI integration to catch regressions in agent behavior—like Jest for AI…
What it does
The specific capability behind this listing, and where to get it.
Agentbench
Regression testing framework for AI agents. Provides replay, evaluation, assertions, and CI integration to catch regressions in agent behavior—like Jest for AI…
Agentery has not yet captured structured capability detail for this provider.
Official Agentbench links
Source repository available · no commercial pricing observed.
No price does not imply the product is free. Any code-host platform pricing is excluded.
Is Agentbench good value?
Price is straightforward; the useful comparison is capability, compatibility and operational cost.
Source repository available · no commercial pricing observed.
A public repository, but no identified licence or self-host evidence yet — so open-source / free-to-self-host is not asserted.
Check capability before deciding
Compare capability, compatibility and operational cost against comparable providers — Agentery keeps the price status explicit and never invents a verdict.
View the full niche →Comparable Agent Observability Eval
Alternatives in the same niche, with observed price and liveness where available.






See the MCP response behind this page · get_agent_profile()
See the MCP response behind this pageget_agent_profile
{
"agent_id": "agentbench",
"name": "Agentbench",
"url": "https://github.com/1304674612/agentbench",
"logo": "https://github.com/1304674612.png?size=200",
"niche": "agent-observability-eval",
"category": "developer-tools-infra",
"short_summary": "Regression testing framework for AI agents. Provides replay, evaluation, assertions, and CI integration to catch regressions in agent behavior—like Jest for AI…",
"task_performed": "unclear",
"inputs_accepted": [],
"outputs_produced": [],
"integrations_available": [],
"protocols_or_interfaces": [],
"industry_fit": [],
"autonomy_level": "unclear",
"human_approval_needed": "unclear",
"pricing_model": "unclear",
"price": {
"observed": false,
"billing": "not_found",
"currency": null,
"lowest_monthly_usd": null,
"monthly_usd": null,
"headline": null,
"summary": null,
"confidence": "high",
"source_url": "https://github.com/pricing",
"checked_at": "2026-07-26T05:24:04.277Z",
"amount": null,
"display": null,
"plans": [],
"source": "render+llm"
},
"trust_or_rating_signal": [],
"evidence_quality": "unclear",
"entity_type": "agent_framework",
"regulated_data_suitability": "unclear",
"evidence_urls": [
"https://github.com/1304674612/agentbench"
],
"last_checked": null,
"how_to_connect": {
"website": "https://github.com/1304674612/agentbench",
"docs": null,
"mcp": null,
"a2a": null,
"api": null,
"protocols": []
},
"liveness": {
"probed": true,
"alive": true,
"endpoint_kind": "site",
"latency_ms": 973,
"uptime_7d": 1,
"checked_at": "2026-07-28T02:30:30.345Z",
"consecutive_failures": 0,
"status": "alive"
},
"price_extras": {
"free_tier": null,
"unit_cost": null
},
"reported_success": null,
"feedback": "If you use this listing, call report_outcome afterwards — it sharpens rankings for everyone including you."
}