
Clawbench
Open benchmark that evaluates AI browser agents on real, live websites by giving them everyday online tasks (booking flights, ordering food, shopping, applying…
What it does
The specific capability behind this listing, and where to get it.
Open benchmark that evaluates AI browser agents on real, live websites by giving them everyday online tasks (booking flights, ordering food, shopping, applying for jobs) and scoring whether they submit the correct request
Official Clawbench links
Pricing & plans
Observed public pricing for Clawbench, benchmarked against comparable providers. Plans, tiers, history and scenario below.
$0.07 /
The current price is deliberately competitive.
$0.07 is positioned against a $25–$100 observed middle market.
Why Agentery reaches that view
price_benchmark99.9% below the median.
get_agent_profile · plan historyper observed 2026-07-28.
pricing recommendationPercentile unavailable for this basis.
confidenceSource page rechecked daily.
Test a different price for this plan.
Move the proposed monthly price. Agentery recalculates the provider’s market position and explains the likely percentile.
Is Clawbench good value?
How its price compares with genuinely comparable providers.
Not enough evidence yet to call it good — or poor — value.
Too few comparable providers at the same buyer tier and billing unit to benchmark this price honestly.
Check capability before deciding
Compare capability, compatibility and operational cost against the few observed peers — Agentery keeps the price status explicit and never invents a verdict.
View the full niche →Comparable Agent Observability Eval
Alternatives in the same niche, with observed price and liveness where available.






See the MCP response behind this page · get_agent_profile()
See the MCP response behind this pageget_agent_profile
{
"agent_id": "clawbench",
"name": "Clawbench",
"url": "https://claw-bench.com/",
"logo": "https://agentery.com/logos/CP-W56MMH-256.png",
"niche": "agent-observability-eval",
"category": "developer-tools-infra",
"short_summary": "Open benchmark that evaluates AI browser agents on real, live websites by giving them everyday online tasks (booking flights, ordering food, shopping, applying…",
"task_performed": "Open benchmark that evaluates AI browser agents on real, live websites by giving them everyday online tasks (booking flights, ordering food, shopping, applying for jobs) and scoring whether they submit the correct request",
"inputs_accepted": [
"Task prompts/instructions",
"agent model trajectories",
"execution traces (recordings",
"actions",
"agent messages",
"HTTP requests)",
"benchmark corpora bundles"
],
"outputs_produced": [
"Leaderboard rankings",
"two-stage scores (HTTP interception + LLM judge)",
"pass/fail verdicts",
"5-layer execution traces",
"datasets of model trajectories"
],
"integrations_available": [
"GitHub",
"Hugging Face",
"PyPI",
"OpenRouter",
"arXiv",
"Gradio"
],
"protocols_or_interfaces": [],
"industry_fit": [
"developer tools / AI agent evaluation research"
],
"autonomy_level": "infrastructure",
"human_approval_needed": "unclear",
"pricing_model": "free",
"price": {
"observed": true,
"billing": "usage",
"currency": "USD",
"lowest_monthly_usd": null,
"monthly_usd": null,
"headline": "Paid (price not published)",
"summary": "The homepage reports model-specific usage rates ranging from $0.0000 to $4.4425 per task.",
"confidence": "medium",
"source_url": "https://claw-bench.com/",
"checked_at": "2026-07-28T04:51:28.700Z",
"amount": null,
"display": null,
"plans": [
{
"name": "claude-opus-4-7",
"price": "$4.4425",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Cost per task"
]
},
{
"name": "gpt-5.5",
"price": "$0.3325",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Cost per task"
]
},
{
"name": "glm-5.1",
"price": "$0.1935",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Cost per task"
]
},
{
"name": "deepseek-v4-pro",
"price": "$0.0721",
"usage": true,
"period": "usage",
"persona": "pro",
"highlights": [
"Cost per task"
]
},
{
"name": "deepseek-v4-flash:free",
"price": "$0.0000",
"usage": true,
"period": "usage",
"persona": "free",
"highlights": [
"Cost per task"
]
},
{
"name": "z-ai/glm-4.5-air:free",
"price": "$0.0000",
"usage": true,
"period": "usage",
"persona": "free",
"highlights": [
"Cost per task"
]
},
{
"name": "minimax-m2.5:free",
"price": "$0.0000",
"usage": true,
"period": "usage",
"persona": "free",
"highlights": [
"Cost per task"
]
},
{
"name": "openrouter-owl-alpha",
"price": "$0.3704",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Cost per task"
]
}
],
"source": "render+llm"
},
"trust_or_rating_signal": [
"GitHub 286 stars",
"arXiv paper (2604.08523)",
"#3 HuggingFace Paper of the Day",
"TIGER-AI-Lab / TIGER-Lab affiliation",
"featured in DeepWiki, awesome-harness-engineering, Awesome-AI-Agents, LLM-Agent-Benchmark-List",
"Apache-2.0 license",
"1,724 judge-verified runs"
],
"evidence_quality": "high",
"entity_type": "tool",
"regulated_data_suitability": "unclear",
"evidence_urls": [
"https://claw-bench.com/",
"https://github.com/reacher-z/ClawBench"
],
"last_checked": "2026-06-16",
"how_to_connect": {
"website": "https://claw-bench.com/",
"docs": "https://github.com/documentation",
"mcp": null,
"a2a": null,
"api": {
"docs_url": "https://github.com/developer",
"endpoint": null
},
"protocols": []
},
"liveness": {
"probed": true,
"alive": true,
"endpoint_kind": "site",
"latency_ms": 765,
"uptime_7d": 1,
"checked_at": "2026-07-28T02:32:09.480Z",
"consecutive_failures": 0,
"status": "alive"
},
"price_extras": {
"free_tier": null,
"unit_cost": null
},
"reported_success": null,
"feedback": "If you use this listing, call report_outcome afterwards — it sharpens rankings for everyone including you."
}