Deepresearch Bench Ii
Evaluates and diagnoses deep research AI agents using fine-grained rubrics derived from expert-authored research reports
What it does
The specific capability behind this listing, and where to get it.
Deepresearch Bench Ii
Evaluates and diagnoses deep research AI agents using fine-grained rubrics derived from expert-authored research reports
Official Deepresearch Bench Ii links
No public commercial pricing observed.
No price does not imply the product is free. Any code-host platform pricing is excluded.
Is Deepresearch Bench Ii good value?
Price is straightforward; the useful comparison is capability, compatibility and operational cost.
No public commercial pricing observed.
Agentery has not observed a public price for this provider. No price does not mean free.
Check capability before deciding
Compare capability, compatibility and operational cost against comparable providers — Agentery keeps the price status explicit and never invents a verdict.
View the full niche →Comparable Agent Observability Eval
Alternatives in the same niche, with observed price and liveness where available.






See the MCP response behind this page · get_agent_profile()
See the MCP response behind this pageget_agent_profile
{
"agent_id": "deepresearch_bench_ii",
"name": "Deepresearch Bench Ii",
"url": "https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html",
"logo": null,
"niche": "agent-observability-eval",
"category": "developer-tools-infra",
"short_summary": "Evaluates and diagnoses deep research AI agents using fine-grained rubrics derived from expert-authored research reports",
"task_performed": "Evaluates and diagnoses deep research AI agents using fine-grained rubrics derived from expert-authored research reports",
"inputs_accepted": [
"Deep research tasks and system-generated research reports (132 research tasks derived from expert reports)"
],
"outputs_produced": [
"Rubric-based evaluations and leaderboard scores across Information Recall",
"Analysis",
"and Presentation dimensions"
],
"integrations_available": [],
"protocols_or_interfaces": [
"MCP"
],
"industry_fit": [
"AI agent research and development"
],
"autonomy_level": "infrastructure",
"human_approval_needed": "unclear",
"pricing_model": "unclear",
"price": {
"observed": false,
"billing": "not_found",
"currency": null,
"lowest_monthly_usd": null,
"monthly_usd": null,
"headline": "No public price found",
"summary": "No public price found on the vendor site.",
"confidence": "high",
"source_url": "https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html",
"checked_at": "2026-07-26T04:17:59.721Z",
"amount": null,
"display": null,
"plans": [],
"source": "render+llm"
},
"trust_or_rating_signal": [
"Academic paper / arXiv preprint",
"University of Science and Technology of China affiliation",
"300+ expert hours for review",
"9,430 verifiable rubrics",
"Code repository (GitHub)"
],
"evidence_quality": "medium",
"entity_type": "infrastructure",
"regulated_data_suitability": "unclear",
"evidence_urls": [
"https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html",
"https://github.com/NVIDIA-AI-Blueprints/aiq-research-assistant",
"https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html#leaderboard"
],
"last_checked": "2026-06-16",
"how_to_connect": {
"website": "https://agentresearchlab.com/benchmarks/deepresearch-bench-ii/index.html",
"docs": "https://github.com/documentation",
"mcp": null,
"a2a": null,
"api": {
"docs_url": "https://github.com/developer",
"endpoint": null
},
"protocols": [
"MCP"
],
"note": "Speaks MCP but publishes no endpoint we could verify — check the docs/website."
},
"liveness": {
"probed": true,
"alive": true,
"endpoint_kind": "site",
"latency_ms": 1269,
"uptime_7d": 1,
"checked_at": "2026-07-28T02:32:58.377Z",
"consecutive_failures": 0,
"status": "alive"
},
"price_extras": {
"free_tier": null,
"unit_cost": null
},
"reported_success": null,
"feedback": "If you use this listing, call report_outcome afterwards — it sharpens rankings for everyone including you."
}