
Wafer
provides fast serverless and dedicated inference for open-source LLMs
What it does
The specific capability behind this listing, and where to get it.
Wafer
provides fast serverless and dedicated inference for open-source LLMs
Official Wafer links
Pricing & plans
Observed public pricing for Wafer, benchmarked against comparable providers. Plans, tiers, history and scenario below.
$1 /token
This plan sits close to the market.
$1 is positioned against a $1–$2 observed middle market.
Why Agentery reaches that view
price_benchmark0% above the median.
get_agent_profile · plan historyper token observed 2026-07-28.
pricing recommendationPercentile unavailable for this basis.
confidenceSource page rechecked daily.
Test a different price for this plan.
Move the proposed monthly price. Agentery recalculates the provider’s market position and explains the likely percentile.
Is Wafer good value?
How its price compares with genuinely comparable providers.
Priced near the market for its buyer tier.
Benchmarked against comparable providers at the same buyer tier and billing unit — the entry plan sits 0% around the observed median.
Compared with LLM Serving Infrastructure
Positioned against the observed p25 / median / p75 of comparable providers at the same buyer tier and billing unit. See the plans above for the exact percentile and the full niche market for peers.
View the full niche →Comparable LLM Serving Infrastructure
Alternatives in the same niche, with observed price and liveness where available.



See the MCP response behind this page · get_agent_profile()
See the MCP response behind this pageget_agent_profile
{
"agent_id": "wafer",
"name": "Wafer",
"url": "https://www.wafer.ai/",
"logo": "https://agentery.com/logos/CP-68DKTK-256.png",
"niche": "llm-serving-infrastructure",
"category": "developer-tools-infra",
"short_summary": "provides fast serverless and dedicated inference for open-source LLMs",
"task_performed": "provides fast serverless and dedicated inference for open-source LLMs",
"inputs_accepted": [
"API calls (OpenAI Chat Completions schema)",
"prompts",
"tool use",
"JSON mode"
],
"outputs_produced": [
"LLM completions/responses via streaming API"
],
"integrations_available": [
"OpenAI SDK",
"LangChain",
"LiteLLM",
"Claude Code",
"Cline",
"OpenAI Chat Completions API"
],
"protocols_or_interfaces": [
"SDK",
"API"
],
"industry_fit": [
"developer tools"
],
"autonomy_level": "infrastructure",
"human_approval_needed": "unclear",
"pricing_model": "usage-based",
"price": {
"observed": true,
"billing": "usage",
"currency": "USD",
"lowest_monthly_usd": null,
"monthly_usd": null,
"headline": "Paid (price not published)",
"summary": "Wafer offers usage-based pricing for serverless LLM inference, charged per million tokens with separate input, output, and cache rates.",
"confidence": "medium",
"source_url": "https://www.wafer.ai/",
"checked_at": "2026-07-28T05:45:27.254Z",
"amount": null,
"display": null,
"plans": [
{
"name": "GLM-5.2-Fast",
"price": "Input: $3.00, Output: $10.25, Cache: $0.50 per M tokens",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Low-latency inference",
"Per-stream throughput SLA"
]
},
{
"name": "GLM-5.2",
"price": "Input: $1.20, Output: $4.10, Cache: $0.20 per M tokens",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"Flagship model",
"Strong coding and reasoning"
]
},
{
"name": "GLM-5.1",
"price": "Input: $1.00, Output: $3.20, Cache: $0.10 per M tokens",
"usage": true,
"period": "usage",
"persona": "individual",
"highlights": [
"General Language Model",
"Strong coding and reasoning"
]
}
],
"source": "render+llm"
},
"trust_or_rating_signal": [
"case studies (Evergrove Labs)",
"public benchmarks",
"SLA-backed uptime",
"DPA available",
"zero data retention available"
],
"evidence_quality": "high",
"entity_type": "infrastructure",
"regulated_data_suitability": "DPA stated, zero data retention available for compliance-bound workloads",
"evidence_urls": [
"https://www.wafer.ai/"
],
"last_checked": "2026-06-16",
"how_to_connect": {
"website": "https://www.wafer.ai/",
"docs": null,
"mcp": null,
"a2a": null,
"api": null,
"protocols": []
},
"liveness": {
"probed": true,
"alive": true,
"endpoint_kind": "site",
"latency_ms": 138,
"uptime_7d": 1,
"checked_at": "2026-07-28T02:39:27.232Z",
"consecutive_failures": 0,
"status": "alive"
},
"price_extras": {
"free_tier": null,
"unit_cost": null
},
"reported_success": null,
"feedback": "If you use this listing, call report_outcome afterwards — it sharpens rankings for everyone including you."
}