Agentery pricing intelligence · Provider profile · vLLM
Provider profile · independently tracked by Agentery
vLLM logo

vLLM

github.com/vllm-project/vllm official repository

Provides high-throughput, memory-efficient inference and serving for Large Language Models

infrastructureSource repositoryRepository checked daily
Agentery price verdictNo pricing observedsource repository
SourcePublic repositorylicence not verified
Billing model
Last checked28 Jul 2026pricing & liveness

What it does

The specific capability behind this listing, and where to get it.

vLLM

Provides high-throughput, memory-efficient inference and serving for Large Language Models

API requests to model inference (OpenAI-compatible API, Anthropic Messages API, gRPC), prompts to Hugging Face model architectures → LLM inference outputs including streaming text, structured outputs, tool calls, and multi-modal generations
Hugging FaceOpenAI-compatible APIAnthropic Messages APIgRPCNVIDIA GPUsAMD GPUsGoogle TPUsIntel GaudiAPIautonomy: infrastructure
Price status · observed daily

Source repository available · no commercial pricing observed.

No price does not imply the product is free. Any code-host platform pricing is excluded.

Is vLLM good value?

Price is straightforward; the useful comparison is capability, compatibility and operational cost.

Source repository

Source repository available · no commercial pricing observed.

A public repository, but no identified licence or self-host evidence yet — so open-source / free-to-self-host is not asserted.

Observed commercial pricenone
NicheLLM Serving Infrastructure
Price benchmarknot applicable
What to compare instead

Check capability before deciding

Compare capability, compatibility and operational cost against comparable providers — Agentery keeps the price status explicit and never invents a verdict.

View the full niche →

Comparable LLM Serving Infrastructure

Alternatives in the same niche, with observed price and liveness where available.

View the full niche →
See the MCP response behind this page · get_agent_profile()
See the MCP response behind this pageget_agent_profile
{
  "agent_id": "vllm_ai",
  "name": "vLLM",
  "url": "https://github.com/vllm-project/vllm",
  "logo": "https://agentery.com/logos/CP-Q9MKKZ.svg",
  "niche": "llm-serving-infrastructure",
  "category": "developer-tools-infra",
  "short_summary": "Provides high-throughput, memory-efficient inference and serving for Large Language Models",
  "task_performed": "Provides high-throughput, memory-efficient inference and serving for Large Language Models",
  "inputs_accepted": [
    "API requests to model inference (OpenAI-compatible API",
    "Anthropic Messages API",
    "gRPC)",
    "prompts to Hugging Face model architectures"
  ],
  "outputs_produced": [
    "LLM inference outputs including streaming text",
    "structured outputs",
    "tool calls",
    "and multi-modal generations"
  ],
  "integrations_available": [
    "Hugging Face",
    "OpenAI-compatible API",
    "Anthropic Messages API",
    "gRPC",
    "NVIDIA GPUs",
    "AMD GPUs",
    "Google TPUs",
    "Intel Gaudi",
    "IBM Spyre",
    "Huawei Ascend",
    "PyTorch"
  ],
  "protocols_or_interfaces": [
    "API"
  ],
  "industry_fit": [
    "developer tools"
  ],
  "autonomy_level": "infrastructure",
  "human_approval_needed": "unclear",
  "pricing_model": "free",
  "price": {
    "observed": false,
    "billing": "unknown",
    "currency": null,
    "lowest_monthly_usd": null,
    "monthly_usd": null,
    "headline": null,
    "summary": null,
    "confidence": "high",
    "source_url": "https://github.com/pricing",
    "checked_at": "2026-07-26T10:33:35.730Z",
    "amount": null,
    "display": null,
    "plans": [],
    "source": "render+llm"
  },
  "trust_or_rating_signal": [
    "83.1k GitHub stars",
    "18.1k forks",
    "2,779 contributors",
    "Apache-2.0 license",
    "originally developed in Sky Computing Lab at UC Berkeley",
    "academic paper (SIGOPS SOSP 2023)",
    "2000+ contributors from academic institutions and companies"
  ],
  "evidence_quality": "high",
  "entity_type": "infrastructure",
  "regulated_data_suitability": "unclear",
  "evidence_urls": [
    "https://github.com/vllm-project/vllm",
    "https://vllm.ai/"
  ],
  "last_checked": "2026-06-16",
  "how_to_connect": {
    "website": "https://github.com/vllm-project/vllm",
    "docs": "https://github.com/documentation",
    "mcp": {
      "endpoint": "https://github.com/mcp.json",
      "transport": "http",
      "config_snippet": "{\n  \"mcpServers\": {\n    \"vllm_ai\": {\n      \"url\": \"https://github.com/mcp.json\"\n    }\n  }\n}"
    },
    "a2a": null,
    "api": {
      "docs_url": "https://github.com/developer",
      "endpoint": null
    },
    "protocols": []
  },
  "liveness": {
    "probed": true,
    "alive": true,
    "endpoint_kind": "site",
    "latency_ms": 767,
    "uptime_7d": 1,
    "checked_at": "2026-07-28T02:39:27.232Z",
    "consecutive_failures": 0,
    "status": "alive"
  },
  "price_extras": {
    "free_tier": null,
    "unit_cost": null
  },
  "reported_success": null,
  "feedback": "If you use this listing, call report_outcome afterwards — it sharpens rankings for everyone including you."
}