LLM gateways pricing compared
One API for several model providers, with routing and logging. Check subscriptions, credit fees and model markups.
7 services · sources last checked 2026-10-06 · how we compare
What the most-compared services cost
Estimated yearly bill at $10,000 model spend / month. Enter your own usage below.
Enter your usage to estimate each bill. Services differ in what they include; check the rates, assumptions and extra charges below.
The estimates below use the default usage. Choose your current provider to compare costs.
| Provider | Parent jurisdiction | Features / pricing source | Est. monthly | Est. yearly |
|---|---|---|---|---|
| Vercel AI GatewayVercel-hosted unified API to many LLM providers with prepaid credits at provider list price, budgets and fallbacks. | 🇺🇸 United States | GDPR DPAPublished rates | $0no markup | $0 |
| PortkeyAI gateway with observability, guardrails and prompt management; open-source gateway plus hosted plans billed by logged requests. | 🇺🇸 United States | open sourceModeled ratesShows the base Production fee only; logged-request overages are not modeled by the model-spend input. | $49no markup | $588 |
| HeliconeOpen-source LLM observability platform and AI gateway with request logging, caching and routing; US or EU cloud regions. | 🇺🇸 United States | EU residencyopen sourceModeled ratesShows the subscription fee; request, storage and other usage-dependent extras are not modeled by the model-spend input. | $79no markup | $948 |
| Cloudflare AI GatewayCloudflare-hosted proxy for LLM APIs with analytics, caching, rate limiting, logging and optional unified billing. | 🇺🇸 United States | GDPR DPAModeled ratesUnified Billing fee only; BYOK avoids it. Log storage and Logpush overages are excluded. | $5005% fee | $6,000 |
| RequestyHosted LLM router with a unified API to 600+ models, fallbacks, caching and spend tracking; EU gateway in Frankfurt. | 🇬🇧 United Kingdom | EU residencyGDPR DPAPublished rates | $5005% fee | $6,000 |
| OpenRouterHosted unified API to hundreds of LLMs from many providers, billed via prepaid credits with a fee on credit purchases. | 🇺🇸 United States | EU plan availableGDPR DPAPublished rates | $550Standard credits · 5.5% fee | $6,600 |
Excludes the model spend itself. Self-hosted options cost the compute you run them on.
Need EU residency? Turn on the EU filter to pick a plan that includes it, which may cost more. “EU plan available” alone does not mean the price shown includes residency. Shared links save your inputs and filters, but prices can change.
Self-hosted options
Operating cost not estimated · servers, licenses and engineering time are yours. We list these separately because there’s no calculated bill to compare with managed services.
- LiteLLM — MIT-licensed, self-hosted proxy exposing an OpenAI-compatible API in front of 100+ LLM providers, by BerriAI.
No self-hosted options match these filters.
Head-to-head comparisons
LLM gateways pricing explained
What drives the bill
Gateways sit between your app and the model APIs and give you one API for many models, fallbacks, caching, logging and spend limits. They charge in one of three ways: a percentage of your model spend, a flat subscription, or nothing (self-hosted open source, or free gateways from platforms that make money elsewhere). A percentage fee looks small but grows with your spend. At $10,000 a month of model usage, 5% is $6,000 a year.
Costs the calculator leaves out
- Fees on buying credits, or on using your own provider keys (BYOK), which some gateways charge even when the per-token markup is zero.
- Log retention and analytics, often priced by request volume.
- Latency. Each hop adds a little, which matters for real-time voice and chat.
- The cost of running a self-hosted gateway such as LiteLLM: a small server and someone to update it.
When paying more makes sense
A single account and bill for several model providers can be useful for prototypes and small teams. As model spend grows, compare percentage fees with flat subscriptions and self-hosting. Include logging, routing features and the work of running a proxy in that comparison.