Technical resources
LA

LLM API cost calculator

active

Enter your token ratio, your server cost, your exchange rate: compare Claude, GPT-5 and Gemini against a self-hosted GPU server, live.

calculator · updated August 24, 2026

costllm-pricingclaude-apigpt-5geminihetznerself-hostingcalculatorbreakeven

The three published numbers on how many tokens per month before self-hosting an LLM is cheaper than the Claude API assume Claude pricing, a Hetzner GEX44 server, and one of three fixed input:output ratios. Real workloads rarely land exactly on one of those nine cells, and Claude is not the only API worth pricing against your own server. This calculator runs the identical formula with your own numbers, across five Claude tiers plus GPT-5 and Gemini.

What it computes

The same blended-price arithmetic as the source page: pick a model, set your actual input:output token ratio (pull it from your API usage dashboard, which reports the two separately), and enter the monthly cost of whatever self-hosted server you are actually pricing against, in EUR. The result is the monthly token volume, in millions, at which that fixed server cost breaks even against paying per token.

The EUR/USD field is pre-filled with the rate used on the source page but is editable: exchange rates move, and the number underneath yours should not be someone else's snapshot.

What it does not do

It does not account for the labor of running and patching a server, the quality gap between a frontier model and whatever open model your server actually runs, or a negotiated enterprise rate that differs from any vendor's public list price. It also does not cover open-weight models (Llama, Qwen, DeepSeek, Mistral...): those are usually the self-hosted side of this comparison, not a metered API to price against it. The source page spells out the Claude-specific caveats in full; read it before treating the number below as a purchase decision rather than a first estimate.

Sources

Claude API pricing: Anthropic's Claude Platform pricing docs. GPT-5 pricing: OpenAI's API pricing page. Gemini pricing: Google's Gemini API pricing page (standard tier, prompts under 200K tokens). Server specification used as the default: Hetzner's GPU dedicated server matrix. Exchange rate default: ECB euro reference exchange rate.

Run your own numbers

The three inputs below drive the same formula as the reproduction script on the source page: model tier, your input:output ratio, and the fixed monthly cost of your self-hosted alternative. The result updates as you type.

Breakeven volume

Enter your email to see the result.