Local LLM vs API break-even calculator
← All tools · Inference cost calculator · VRAM calculator
Cheaper to buy a GPU and run local, or pay per token? Plug in hardware cost, power draw, daily output volume, and an API price — get monthly costs and a break-even timeline.
Hardware
Electricity
Usage
Compute time: h/day ( h/mo at this volume).
API comparison
Output-token pricing only. Compare with the inference cost calculator for full input+output workloads. snapshot.
Local / month
API / month
output tokens/mo
Monthly savings vs API:
Show your work
Caveats: local open-weight models still trail frontier APIs on hard tasks. Your time maintaining drivers, quantizations, and crashes is a real cost this sheet ignores. API prices keep falling. Privacy, offline access, and zero per-token anxiety are non-dollar benefits that may dominate the decision — use the VRAM calculator to confirm the model actually fits your hardware first.
About this tool
This is a rough TCO sketch: amortized hardware plus electricity for the hours you are actually generating tokens, compared to a single output-token API rate. It does not model input tokens, batch APIs, or subscription plans.
Read more
Free books and courses
Get my programming books and courses as PDF and EPUB files.
Get the download library →