AI Serving Cost 0.1.0

An AI API bill is a quote you reverse-engineer after the fact. Ai-serving-cost does the reverse: it restates a provider's rates as one number, cost per request, then multiplies it out to the monthly bill at your traffic. The math follows the bill. Input tokens price at their per-million rate with the cache discount applied to the cached share, output tokens price at theirs, and the per-request fee sits on top. At 10,000 input and 500 output tokens per request, $3 and $15 per million rates, a 30% cached share, and 20,000 requests a day, it prints $0.029 per request, a $17,400 monthly bill, and $5,400 a month in input-cache savings. --compare scores two cache shares head to head; the sample run shows 60% caching saving $0.009 per request over 30%. --json emits machine output for pipelines. It treats rates as flat inputs: tiered pricing, batch windows, and context re-send costs stay out of scope. One Python file, standard library only, no network calls, MIT licensed. Runs on Python 3.8+ with nothing to install.

Tags python calculator api costs ai
License MITL
State stable

Recent Releases

0.1.011 Oct 2026 04:07 minor feature: Initial release: per-token and per-request serving math with cache discount handling, monthly bill projection, cache share comparison, JSON output.