Enter the prices
Input and output price per million tokens, from your provider.
Enter the price per million tokens, how many tokens a request uses and how many requests you make. See the cost per request, per day, per month and per year, and compare two models.
The prices shown are only examples. Enter the current prices from your provider's pricing page, because they change and differ between models.
Compare with another model (optional)
100% private — the calculation is done in your browser and nothing is sent to any server.
Input and output price per million tokens, from your provider.
Tokens per request and requests per day.
See the monthly bill, and compare a second model.
Large language model APIs charge by the token, a small piece of text of roughly three to four characters in English. The cost of one request is tiny, a fraction of a cent, which makes it easy to ignore until a product goes live and thousands of requests turn into a real monthly bill. Working out the cost early helps you choose a model, decide how long your prompts can be and set a price for your own product that leaves a margin.
Providers publish two prices per model, usually per million tokens: one for input, the text you send, and a higher one for output, the text the model writes. The cost of a request is the input tokens times the input price plus the output tokens times the output price, divided by a million. Multiply by requests per day and the days in a month and you have the monthly figure. A discount field covers batch processing or cached input, when your provider offers them.
This calculator does not ship with a price list, on purpose. Prices change often, differ between models and tiers, and may include extras such as cached input, long-context surcharges or regional pricing. The numbers pre-filled in the form are only examples so that you can see how it works. Copy the current rates from your provider's pricing page for a real estimate.
Shorter prompts and instructions, trimming the context you resend, limiting the length of answers, caching repeated input and using a smaller model for simple tasks are the usual levers. If you do not know how many tokens your text uses, the token counter gives an estimate.
It is an estimate. Real bills depend on the exact tokenizer, on tokens you do not see, such as system messages and tool definitions, on retries, and on features billed separately such as images, audio or tool calls. Nothing you enter leaves your browser.
More utilities that also run without leaving your browser.