An API pricing table might say “US dollars per 1 million tokens.” That can look intimidating until you translate it into the cost of one real request.
The first rule is: do not treat “one request” as a fixed unit of cost.
A request containing a short sentence and a request containing a 100-page document are both one request, but their usage is completely different.
What does “per 1M tokens” mean?
“1M” means one million.
If a provider lists a price of US$2 per 1 million input tokens, that means 1,000,000 input tokens cost US$2 at that listed rate.
The word US$ matters: it means United States dollars, not necessarily your local currency.
Prices also change, so always check the provider’s current official pricing page before making a budget decision. The source links attached to this lesson are first-party references for that reason.
Input and output are often different rates
Text-model billing commonly separates:
- input tokens — the text/context you send to the model,
- output tokens — the text the model generates back,
- cached input tokens — reused input that some providers can process at a lower rate when caching rules are met.
Output can cost more than input because generation requires sequential computation. The exact ratio depends on the model and provider.
A worked example with hypothetical prices
The following numbers are an example only, not a current quote from any provider.
Suppose a model costs:
- US$2 per 1M input tokens,
- US$8 per 1M output tokens.
Your request uses 5,000 input tokens and produces 1,000 output tokens.
Input cost:
5,000 / 1,000,000 × US$2 = US$0.01
Output cost:
1,000 / 1,000,000 × US$8 = US$0.008
Total for that request:
US$0.018
Again, this is only a calculation example. Replace the rates with the current official prices for the model you actually use.
Small costs become large at scale
US$0.018 sounds tiny.
But 100,000 similar requests would be approximately:
100,000 × US$0.018 = US$1,800
That is why developers monitor average input size, average output size and request volume rather than only looking at the price table once.
Context can quietly become the expensive part
If every request resends a long conversation history, document set or system prompt, input usage grows.
For an agent or RAG application, the model may be called several times for one user action. The cost visible to the user as “one click” may actually contain multiple API calls.
Caching is not “free memory”
Some providers discount eligible repeated input through prompt or context caching. That can be useful when the same large prefix is sent repeatedly.
But caching has provider-specific rules: minimum sizes, retention windows, supported models and cache-hit conditions can differ.
Always read the current documentation rather than assuming every repeated token receives a discount.
Budget with a formula
A simple estimate is:
cost ≈ input tokens × input rate + output tokens × output rate + other metered features
When rates are quoted per million tokens, divide token counts by 1,000,000 before multiplying.
Image generation, audio, video, tool calls or storage may use additional billing units, so a text-token calculation may not cover the entire product.
One thing to remember
AI API cost is usually a usage equation, not a fixed price per request. Read the unit, separate input and output, then multiply by real traffic.
Lesson 020 introduces AI agents, where one user request can trigger several model calls and tools—making this cost model even more important.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.