← Back home
LESSON 019AI Development10 min

How does AI API pricing work? Read input, output and cached tokens before estimating cost

AI API prices are usually usage-based. Lesson 019 explains per-million-token units, input versus output pricing, cached input, worked examples and scaling.

Today’s analogya utility bill where different kinds of usage can have different rates instead of one flat price per request

An API pricing table might say “US dollars per 1 million tokens.” That can look intimidating until you translate it into the cost of one real request.

The first rule is: do not treat “one request” as a fixed unit of cost.

A request containing a short sentence and a request containing a 100-page document are both one request, but their usage is completely different.

What does “per 1M tokens” mean?

“1M” means one million.

If a provider lists a price of US$2 per 1 million input tokens, that means 1,000,000 input tokens cost US$2 at that listed rate.

The word US$ matters: it means United States dollars, not necessarily your local currency.

Prices also change, so always check the provider’s current official pricing page before making a budget decision. The source links attached to this lesson are first-party references for that reason.

Input and output are often different rates

Text-model billing commonly separates:

Output can cost more than input because generation requires sequential computation. The exact ratio depends on the model and provider.

A worked example with hypothetical prices

The following numbers are an example only, not a current quote from any provider.

Suppose a model costs:

Your request uses 5,000 input tokens and produces 1,000 output tokens.

Input cost:

5,000 / 1,000,000 × US$2 = US$0.01

Output cost:

1,000 / 1,000,000 × US$8 = US$0.008

Total for that request:

US$0.018

Again, this is only a calculation example. Replace the rates with the current official prices for the model you actually use.

Small costs become large at scale

US$0.018 sounds tiny.

But 100,000 similar requests would be approximately:

100,000 × US$0.018 = US$1,800

That is why developers monitor average input size, average output size and request volume rather than only looking at the price table once.

Context can quietly become the expensive part

If every request resends a long conversation history, document set or system prompt, input usage grows.

For an agent or RAG application, the model may be called several times for one user action. The cost visible to the user as “one click” may actually contain multiple API calls.

Caching is not “free memory”

Some providers discount eligible repeated input through prompt or context caching. That can be useful when the same large prefix is sent repeatedly.

But caching has provider-specific rules: minimum sizes, retention windows, supported models and cache-hit conditions can differ.

Always read the current documentation rather than assuming every repeated token receives a discount.

Budget with a formula

A simple estimate is:

cost ≈ input tokens × input rate + output tokens × output rate + other metered features

When rates are quoted per million tokens, divide token counts by 1,000,000 before multiplying.

Image generation, audio, video, tool calls or storage may use additional billing units, so a text-token calculation may not cover the entire product.

One thing to remember

AI API cost is usually a usage equation, not a fixed price per request. Read the unit, separate input and output, then multiply by real traffic.

Lesson 020 introduces AI agents, where one user request can trigger several model calls and tools—making this cost model even more important.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. OpenAI — API Pricing ↗
  2. Google — Gemini Developer API Pricing ↗
  3. xAI — API Pricing ↗
← Previous018What is an API? Think of a restaurant order window between your app and an AI model
Next →020What is an AI agent? A model that can decide steps and use tools toward a goal
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200