← Back home
LESSON 050AI Basics10 min

What are thinking and reasoning modes? The same AI can trade speed for more inference-time work

AI products often offer fast responses and higher reasoning effort. Learn how inference compute, latency, cost, tools and verification relate—and why more thinking is not always better.

Today’s analogyThe same person answers the time instantly but spends longer checking a proof or contract; reasoning effort is a larger problem-solving budget.

Ask an AI:

What is 7 × 8?

Spending a minute planning and re-checking the answer would add little value.

Ask instead:

Find the cause of this intermittent distributed-system deadlock.
Form three hypotheses, test each against the logs and propose the smallest fix.

A rushed first answer may be exactly what you do not want.

This is why modern AI products and APIs expose thinking / reasoning modes or reasoning effort controls.

Separate training from inference

Lesson 003 introduced training, the process that adjusts model parameters using data and optimization.

When you ask a trained model a question, it performs inference: it uses those learned parameters to produce this particular response.

Reasoning controls mainly change how much work the system is willing to do during inference for the current task.

Think of effort as a problem-solving budget

A person does not allocate the same amount of thought to every question.

“What time is it?”
→ answer directly

“Do these 80 pages of contract clauses contradict each other?”
→ read, take notes, cross-check, reconsider

Reasoning effort is a useful way to think about allocating a larger or smaller problem-solving budget.

Different providers implement labels such as low, high or thinking differently. Do not assume identical internal mechanisms simply because the names look similar.

What can more reasoning help with?

Complex tasks that may benefit include:

More inference-time work gives the system more opportunity to handle intermediate steps rather than emitting the first plausible response.

The trade-off is latency and compute

Latency is the delay between sending a request and receiving the result.

Heavier reasoning may increase latency and consume more computation. Depending on the provider and pricing model, that can also affect cost.

Therefore “always use the maximum setting” is not automatically the best product design.

Simple tasks may not benefit much

Translation of a short sentence, formatting, straightforward extraction and other highly constrained tasks may gain little from the heaviest reasoning mode.

Real systems balance:

Quality
Latency
Cost

Lesson 043 and Lesson 044 introduced external information retrieval.

Keep the concepts separate:

Reasoning
= work more deeply with available context

Web search
= retrieve new external information

Deep Research
= perform repeated search, reading and synthesis

If a model does not have today’s stock price, reasoning for three times longer will not magically create the correct live value.

More reasoning does not eliminate hallucination

The lesson from Lesson 006 still applies. A model may catch an error after more checking, but it can also develop an incorrect assumption into a more elaborate answer.

High-stakes work still needs evidence, tools, tests and human review.

Agents use reasoning to decide the next action

In a coding agent or computer-use workflow, reasoning is not only about the final paragraph. The model may need to decide which file to inspect, whether to run a test, which hypothesis matches an error and which tool to call next.

That is why complex agent tasks can benefit more from reasoning effort than a simple chat response.

A simple beginner rule

Simple, reversible, formatting-heavy task
→ fast mode

Multiple constraints or nontrivial deduction
→ medium/high reasoning

High-risk debugging, long workflows or complex research
→ high reasoning + tools + verification

Lesson 050: the map is getting connected

From Lesson 001 to Lesson 050, the ideas now form one larger system: models understand and generate; tools connect them to external systems; reasoning helps plan difficult work; search and research bring in evidence; and safety plus verification determine whether the resulting system is trustworthy enough for real use.

One thing to remember

Thinking or reasoning mode is best understood as changing the inference-time problem-solving budget for a task. More effort can help difficult work, but it costs time and compute and does not replace fresh information, tools, tests or verification.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. OpenAI API — Reasoning guide ↗
  2. OpenAI API — Models ↗
  3. Google AI — Thinking ↗
← Previous049What is vibe coding? Natural-language app building is fast until the prototype becomes a real system
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200