You hand a store assistant the same cream dress on two different days and ask, “Style this for work.”
On Monday she pairs it with a berry-red bag and nude heels.
On Friday she chooses dark loafers and a black tote.
Both outfits can make sense.
Generative artificial intelligence works a little like that.
The model does not begin with one complete hidden answer
A common mental model is: question goes in, the system looks up the matching answer, and the answer comes back out.
A large language model (LLM) usually works differently. It predicts a plausible next token from the text it has so far, then predicts the next one again, and continues step by step.
After “Tonight I want to eat…,” several continuations might be reasonable: pasta, hot pot, something light, or Japanese food nearby.
The model gives different possibilities different probabilities. Once one path is chosen, that new text becomes part of the context for the next step.
An entire answer is therefore a long chain of small choices.
Open questions have more possible paths
“Write a birthday message” can be warm, funny, restrained or sentimental.
“Plan a weekend trip” can begin with transport, hotels or places to visit.
There may be many valid routes, so different does not automatically mean wrong.
The useful questions are: Did it follow your constraints? Are the facts correct? Is the tone what you wanted?
Where does variation come from?
One common generation method is sampling: choosing from several plausible next tokens, with more likely ones usually receiving a higher chance of selection.
Think of five pairs of shoes that all work with the dress. The assistant is not choosing completely at random, but she is not forced to pick the exact same pair every time either.
Some model APIs expose a setting called temperature. At a high level, lower temperature tends to favor more predictable high-probability choices; higher temperature allows more unusual choices to appear more often.
Consumer chat apps do not always show this control, but generation can still contain variation behind the interface.
Why are factual questions often more stable?
If the space of reasonable answers is narrow, repeated outputs are likely to look similar.
Ask, “Is Tokyo the capital of Japan?” and there is little room for creative variation.
Ask, “Give me three autumn date ideas,” and the possibility space is enormous.
A useful rule of thumb is: the more open the task, the more the answer can branch.
How do you make answers more consistent?
Reduce the possibility space.
Instead of:
Write a nice birthday message.
Try:
Write a birthday message under 80 words for a female friend I have known for ten years. Keep it warm but not sentimental, use no emoji, and end with the idea that we should travel together again.
That is like telling the store assistant: flat shoes, black, office-safe, closed toe, under a fixed budget.
Clear constraints and fixed output formats usually make repeated runs more similar.
Variation can also be useful
For headlines, outfits, brainstorming and creative planning, you often do not want the same answer every time.
You can deliberately ask:
Give me five directions that feel genuinely different from one another.
Then variability becomes a feature.
So when the same question produces a different answer, you do not need to imagine the AI changed its mood.
A better model is simpler:
It is generating a path through many possible next steps.
The same dress can support more than one good outfit.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.