When you call an AI API, your application sends data to a provider’s infrastructure and the model runs there.
A local Large Language Model (local LLM) runs on hardware you control: a laptop, desktop, workstation or server.
The difference is similar to ordering food versus cooking in your own kitchen. More control also means more responsibility.
You need the model weights
To run a model locally, you need access to its model weights—the learned numerical parameters produced by training.
Models whose weights are downloadable are often called open-weight models. This is not exactly the same as “open source,” because training data, code or license terms may still be restricted.
Always read the model license before commercial use.
Memory is often the first hardware limit
Large models contain billions of parameters. Those parameters need memory.
For GPU inference, VRAM (Video Random Access Memory) is often a key constraint.
A rough idea: storing 8 billion parameters at 16 bits would require about 16 GB just for the weights, before considering runtime overhead, caches and other memory.
That is why local inference often uses quantization.
What is quantization?
Quantization stores or computes model values with fewer bits.
For example, a model may be converted from 16-bit weights to 8-bit or 4-bit representations.
This can greatly reduce memory use and sometimes improve speed, but it may reduce accuracy or quality depending on the method and model.
“4-bit” is not automatically bad; the only reliable answer is to evaluate the specific quantized model on your tasks.
Why run locally?
Common reasons include:
- keeping data inside your environment,
- working offline,
- avoiding per-request API fees,
- pinning a model version,
- experimenting with open-weight models,
- integrating tightly with local files or systems.
Local does not automatically mean private or free
If your local app sends telemetry, calls external search services or uploads prompts elsewhere, data can still leave the machine.
And although there may be no API bill, hardware, electricity, maintenance and engineering time still have costs.
Performance depends on the whole stack
Tokens per second can depend on:
- model size,
- quantization,
- GPU and VRAM bandwidth,
- CPU and system RAM,
- context length,
- inference engine,
- batch size.
A smaller well-optimized model can feel faster and more useful than a larger model that barely fits in memory.
Hosted API or local model?
Hosted APIs are convenient when you want strong models without maintaining hardware.
Local models are attractive when control, privacy, offline operation or predictable infrastructure matters.
Many systems use both.
One thing to remember
A local LLM is a language model whose weights are loaded and executed on hardware you control, trading cloud convenience for more control and operational responsibility.
Lesson 026 introduces the language you will see everywhere in AI experiments and automation: Python.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.