← Back home
LESSON 025AI Development10 min

What is a local LLM? Run the model on your own machine instead of calling a hosted API

A local LLM runs on hardware you control. Lesson 025 explains weights, VRAM, quantization, privacy, performance and the trade-offs versus hosted APIs.

Today’s analogycooking in your own kitchen instead of ordering every meal from a restaurant

When you call an AI API, your application sends data to a provider’s infrastructure and the model runs there.

A local Large Language Model (local LLM) runs on hardware you control: a laptop, desktop, workstation or server.

The difference is similar to ordering food versus cooking in your own kitchen. More control also means more responsibility.

You need the model weights

To run a model locally, you need access to its model weights—the learned numerical parameters produced by training.

Models whose weights are downloadable are often called open-weight models. This is not exactly the same as “open source,” because training data, code or license terms may still be restricted.

Always read the model license before commercial use.

Memory is often the first hardware limit

Large models contain billions of parameters. Those parameters need memory.

For GPU inference, VRAM (Video Random Access Memory) is often a key constraint.

A rough idea: storing 8 billion parameters at 16 bits would require about 16 GB just for the weights, before considering runtime overhead, caches and other memory.

That is why local inference often uses quantization.

What is quantization?

Quantization stores or computes model values with fewer bits.

For example, a model may be converted from 16-bit weights to 8-bit or 4-bit representations.

This can greatly reduce memory use and sometimes improve speed, but it may reduce accuracy or quality depending on the method and model.

“4-bit” is not automatically bad; the only reliable answer is to evaluate the specific quantized model on your tasks.

Why run locally?

Common reasons include:

Local does not automatically mean private or free

If your local app sends telemetry, calls external search services or uploads prompts elsewhere, data can still leave the machine.

And although there may be no API bill, hardware, electricity, maintenance and engineering time still have costs.

Performance depends on the whole stack

Tokens per second can depend on:

A smaller well-optimized model can feel faster and more useful than a larger model that barely fits in memory.

Hosted API or local model?

Hosted APIs are convenient when you want strong models without maintaining hardware.

Local models are attractive when control, privacy, offline operation or predictable infrastructure matters.

Many systems use both.

One thing to remember

A local LLM is a language model whose weights are loaded and executed on hardware you control, trading cloud convenience for more control and operational responsibility.

Lesson 026 introduces the language you will see everywhere in AI experiments and automation: Python.

Primary sources

Analogies build intuition; use the original sources for formal definitions and technical detail.

  1. Ollama — Documentation ↗
  2. llama.cpp — LLM inference in C/C++ ↗
← Previous024What is fine-tuning? Continue training a model so a behavior becomes part of the model itself
Next →026Why is Python used so much in AI? It is the glue between models, data and experiments
COMMUNITY

Comments

Questions, reactions and useful additions are welcome here.

0 / 1200