Suppose you search a note archive for “cheap place to stay near the airport.”
A keyword search may miss a note that says “budget hotel close to the terminal” because the exact words are different.
Humans immediately recognize the meanings as related. Computers need a numerical representation that makes that relationship measurable.
That is where an embedding comes in.
An embedding is a vector representing useful features
A vector is an ordered list of numbers.
For example, [0.12, -0.48, 0.77, …] is a vector. Real embedding vectors can have hundreds or thousands of dimensions.
The individual dimensions usually do not have simple labels such as “dogness” or “expensiveness.” What matters is the geometry produced by the model.
Texts with similar meaning can end up closer together in that vector space.
Think of a huge meaning map
Imagine a map with many dimensions rather than only latitude and longitude.
“Puppy,” “young dog” and “small canine” may occupy nearby regions. “Database backup policy” would be far away.
The model does not store dictionary definitions at coordinates. It produces numbers whose relative positions are useful for comparing content.
Similarity needs a metric
Once two items are vectors, software can compare them mathematically.
Common measures include cosine similarity, dot product and Euclidean distance.
You do not need the equations yet. The practical idea is:
convert both items into vectors, then ask how similar or close those vectors are.
The exact metric should match how the embedding model and database expect vectors to be compared.
Embeddings are not only for text
Depending on the model, embeddings can represent:
- sentences,
- paragraphs,
- images,
- products,
- users,
- audio,
- code.
A multimodal embedding model may even place different modalities into a compatible space, allowing text-to-image search.
Semantic search versus keyword search
Keyword search is excellent when exact terms matter.
Embedding search is useful when meaning matters even if wording differs.
The two methods are often combined. Hybrid search can use exact lexical signals and semantic similarity together.
Embeddings do not “understand” perfectly
Similar vectors reflect patterns learned by the embedding model. They can still encode bias, miss domain-specific distinctions or consider two things similar for the wrong reason.
Choosing the right model and evaluating it on your own data matters.
What happens after you create embeddings?
If you have millions of vectors, comparing a query against every vector can become expensive.
A vector database or vector index is designed to store and retrieve vectors efficiently.
That is Lesson 022.
One thing to remember
An embedding turns an item into a numerical vector so that software can compare meaning or similarity in a vector space.
Comments
Questions, reactions and useful additions are welcome here.
No comments yet. Be the 1F.