Type "kitten" into a search box and a good one will also show you pages about cats. Yet to a computer, kitten and cat are just two different strings of characters, and kitten and kitchen look more alike than kitten and cat. How does software get from letters to meaning?

The trick: turn meaning into numbers you can measure

An embedding is a list of numbers that stands for a piece of text (a word, a sentence, even a whole paragraph). The list is built so that things with similar meaning get similar numbers. Once meaning is numbers, "are these two related?" becomes a question about distance, and computers are very good at measuring distance.

Think of a map. Two cafés on the same street have nearby coordinates, and you can tell they are close without knowing anything else about them. An embedding is coordinates on a map of meaning. The difference is that a map of a city needs two numbers (latitude and longitude), while real embeddings use many more: often hundreds to a few thousand numbers per piece of text (for example, one popular OpenAI model returns 1,536 of them). We can't draw 1,536 dimensions, so here is a tiny two-number version.

Try it: a map of 12 words

The map below is hand-made for this article: we placed each word ourselves so animals, vehicles and foods sit in separate neighbourhoods. Drag the white ? point (or tap anywhere on the map) and see which words are closest to it.

    Hand-placed toy map. Real embeddings are learned by a model, not placed by a person, and have far more than two numbers.

    Try the middle of the map. There the nearest words come from different groups, and the distances to them are all long. That is a useful property: the numbers don't only say "these two are related", they also say how related. A far-away nearest neighbour means nothing here is a good match.

    Where do the numbers come from?

    Nobody writes them by hand. A model (a program that has learned patterns from data) is trained on huge amounts of text. During training it is nudged, over and over, so that words used in similar contexts end up with similar numbers. "Cat" and "kitten" appear near words like "purr", "whiskers" and "adopt", so they drift together. "Car" appears near "engine" and "drive", so it settles somewhere else. An early well-known method called word2vec (published in 2013) worked this way. Today's models embed whole sentences and paragraphs, but the idea is the same.

    One honest limit: the model has no idea what a cat is. It has only learned which words keep company with which. That is enough to be very useful, and it is also why embeddings can carry biases found in the text they were trained on.

    Measuring "close" in code

    The map used ordinary straight-line distance. A very common alternative for embeddings is cosine similarity. It looks at the angle between two lists of numbers rather than the gap between them: 1 means they point the same way, 0 means they are unrelated, and negative values mean they point in opposite directions. Here it is in Python, with three made-up scores per word (animal-ness, vehicle-ness, food-ness):

    import math
    
    def cosine(a, b):
        dot = sum(x * y for x, y in zip(a, b))
        size_a = math.sqrt(sum(x * x for x in a))
        size_b = math.sqrt(sum(y * y for y in b))
        return dot / (size_a * size_b)
    
    # made-up scores: [animal-ness, vehicle-ness, food-ness]
    words = {
        "cat":    [0.9, 0.0, 0.1],
        "kitten": [0.8, 0.2, 0.1],
        "car":    [0.0, 0.9, 0.0],
        "pizza":  [0.0, 0.0, 0.9],
    }
    
    for name in ["kitten", "car", "pizza"]:
        print(f"cat vs {name}: {cosine(words['cat'], words[name]):.2f}")

    Running it prints:

    cat vs kitten: 0.97
    cat vs car: 0.00
    cat vs pizza: 0.11

    zip pairs up the numbers position by position, dot multiplies and adds them, and dividing by the two lengths removes the effect of how long each list is. Real systems do exactly this, just with 1,536 numbers instead of 3.

    What this lets software do

    Once text is a list of numbers, a few useful things become easy. Search by meaning: embed the query, embed every document once, return the nearest ones, so "how do I fix a flat" can find a page titled "repairing a punctured bicycle tyre". Recommendations: show items whose numbers sit near things you liked. Grouping: items close together form a topic. Many chatbot features that "look things up in your documents" first use embeddings to pick which documents to show the model.

    What to remember

    An embedding is coordinates on a map of meaning. Close points mean related things, and "close" is measured with plain arithmetic like the code above. When you meet the word in the wild, you now know what it is: a list of numbers, learned from text, that you can compare.