Back to journal How Transformer Models Work, Explained with a Small PyTorch Example

How Transformer Models Work, Explained with a Small PyTorch Example

A beginner-friendly guide to tokens, attention, and how transformer models turn a prompt into a reply.

From words to numbers

A transformer is a kind of computer model that is very good at working with sequences, especially text. Chatbots and many writing tools use transformers. The name sounds complicated, but the first steps are easier to picture: turn text into small pieces, turn those pieces into numbers, and let the model look for patterns.

Those small text pieces are called tokens. A token might be a whole short word, part of a longer word, or punctuation. The sentence “cats chase a ball” might be split into tokens like cats, chase, a, and ball. The model then maps each token to a list of numbers. Those lists give the model a way to do maths with language.

Word order matters. “A dog chased a cat” does not mean the same thing as “a cat chased a dog.” Transformers add information about where each token appears so the model can tell the difference.

Text→Tokens→Numbers the model can process

Attention: choosing what matters

Here is the useful trick behind the name. When processing one token, the model can look at other tokens in the same sentence and decide which ones matter most to it. This process is called attention.

Consider: “Maya put the book on the table because it was flat.” To work out what “it” refers to, you need to look back at the other words. Attention is a maths-based way for the model to connect related pieces of text. It does not make a human-like judgement; it calculates which connections are useful for its next step.

The model uses many attention calculations at once. Different calculations can notice different kinds of patterns, such as which names a pronoun might refer to or which words belong together. Their results are mixed and passed through more layers that refine the numbers.

How does it make a reply?

A text-generating model usually predicts one next token at a time. If the prompt is “The opposite of hot is”, it may give “cold” a high chance. It picks a next token, adds it to the text, and predicts again. Repeating this produces a sentence.

The model is not opening a dictionary and copying the perfect answer. It is calculating likely next pieces based on patterns it learned during training and the words in its current context. That is why changing the prompt can change the reply—and why a fluent reply can still contain a mistake.

What my PyTorch exercise showed

I used a small PyTorch example to study the parts that make a transformer work. PyTorch helps describe the number operations and run them on example data. Seeing tokens move through steps made the diagram feel less mysterious than just reading the word “attention.”

This was a learning exercise, not a model trained to chat. A real large language model has an enormous number of learned values and is trained on far more data and computing power. The small example helped me understand the shape of the idea, even though it does not behave like a production chatbot.

One takeaway: a transformer does not read a sentence the way you do. It turns text into numbers, calculates connections between parts, then uses those patterns to choose what comes next.