GenAI & LLMs · 1 min read · Updated 2026-10-05
Transformers and attention
The old problem
Earlier AI read one word at a time and often forgot the start of a long sentence. It was also slow to train.
What attention does
Take 'The cup would not fit in the bag because it was too big'. What is 'it'? You know it is the cup. Attention helps the computer work that out by linking 'it' back to 'cup'.
Why it won
Transformers read all the words together, so they train fast on powerful chips. Researchers at Google introduced them in 2017. GPT, Claude, Gemini and Llama all use them.
Key takeaways
- Attention connects each word to the words that explain it.
- Reading everything at once makes training fast.
- Almost every big chatbot is a transformer.
Quick questions
Who invented it?
A team at Google, in a 2017 paper called 'Attention Is All You Need'.
What does GPT stand for?
Generative Pre-trained Transformer.
Are all chatbots transformers?
Nearly all the big ones are. Researchers are trying other designs too.
Keep learning
What is a large language model (LLM)?An LLM is a computer program trained on a huge amount of writing. Its main skill is guessing the next word, again and…1 min readRead →Tokens and the context windowA token is a small piece of text, about three-quarters of a word. The context window is how much text the AI can keep in…1 min readRead →