How does ChatGPT know where the ketchup is?

Or: a gentle introduction into the architecture of GPT-style Large Language Models (LLM).

What does GPT stand for?

GPT stands for Generative Pre-trained Transformer. Let’s break that down:

Generative: This means it can generate or create new text, like writing a story, answering questions, or even composing emails.

Pre-trained: Before you use it, the model has already learned a lot by reading a vast amount of text from the internet, books, and other sources. It’s like a student who has studied a lot before the exam, so when you ask it a question, it already knows a bit about many things.

Transformer: This is the type of architecture or structure of the model. Think of it as the “brain” of the system, designed to understand and generate language in a very sophisticated way.

How does it work?

Imagine you’re trying to learn a new language by reading a giant library of books:

Learning from examples: The model reads through texts, learning patterns in how words connect, the structure of sentences, and the context in which words are used. It’s like learning the rules of grammar and vocabulary by osmosis, but in a very digital way.

Analogy: Think of it like a chef tasting hundreds of dishes to learn what flavors work together. The chef (GPT) doesn’t just memorize recipes; it learns how ingredients (words) complement each other.

Predicting the next word: When you type or ask something, the model tries to guess what word should come next based on what it’s learned. This is similar to how you might predict the next word in a sentence you’re reading.

Example: If you type “The cat sat on the…”, the model might suggest “mat” because it has seen this pattern many times in texts it has read.

Context is king: The model doesn’t just look at the last word but considers the entire sentence or even paragraphs to understand the context. This is what makes it good at understanding questions or continuing stories.

Analogy: Like guessing the plot twist in a movie by understanding the entire storyline, not just one scene.

Generating responses: Once it has a context, it generates responses by piecing together words in a way that makes sense, aiming to be coherent and relevant.

Example: If asked “What do cats like?”, it might respond with “Cats generally enjoy playing with toys, napping in warm spots, and chasing after small creatures or objects.” It forms this sentence from bits of knowledge it has absorbed.

In essence, GPT is like an extremely well-read friend who can write or speak on almost any topic you bring up, based on what it has learned from reading. However, remember, it doesn’t have personal experiences, emotions, or real understanding; it’s all about patterns and statistics of how words are used together in the texts it has seen.

GPT models are word predictors

Think of GPT-based Language Models (LLMs) as very sophisticated “word predictors.” Here’s how it works:

When you have a sentence like “The ketchup was in the…”, the LLM looks at this partial sentence and tries to guess what word should come next based on what it has learned from vast amounts of text it was trained on. Here’s how it might work with the ketchup example:

Most probable word: The LLM might predict “bottle” because:
It’s very common for ketchup to be in a bottle.
In many sentences it has seen, “bottle” frequently follows phrases like “The ketchup was in the.”
Other possible words:
“Kitchen”: This could be predicted because kitchens are where you commonly find ketchup, even if not directly in the sentence context.
“Cupboard”: Similar logic, as ketchup is often stored in cupboards.
“Fridge”: Although less common for ketchup, some people do refrigerate it, so this word has been seen in this context before.

The LLM essentially:

Analyzes context: It looks at the words already used in the sentence (“The”, “ketchup”, “was”, “in”, “the”) and considers what usually comes after such sequences.

Assigns probabilities: Based on its training data, it calculates the likelihood of each word following the given text. “Bottle” might have the highest probability, but “kitchen”, “cupboard”, or “fridge” are also plausible.

Chooses or suggests: If generating text, it might choose “bottle” for a response because it’s the most probable.
If it’s assisting in writing or auto-completing, it might show a list of these options in order of likelihood.

The key here is that these models don’t just memorize sentences but learn patterns in language use. They predict based on statistical likelihood from patterns seen in their training data rather than understanding the actual content or the world. So, while “bottle” might be the best guess, other contextually relevant words like “kitchen”, “cupboard”, or “fridge” are considered because they’ve been seen in similar contexts in the texts the model has processed.

This word prediction ability is what makes these models useful for tasks like text completion, translation, or even creative writing, as they can suggest or generate text that feels natural and coherent based on the input given.

A big disadvantage to this type of LLM is that it can sometimes “hallucinate“. More about that in another blog post.

Leave a comment

Your email address will not be published. Required fields are marked *