Everyone talks about ChatGPT, Claude, Gemini. Few people know what GPT stands for, or what actually happens when you type a question. It's not magic. It's not intelligence in the human sense either. It's something else. And understanding this "something else" changes the way you use these tools. Since 24 September 2026, this article also brings together our explanation of the GPT acronym, previously published separately.
What GPT stands for
GPT stands for Generative Pre-trained Transformer. Three words, three ideas:
GPT is the brand of OpenAI's models, but the word has become generic. The rest of this article explains what each of these three ideas covers.
The stochastic parrot: an imperfect but useful metaphor
We've often heard that LLMs are "stochastic parrots": they repeat patterns without understanding. It's partially true, but it's reductive.
An LLM doesn't store text that it regurgitates. It has learned statistical relationships between words. When it generates "The cat is on the...", it doesn't search a database. It calculates that "mat", "sofa" or "bed" have a high probability of following, based on billions of examples.
What this actually means: An LLM doesn't "know" anything in the strict sense. It predicts. Very well. But predicting isn't knowing.
Attention: the 2017 invention that changed everything
Before 2017, language models read text word by word, in order. It was slow, and they lost the thread of long sentences.
The Transformer, introduced that year by Google researchers, brought in a mechanism called "attention": the model looks at every word in a sentence at once and weighs the links between them.
Example: in "The cat that was on the mat fell asleep", it links "fell asleep" to "cat", not to "mat". Earlier models stumbled over sentences like this.
That ability to hold the overall context is what makes today's LLMs so fluent.
Tokens: why AI chops up your words oddly
LLMs don't read words. They read "tokens": chunks of words.
"Unconstitutionally" becomes several tokens: "Un", "constitu", "tion", "ally". The word "the" in English is a single token. The word "Sion" (the Swiss city) might be split into "S" + "ion".
Why this matters:
An answer written one token at a time
An LLM doesn't think before it writes. It produces its answer one token at a time: it reads your request, predicts the first token, adds it to the context, predicts the next one, and repeats until the end.
That's why:
Training: billions of texts, zero truth
An LLM is trained on massive quantities of text: books, websites, articles, forums, source code. Everything written by humans.
The problem: The internet contains both true and false information. The LLM learns both without distinction. It learns that "the earth is round" AND "the earth is flat" exist as sentences. It doesn't know which one is true; it only knows which one is more frequent in certain contexts.
This is why LLMs "hallucinate": they generate plausible text, not true text. If you ask for an obscure fact, the model will produce something credible, whether it's accurate or invented.
And a cut-off date: training stops at a given point. The model knows nothing of what happened afterwards, unless it has access to an online search. Biases present in the training texts carry over into its answers.
Temperature: the creativity/precision slider
When an LLM generates text, it can be more or less "creative". This is controlled by a parameter called temperature.
In practice:
GPT, Claude, Gemini, Llama: cousins
Almost every major model is built on the same Transformer architecture:
What sets them apart isn't the basic principle, but the training data, the adjustments made afterwards, the safety choices and the size of the model.
When a company boasts about "its revolutionary proprietary AI", chances are it's a Transformer like the others, with an interface on top.
What this changes for you
Understanding these mechanisms changes the way you interact with an LLM:
1. Don't trust it for facts
Always verify. Especially dates, numbers, and citations.
2. Be precise
The clearer your request, the less the model has to "guess". The less it guesses, the less it invents.
3. Context matters
An LLM with your conversation history performs better than a "cold start" LLM.
4. The limits aren't bugs
Hallucination isn't a problem OpenAI will "fix". It's inherent to how it works.
Key takeaways
GPT = Generative Pre-trained Transformer: a model that generates, pre-trained, built on the Transformer architecture.
An LLM predicts probable text, one token at a time; it doesn't "understand" in the human sense.
Attention, invented in 2017, lets it hold the context of a whole sentence.
Training on the internet includes both true and false, and stops at a date: hence hallucinations and gaps.
Temperature sets the slider between creativity and reliability.
GPT, Claude, Gemini and Llama are cousins: same architecture, different settings.