Every article about AI throws around four terms as if they meant the same thing: artificial intelligence, machine learning, deep learning and large language models. They do not. They are four boxes, each sitting inside the previous one. Once you see how they nest, most of the noise about AI becomes much easier to follow. This is the first post in our AI basics series, which builds up from here to how a chatbot is actually trained.
The four boxes
Picture four boxes, one inside another:
- Artificial intelligence is the biggest box. It means any program that does something we would normally call intelligent: playing chess, recognising a face, translating a sentence, driving a car. How the program does it does not matter.
- Machine learning is a box inside AI. It means the program was not written rule by rule. Instead it was given examples and it worked out the rules itself.
- Deep learning is a box inside machine learning. It means the learning is done by a neural network with many layers. Almost every AI headline of the last ten years is deep learning.
- Large language models, or LLMs, are a box inside deep learning. They are very large neural networks trained on text to predict the next word. ChatGPT, Claude and Gemini are all LLMs.
So an LLM is a kind of deep learning, which is a kind of machine learning, which is a kind of AI. When someone says "AI wrote this email" they mean an LLM. When someone says "AI spotted the tumour" they probably mean a deep learning model that is not a language model at all.
AI before machine learning: writing the rules by hand
For its first few decades, AI mostly meant people writing rules. A chess program had rules like "a queen is worth nine pawns" and searched through millions of positions applying them. A spam filter had a list of suspicious words. A medical expert system had thousands of if-then statements collected from doctors.
This works for problems where humans can write down the rules. It fails badly for problems where we cannot. Try writing rules that tell a cat photo from a dog photo. What pixel pattern is "cat"? Nobody can write that down, and yet a three year old does it instantly. That gap is what machine learning fills.
Machine learning: show, do not tell
In machine learning you do not write the rules. You collect examples, each labelled with the right answer, and you let an algorithm find a function that maps inputs to answers. Show it 10,000 cat photos labelled "cat" and 10,000 dog photos labelled "dog", and it finds patterns in the pixels that separate the two. You never say what those patterns are. Often you cannot even inspect them afterwards.
The examples are called training data. The process of finding the function is called training. The finished function is called a model. Using the model on new data is called inference. Those four words will come up in every post of this series.
There are three broad flavours:
- Supervised learning. Every example has a label. Photo and "cat". House details and its sale price. Email and "spam". This is the most common kind and the easiest to understand.
- Unsupervised learning. No labels. The algorithm looks for structure on its own, such as grouping customers into clusters that behave alike.
- Reinforcement learning. No labels either, but a reward signal. The program tries things, gets points for good outcomes, and adjusts. This is how game playing systems learn, and it is also the last stage of training a chatbot, which we cover in part five.
Machine learning is not new. Linear regression, the line of best fit you may have met in a statistics class, is machine learning. Spam filters have used it since the 1990s. Netflix recommendations, credit scoring and fraud detection are all machine learning that predates the current AI wave.
Deep learning: many layers of simple units
Machine learning has many algorithms: decision trees, support vector machines, nearest neighbour lookups. Deep learning is one of them, and since about 2012 it has been the one that wins on images, speech and text.
A neural network is a stack of layers. Each layer takes a list of numbers, multiplies them by some weights, adds them up, applies a simple squashing function, and passes a new list of numbers to the next layer. "Deep" just means there are many layers, sometimes hundreds. The word neural is a loose nod to brain cells. Nothing about a modern network works like a brain, and you can safely ignore the biology.
What made deep learning take off was not a new idea. The maths dates from the 1980s. Three things arrived together: huge labelled datasets from the internet, graphics cards that happen to be very good at the multiplications involved, and a few engineering tricks that let very deep stacks train without falling apart. We walk through exactly how a network learns from examples in part two.
The key property of deep learning is that the layers learn their own features. In older machine learning a human had to decide what to measure: edge counts, colour histograms, word frequencies. A deep network starts from raw pixels or raw text and works out for itself which patterns matter. Early layers find edges, later layers find eyes and ears, the last layers find "cat".
Large language models: predict the next word, at enormous scale
An LLM is a deep neural network with one training task: given some text, guess the next piece of text. Feed it "The capital of France is" and it should output "Paris". Do this over trillions of words from books, websites and code, and something surprising happens. To predict the next word well, the model has to absorb grammar, facts, reasoning patterns, coding conventions and the style of every kind of writing. It ends up with a compressed statistical picture of most of what humans have written down.
"Large" is not a figure of speech. The weights in these networks, the numbers being tuned during training, run to hundreds of billions. Training one takes thousands of graphics cards running for months. That is why a handful of companies build them and everyone else rents access.
Two things about LLMs are worth fixing in your head now, because they explain most of their odd behaviour:
- They generate one piece at a time. A model does not plan a paragraph and then write it. It picks the next token, appends it, and picks again. Part three explains what a token is and part four explains how the model looks back at everything so far to choose the next one.
- They are trained to be plausible, not to be true. The training signal is "what text usually comes next", not "what is correct". Correct text is usually plausible, so the model is usually right. But when it does not know, it produces something that sounds right anyway. That is a hallucination, and part five covers why it happens and what reduces it.
Where the popular products fit
| Product | What it is underneath |
|---|---|
| ChatGPT, Claude, Gemini | LLMs with an extra training stage to make them follow instructions and hold a conversation |
| GitHub Copilot, Claude Code, Cursor | LLMs trained heavily on code, wrapped in tools that let them read and edit files |
| Midjourney, DALL-E, Stable Diffusion | Deep learning, but diffusion models rather than language models. They learn to turn noise into images guided by a text description |
| Face unlock on your phone | Deep learning image model, small enough to run on the device |
| Netflix recommendations | Classic machine learning, mostly not deep |
| Chess engines like Stockfish | Mostly hand written search, with a small neural network for position evaluation added in recent years |
| A thermostat that learns your schedule | Simple statistics. Marketing calls it AI because everything is AI now |
Why this matters for how you use these tools
Knowing which box a tool sits in tells you what to expect from it. A supervised model is only as good as its labels and only works on inputs like the ones it saw. A language model is fluent in everything but certain of nothing, so its output needs checking in proportion to how much is riding on it. A recommendation system is optimising for your clicks, not your interests. None of this is mysterious once you know what the machine was actually trained to do.
What comes next
This series builds up in order. Part two opens the box and shows, with a network small enough to compute by hand, how weights, a loss function and gradient descent turn examples into a working model. After that we get to text: tokens and embeddings, then the transformer architecture, then the training pipeline that turns a raw model into a chatbot.