Supervised vs Unsupervised Learning: What’s the Difference?

Supervised learning trains on labelled examples — inputs paired with the correct output — and learns to predict the label for new inputs. Unsupervised learning is given data with no labels and finds structure in it: clusters, density, compressed representations. The practical divide is the cost and availability of labels: labels are expensive and often scarce, which is why the largest models are pre-trained unsupervised on raw text and only then fine-tuned, supervised, on far less labelled data.

Supervised
Learns from labelled input→output pairs
Unsupervised
Finds structure in unlabelled data
Why it matters
Labels are the scarce, costly ingredient

Why the frontier is mostly unsupervised

Supervised learning is powerful but bounded by labelled data, which humans must produce. Unsupervised (and the self-supervised variant that invents its labels from the data itself, like predicting the next token) can consume effectively unlimited raw text — which is why foundation models are pre-trained that way, learning language and world structure before any labels are involved.

The modern pipeline is therefore a sequence, not a choice: self-supervised pre-training on raw data to build general capability, then a comparatively small supervised and preference-based stage to shape behaviour. Framing the two as rivals misses that today’s systems depend on both in order.

Related standards

Goodfellow, Bengio & Courville, 2016 — Deep Learning

Questions

Where does self-supervised learning fit?

It is unsupervised in needing no human labels, but creates a prediction target from the data itself — the basis of language-model pre-training.

Is one better?

Neither — they solve different problems and are used in sequence: unsupervised pre-training, then supervised and preference fine-tuning.

Keep reading

related
What Is Model Evaluation (Benchmarks and Evals)?
related
What Is Fine-Tuning?
related
What Is RLHF (Reinforcement Learning from Human Feedback)?
related
What Is an Embedding?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
What Is Agent Context (and the Context Window)?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
Managing an Agent’s Context Window
Agent Data & Memory
What Is a Vector Database?
Agent Data & Memory
What Is a Mixture-of-Experts Model?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is a Transformer (Neural Network Architecture)?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Model Distillation?
Agent Data & Memory
What Is Model Quantization?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
Transformer vs RNN: What Changed?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.