Unsupervised & Self-Supervised
Two different answers to "how do you learn from data with no labels" — find the structure that's already there, or invent a supervised task out of the data itself. Interviewers test whether you know which is which.
Unsupervised learning finds structure in data with no labels anywhere in the loop: clustering, dimensionality reduction, density estimation. Self-supervised learning is a specific strategy inside that same unlabeled-data world. It manufactures a supervised-looking task directly from the data itself. Mask a word and predict it. Make two crops of an image and decide they're the same thing.
You get a real loss function and a real gradient without a human ever providing a label. Almost every foundation model trained since 2018 starts with a self-supervised pretraining phase.
One sentence for an interview: unsupervised learning asks "what structure is already here?" while self-supervised learning asks "what free prediction task can I invent from this data to force a network to learn something useful?" Conflating the two is one of the most common stumbles in this area.
01Intuition
The single most common interview trap in this area: "unsupervised" and "self-supervised" are not the same thing, even though both work on unlabeled data.
Unsupervised learning has no notion of a target at all. Feed k-means a pile of customer purchase vectors and it partitions them into groups. There is no correct grouping it checks its answer against, and no loss function comparing a prediction to a ground truth. The goal is to expose structure already sitting in the data: clusters, a lower-dimensional manifold, a density estimate.
Self-supervised learning still has zero human-provided labels. But it does have a target. The model generates that target itself, from the data, then trains with an ordinary supervised loss against it.
Take a sentence and hide one word. Ask a model to predict it from the rest. The label is the word you already had and chose to hide.
Take one image and produce two randomly cropped, colour-jittered versions. Ask a model to recognise they came from the same source. The label is the fact that you know which crops came from which image, because you made them.
Both are ordinary gradient descent on an ordinary cross-entropy or contrastive loss. The only unusual part is where the label came from.
Why the distinction matters in an interview. If someone asks how you would learn from images you have no labels for, the honest answer branches immediately:
- Explore or compress what you already have. This is an unsupervised question, so reach for PCA or clustering.
- Get a representation that transfers to a task you haven't seen yet. This is almost always a self-supervised question now, so reach for a contrastive or masked-prediction pretext task, because those are the objectives empirically shown to produce features that generalize.
There is a unifying way to think about what a good learned representation is doing, covered from the information-theory side in 01. A representation should keep high mutual information with whatever you care about downstream, while discarding nuisance variation: sensor noise, the exact pixel values of a background, the specific synonym a sentence happened to use.
Every method on this page approximates that same goal. Clustering, PCA, contrastive learning, masked prediction. None of them is ever told explicitly what "what you care about" is.
Unsupervised learning finds structure without labels. Here are six numbers with no labels at all, and a deliberately poor starting guess for two group centers.
Try it Find two groups in six unlabeled numbers with k-means
import numpy as np
x = np.array([1.0, 1.5, 2.0, 8.0, 9.0, 9.5]) # no labels, just numbers
centers = np.array([1.0, 2.0]) # a poor starting guess
for step in range(4):
dist = np.abs(x[:, None] - centers[None, :]) # distances
group = dist.argmin(axis=1) # nearest center wins
centers = np.array([x[group == k].mean() for k in (0, 1)])
print(f"step {step}: groups {group} centers {np.round(centers, 2)}")
1.25 and 7.12. The second moves them to 1.5 and 8.83, and nothing changes after that. Nobody told the algorithm which numbers belong together; it found the two clumps from distances alone.02Timeline
For decades "no labels" meant clustering and linear dimensionality reduction such as k-means and PCA: useful for exploration and compression, but built on hand-designed features and unable to produce representations a downstream network could reuse.
BERT (2018) masked words and predicted them, and SimCLR and MoCo (2020) trained an encoder to place two augmentations of the same image close together, so both manufactured a training signal from unlabeled data.
By the mid-2020s self-supervised pretraining is the standard pipeline: GPT-style next-token prediction, BERT-style masked prediction and CLIP's image-text contrast, with labeled fine-tuning a small final step.
03Architecture: how the data flows
This section is a tour of several methods, since "unsupervised and self-supervised" spans a wide family that shares no single architecture. Grouped by what they do, the space gets much smaller:
- Clustering. Partition or model the data directly.
- Dimensionality reduction. Find a lower-dimensional description of it.
- Autoencoders. Compress it, then reconstruct it.
- Self-supervised learning. Invent a prediction task and solve it with an encoder.
That last bucket is where most of 2020s representation-learning research lives. It gets the most depth here, and the one diagram.
Classical clustering: k-means, k-medoids, hierarchical, DBSCAN
k-means starts by picking k initial cluster centers (centroids), then alternates two steps until nothing changes. Assign every point to its nearest centroid. Recompute each centroid as the mean of the points now assigned to it.
It converges, because total within-cluster distance can only decrease or stay flat each round. But not necessarily to the best possible partition. A bad random initialization can trap it in a poor local optimum, which is why in practice you run it from several starting points (or use k-means++, a smarter initialization that spreads the starting centroids out) and keep the best result.
k-medoids solves the same problem but forces each cluster's center to be an actual data point, a medoid, and not a computed mean, which makes it far less sensitive to outliers. One extreme point can drag a k-means centroid a long way, but it cannot become the medoid unless it is central.
Hierarchical (agglomerative) clustering skips picking k at the start entirely. It begins with every point as its own cluster, then repeatedly merges the two closest clusters until everything is one cluster. "Closest" is set by a linkage rule: closest pair of points, farthest pair, or average distance across all pairs.
The result is a dendrogram, a tree recording every merge. You choose k after the fact by cutting the tree at whatever height gives you the number of groups you want. The cost is scale, because computing and updating all pairwise distances is quadratic in the number of points in the naive form.
DBSCAN drops the idea of a cluster center altogether and defines clusters by density. A point is a core point if enough other points (minPts) sit within a small radius (eps) of it, and clusters grow by chaining together core points that are close to each other. Points that don't belong to any dense region get labeled noise and are not forced into a cluster they don't belong in.
That buys two things k-means structurally cannot:
- Arbitrary cluster shapes, beyond roughly spherical blobs.
- Automatic outlier detection, with no need to specify k in advance.
The tradeoff is that eps and minPts aren't always easy to pick. DBSCAN also struggles badly when different real clusters in the same dataset have very different densities, since one eps cannot be simultaneously right for a tight cluster and a sparse one.
Gaussian mixture models and EM
A Gaussian Mixture Model is what you get if you make k-means probabilistic. Instead of assuming every cluster is a hard, spherical blob with every point belonging to exactly one, a GMM assumes the data was generated by a mixture of k Gaussians.
Each has its own mean, covariance (so clusters can be elliptical as well as round), and mixing weight. Every point gets a soft, probabilistic membership across all k of them, where k-means gives a hard assignment to one.
You can't fit a GMM's parameters with a closed-form formula. You don't know which Gaussian generated which point, and you don't know the Gaussians' parameters. Each depends on the other.
Expectation-Maximization breaks that circularity by alternating, like k-means, with soft numbers in place of hard ones:
- E-step. It fixes the current parameters and, for every point, computes the responsibility: the posterior probability it came from each of the k Gaussians, given where the Gaussians currently sit.
- M-step. It fixes those responsibilities and re-fits each Gaussian's mean, covariance, and mixing weight, using the responsibilities as per-point soft weights in place of hard 0/1 membership.
Repeat, and the data's likelihood under the model provably increases every round until it converges.
k-means is the special case you get by forcing every Gaussian to be spherical, tiny, and equally weighted. The soft responsibilities collapse into hard 0/1 assignments, and the M-step's weighted mean becomes an ordinary mean. Same alternating logic: one is a limit of the other.
Like k-means, EM only guarantees a local optimum of the likelihood. In practice you run it from several starting points (often seeded with a quick k-means pass) and keep the best result by final log-likelihood.
Dimensionality reduction: PCA, ICA, NMF, spectral methods
PCA, ICA, and NMF all answer one question: can this high-dimensional data be described with far fewer numbers? They differ in what counts as a good small description. The eigenvector and SVD machinery behind PCA is covered in full in 01, and the Essence of Linear Algebra series is the best visual build-up of it if eigenvectors still feel abstract.
- PCA finds the orthogonal directions of maximum variance and projects onto the top few. It decorrelates the data, and it is the optimal linear projection if the only goal is preserving variance, equivalently minimizing squared reconstruction error.
- ICA (Independent Component Analysis) asks for something strictly stronger: statistically independent components, which is a stronger condition than uncorrelated.
- NMF (Non-negative Matrix Factorization) adds a different constraint. Every factor and every coefficient must be non-negative.
The ICA distinction matters because decorrelated is a much weaker condition than independent. Two variables can have zero correlation and still be tightly, nonlinearly dependent. ICA needs the underlying sources to be non-Gaussian to be identifiable at all. It is the tool behind the classic cocktail-party problem: given several microphone recordings, each a different mixture of the same underlying voices, recover the original independent voice signals.
NMF's constraint sounds small but has a big effect. Because components can only add, never cancel, NMF tends to discover parts-based, additive representations. Facial features that combine into a face. Topics that combine into a document. Those are often far more interpretable than PCA's components, which mix positive and negative contributions and are hard to read individually.
Spectral methods take a different starting point: build a graph where nearby data points are connected, weighted by similarity, compute that graph's Laplacian, and use its eigenvectors as a new coordinate system for the data.
Spectral clustering then runs ordinary k-means in that new coordinate system instead of raw feature space. This lets it separate clusters that are non-convex or intertwined, such as two concentric rings, which k-means cannot do directly because it only draws straight-line, Euclidean boundaries.
Autoencoders, denoising autoencoders, and VAEs
A plain autoencoder is an encoder-decoder pair trained end to end to reconstruct its own input through a bottleneck. Compress x down to a low-dimensional code z, then decode z back to something close to x, minimizing reconstruction error. It is unsupervised in the sense defined above: no labels, just x predicting x.
But its latent space has no guaranteed structure beyond whatever shape makes reconstruction loss small. So a plain autoencoder is not a generative model: sample a random point in its latent space and, unless it happens to land near where a real training example landed, the decoder produces garbage. Nothing during training ever asked the latent space to be smooth or densely packed.
A denoising autoencoder makes one small, effective change. Corrupt the input before it reaches the encoder (mask patches, add noise), but still ask the decoder to reconstruct the clean original. The network can no longer get away with learning something close to the identity function. It has to learn the real structure and correlations in the data, since the corrupted parts have to be inferred from everything else.
The idea is to corrupt the input and then predict the clean version. It reappears further down this page as BERT-style masked language modeling and MAE-style masked image modeling, with the same mechanism in a different modality.
A Variational Autoencoder fixes the plain autoencoder's ungenerative latent space directly. Instead of encoding x to a single point z, the encoder outputs the parameters of a distribution over z, a mean and variance. A sample is drawn from that distribution, and a KL-divergence term in the loss pulls every input's latent distribution toward a fixed, simple prior, almost always a standard Gaussian.
That regularization term forces the latent space to be smooth and densely packed with no gaps, so sampling from the prior (not just encoding a real example) reliably decodes into something plausible. The full latent-variable setup, and the ELBO derivation behind it, gets its complete treatment in 12. The equation itself is recapped below.
Self-supervised learning: manufacturing a task, then solving it with an encoder
Everything from here on is self-supervised in the specific sense from the intuition section: no human label, but a real, invented target. There are two broad families. The distinction between them (joint-embedding versus generative objectives, in the curriculum's own language) matters more than the list of model names attached to each.
Joint-embedding methods are the ones this page's diagram covers, and contrastive learning is the clearest version of the idea. It runs in four moves:
- Take one unlabeled image. Apply two different random augmentations, such as a random crop, a colour jitter or a horizontal flip. You get two different-looking images that are still, semantically, the same underlying thing.
- Run both through the same encoder. Literally the same weights, applied twice.
- Those two embeddings are a positive pair. The objective pulls them together.
- Every other image in the same batch gives embeddings that are negatives. The objective pushes those apart.
No human ever labeled anything. The label is just "these two crops came from the same source image and these others didn't," a fact you know for free because you generated the crops yourself.
This whole setup, two or more branches sharing weights and trained so that distance in embedding space means something, is a Siamese network. It predates SimCLR by decades. It is the same architecture behind face-verification systems (are these two photos the same person?) and one-shot signature or fingerprint matching.
Metric learning is the general name for training an embedding space where distance is meaningful this way. Usually with a contrastive loss (pull positives together, push negatives apart by at least some margin) or a triplet loss (anchor, positive, negative: push the anchor-negative distance to exceed the anchor-positive distance by a margin).
Three families of joint-embedding method, in order of how much machinery they need:
- SimCLR (2020) is close to the diagram above in its purest form. Large batches supply the negatives directly. In the original paper's own ablations, augmentation strategy turned out to be a bigger lever than architecture.
- MoCo (2020) gets a large number of negatives without needing enormous batches. It keeps a running queue of embeddings from recent batches, computed by a slowly-updated momentum encoder, as a large cheap negative pool.
- BYOL and SimSiam, both 2020, asked a surprising question: do you need negative pairs at all? Both train with only positive pairs and still avoid collapsing to a trivial solution.
How that last one works, and why it is a nontrivial claim, is worth its own discussion in Failure modes below.
The other broad family generates its supervision by masking, then reconstructing, in input space and not in embedding space.
BERT-style masked language modeling (2018) hides about 15% of the tokens in a sentence and trains the model to predict them from the surrounding context. The prediction is bidirectional: the model can look both left and right of the masked token, which a left-to-right autoregressive model cannot do.
MAE-style masked image modeling (2021) does the visual analogue. Mask out a large majority of an image's patches (MAE's own default is a striking 75%) and train a decoder to reconstruct the missing pixels from the small number of visible patches an encoder was allowed to see.
That very high masking ratio is deliberate. Images have far more spatial redundancy than text, so a much more aggressive mask is needed to make the reconstruction task hard enough to force useful learning, since otherwise it could be solved by copying nearby pixels.
Both families are self-supervised and differ in where the prediction target lives:
- Joint-embedding (SimCLR, MoCo, BYOL, SimSiam) never reconstructs anything in input space. The entire objective lives in embedding space, comparing representations to each other.
- Generative or reconstruction (denoising autoencoders, BERT, MAE) predicts directly in the original input space: actual tokens, actual pixels.
The practical tradeoff cuts both ways. Reconstruction objectives are simple and stable to train, being ordinary regression or classification against a well-defined target. A trivial constant-output encoder fails at reconstructing real detail, so collapse isn't a live risk.
But they can spend a lot of model capacity getting low-level detail right that may not matter downstream. The exact texture of grass in the background doesn't matter for classifying "dog."
Joint-embedding objectives sidestep that waste, since they only ever have to make the right high-level things end up close together. But that freedom is exactly what opens the door to the failure mode covered next: representation collapse.
04The equations
InfoNCE
InfoNCE turns the pull/push picture from the diagram above into one loss function. It was introduced for representation learning by van den Oord et al. (2018), and it is the loss SimCLR uses directly, sometimes called NT-Xent (normalized temperature-scaled cross-entropy). For a positive pair (i, j) drawn from a batch of 2N augmented examples, meaning N images with two augmentations each:
- z_i, z_j the two embeddings of the positive pair — same source image, two different augmentations — both L2-normalized so their dot product is a cosine similarity.
- sim(·,·) cosine similarity between two embeddings, ranging from -1 to 1: how aligned the two vectors are.
- τ temperature, a scalar that rescales similarities before the softmax. Small τ sharpens the distribution so the loss reacts much more strongly to hard, nearby negatives; large τ flattens it so every negative contributes more evenly. The project in section 07 has you sweep τ and watch the embedding clusters tighten or smear.
- numerator the positive pair's similarity, exponentiated — the one term gradient descent is trying to make as large as possible relative to everything else.
- denominator the sum over every other embedding in the batch except i itself — normalizing against the whole batch is what pushes every negative's similarity down, since increasing the numerator's share of that sum requires decreasing everyone else's relative share.
The ELBO, recapped
A Variational Autoencoder's loss is built from the evidence lower bound, or ELBO. It is the standard variational-inference objective from Kingma & Welling (2013), and it is what makes maximizing an otherwise-intractable log p(x) practical to train by gradient descent. The full derivation, and the reparameterization trick that makes sampling z differentiable, lives in 12. Here is the equation and what each term does:
- q_φ(z|x) the encoder's output distribution over latents given the input — typically a Gaussian whose mean and variance are predicted by the encoder network.
- E[log p_θ(x|z)] expected log-likelihood of reconstructing x from a sampled z — in practice, this is just the reconstruction loss.
- p(z) the fixed prior being regularized toward, almost always a standard Gaussian.
- D_KL(·‖·) is the penalty for how far the encoder's per-input distribution strays from that prior. A plain autoencoder doesn't have this term, and it is what makes the latent space smooth and samplable.
05Why self-supervised pretraining won
The obvious alternative to inventing a pretext task is labeling more data. It doesn't scale the way self-supervision does, mostly for economic reasons and not algorithmic ones.
Raw text, images, audio, and video are abundant. The internet produces them continuously at a cost close to zero per example. Labels are not cheap, since every one requires a human decision, at a rate and cost that scales with dataset size. Doubling an unlabeled dataset is usually a storage and compute problem. Doubling a labeled dataset is a headcount and time problem, and headcount doesn't scale the way GPUs do.
Self-supervised pretraining sidesteps the bottleneck. The model extracts structure from the raw, abundant data first (what a well-formed sentence looks like, what a real natural image looks like, what tends to co-occur with what) before a single labeled example is involved.
Fine-tuning on labels then only has to teach the specific task on top of representations that already understand the domain. So fine-tuning typically needs orders of magnitude fewer labeled examples than training the same architecture from a random initialization would.
It is also why label-quality issues matter less than they used to. Inconsistent annotators, ambiguous categories, ordinary noise in any large human-labeled dataset: the representation doing most of the heavy lifting was never trained on those noisy labels in the first place.
This doesn't make labels obsolete. You still need them to point the model at the specific task you want, and to check whether pretraining worked at all. What changed is the ratio.
A small, carefully curated labeled set on top of a good self-supervised backbone now reliably beats a much larger labeled set feeding a model trained from scratch. That ratio shift is the practical reason this happened as fast as it did.
06Failure modes, evaluation, and what breaks
This is where "no labels" becomes a problem. Without a ground-truth answer to check against, most of what can silently go wrong here doesn't throw an error. It quietly produces a plausible-looking, wrong result. Picking a clustering algorithm, in particular, is mostly picking which assumption you are comfortable making:
| k-means | GMM (EM) | Hierarchical | DBSCAN | |
|---|---|---|---|---|
| Cluster shape assumed | Roughly spherical, similar-sized | Gaussian — can be elliptical | None explicit — driven by linkage rule | Arbitrary — defined by density |
| Need to pick k upfront | Yes, fixed in advance | Yes, number of components | No — cut the dendrogram after the fact | No, but needs eps/minPts |
| Handles outliers | Poorly — every point forced into a cluster | Soft assignment helps some, still assumes Gaussian shape | Sensitive — outliers can distort early merges | Explicitly labels outliers as noise |
| Rough cost | Fast, scales well | Heavier — fits full covariances | Quadratic in n, naively — doesn't scale to huge n | Near-linear with a spatial index; struggles with mixed densities |
Representation collapse
Make the encoder output the exact same constant vector for every input, regardless of content.
Positive pairs are trivially identical, which gives perfect similarity and a minimised loss.
A naive loss function cannot tell that apart from good representations for free. You have to build in a mechanism that specifically forbids it. That failure mode is representation collapse, and avoiding it is the central design problem in this whole family of methods.
Contrastive methods like SimCLR and MoCo avoid it directly. Negative pairs are explicitly pushed apart, so the collapsed constant-output solution scores as badly as possible on every negative term in the loss. Collapsing destroys the part of the loss pushing negatives apart just as much as it satisfies the part pulling positives together, so there is no net incentive to take the shortcut.
BYOL and SimSiam remove negative pairs entirely and still, empirically, don't collapse, which is surprising. Both rely on an asymmetry between the two branches instead.
- BYOL. One branch, the online network, trains normally by gradient descent and ends with an extra predictor layer that tries to predict the other branch's output. The other branch, the target network, never receives a gradient at all. Its weights are only ever updated as a slow exponential moving average of the online network's weights, with gradients explicitly stopped from flowing into it.
- SimSiam. simplifies further: it drops the momentum-averaged target network altogether and shows a predictor plus stop-gradient alone is enough.
The stop-gradient is the load-bearing piece in both, because it breaks the trivial path where both branches collapse to the same constant together and call it solved, because one branch is deliberately prevented from copying or trivially chasing the other.
The full theoretical picture of exactly why stop-gradient plus a predictor reliably prevents collapse is still an active research question. The SimSiam paper offers an EM-like interpretation, but "guaranteed by a clean proof" would be an overclaim. What is solid is the empirical result, replicated many times over, that this specific asymmetric architecture reliably avoids collapse without ever needing a negative pair.
"Two clusters that are far apart in a t-SNE or UMAP plot are far apart in the real, high-dimensional data; clusters close together are similar." This is only reliable for which points fall into which local neighborhood, because both algorithms are designed to preserve local structure and make no promise about global distances.
Cluster size, the distance between two separate clusters, and even the density of points within a cluster can all be artifacts of the algorithm's hyperparameters (perplexity for t-SNE, n_neighbors for UMAP) and not anything real in the underlying data.
Treat these plots as evidence for "these points are each other's neighbors," never as evidence for exactly how different two well-separated blobs are. van der Maaten & Hinton flag this limitation directly in the original t-SNE paper, and citing it is a strong answer if this comes up in an interview.
Curse of dimensionality, for distance-based methods specifically
k-means, DBSCAN, and anything built on Euclidean distance share a specific, well-documented failure as dimensionality grows. In high dimensions, the ratio between the distance to the nearest point and the distance to the farthest point tends toward 1. Everything ends up roughly equidistant from everything else, and "nearest neighbor" stops being a meaningful concept.
This is exactly why clustering raw pixel-space or raw high-dimensional feature vectors directly is usually a bad idea. The standard fix is to reduce dimensionality first (PCA, or a learned embedding from an autoencoder or a self-supervised encoder) and cluster in that lower-dimensional, more meaningful space instead.
Why unsupervised evaluation is hard
Supervised learning has a built-in success metric: does the prediction match the label. Unsupervised and self-supervised methods have no such thing by construction. There is no ground truth to check a clustering or a learned embedding against, which makes "is this good" a harder question here than almost anywhere else in ML.
In practice you fall back on proxy metrics. Each has specific limits:
- Silhouette score measures, for each point, how much closer it is to its own cluster than to the next-nearest one, averaged over all points and ranging from -1 to 1. Useful for comparing different values of k on the same algorithm. But it still assumes roughly convex, well-separated clusters, so it can penalize a correct DBSCAN-style non-convex clustering for the wrong reasons.
- Linear probe is the standard proxy for self-supervised representations. Freeze the pretrained encoder entirely, train only a small linear classifier on top using a labeled dataset, and report that accuracy.
The linear probe is literally the evaluation protocol SimCLR, and nearly every contrastive-learning paper since, uses to report results. There is no other honest way to say "these representations are good" without eventually testing them against a real task.
07Build this
Representation collapse is the failure this whole family of methods is designed around. Reading about it is one thing. Watching every embedding in your batch slide onto the same point is another.
Train a SimCLR-style encoder on a toy image set where you already know what the right answer looks like. Then delete the negatives and watch the representation die. The embedding is two-dimensional, so it plots directly. It avoids t-SNE on purpose, since section 06 warns that distances in a t-SNE plot are not distances in your data, and this project needs a picture you can trust.
- Generate the data: a few hundred 32×32 images, each one solid shape on a black background, in one of four colors. Color is the only semantic content. Keep the color labels aside, for plotting and nothing else.
- Define two augmentations that preserve color: random resized crop and horizontal flip. Two views per image, one shared encoder, exactly the diagram in section 03.
- Encoder: three small conv layers down to a 2-D output, L2-normalized. Two dimensions means every embedding is plottable with no projection step in between.
- Write InfoNCE by hand from the equation in section 04. Cosine similarity matrix over the 2N views, mask the diagonal, divide by $\tau$, cross-entropy against the index of the matching view. Train at $\tau = 0.1$, then again at $\tau = 1.0$.
- Plot the trained embeddings on the unit circle, colored by the labels you set aside. Log the standard deviation of the embeddings across each batch as training runs.
- Now break it by dropping the negatives entirely: replace the loss with the negative cosine similarity of the positive pair alone, and retrain from scratch.
Where this runs in production
This is a common real pipeline. You have a large pool of unlabeled images from your own domain (product photos, medical scans, satellite imagery, whatever is abundant) and only a small, expensive-to-produce labeled set for the task you care about, say defect classification. Three steps:
- Pretrain a vision encoder on the full unlabeled pool, using a contrastive objective (SimCLR/MoCo-style) or a masked-reconstruction objective (MAE-style). No labels are used yet.
- Attach a small task-specific head to that pretrained encoder and fine-tune on the small labeled set.
- Or, cheaper still, train only a linear probe on the frozen features.
In practice this consistently beats training the same architecture from a random initialization directly on the small labeled set alone. The encoder already learned general visual structure (edges, textures, object parts, shape) from the much larger unlabeled pool before ever seeing a single label. Fine-tuning only has to adapt that structure to the specific task, and does not have to learn vision from scratch on too little data.
08Interview questions
BeginnerWhat's the actual difference between unsupervised and self-supervised learning?
Both use unlabeled data, but only one produces a real training target. Unsupervised learning has no target at all — clustering and PCA just expose structure already in the data, with no loss comparing a prediction to a ground truth. Self-supervised learning manufactures a target directly from the data itself — mask a word and predict it, compare two augmented views of an image — and then trains with an ordinary supervised loss against that manufactured target.
BeginnerWalk me through the k-means algorithm.
Pick k initial centroids. Alternate two steps until nothing changes: assign every point to its nearest centroid, then recompute each centroid as the mean of the points now assigned to it. It's guaranteed to converge but only to a local optimum, so in practice you run it from several random initializations (or use k-means++ for a smarter start) and keep the best result.
IntermediateExplain what happens in the E-step and M-step of EM for a Gaussian mixture model.
The E-step fixes the current Gaussian parameters and computes, for every point, the responsibility — the posterior probability it came from each of the k Gaussians. The M-step fixes those responsibilities and re-fits each Gaussian's mean, covariance, and mixing weight, using the responsibilities as soft, per-point weights. Repeating this alternation increases the data's likelihood under the model monotonically, but only to a local optimum — k-means is the hard-assignment special case of exactly this same alternation.
IntermediateWhy does temperature matter in the InfoNCE loss?
Temperature rescales similarities before the softmax inside the loss. A small temperature sharpens the distribution, so the loss and gradient become dominated by whichever negative is currently closest to the positive — forcing the encoder to focus on separating confusable, hard negatives. A large temperature flattens the distribution, treating every negative roughly equally regardless of how close it actually is, which is a much weaker training signal.
IntermediateWhat's the practical difference between a plain autoencoder and a VAE?
A plain autoencoder's latent space has no guaranteed structure beyond whatever shape minimizes reconstruction loss, so sampling a random point from it and decoding usually produces garbage — it's not a generative model. A VAE's encoder outputs a distribution over the latent instead of a single point, and a KL-divergence term pulls that distribution toward a fixed prior (usually a standard Gaussian), which forces the latent space to be smooth and densely packed — that's what makes sampling from the prior actually generate plausible outputs. Full ELBO derivation is on 12.
DeepWhat is representation collapse, and how do BYOL/SimSiam avoid it without negative pairs?
Representation collapse is the trivial failure mode where the encoder outputs the same constant vector for every input, which vacuously satisfies "positive pairs are close" without learning anything useful. Contrastive methods prevent it with negative pairs — collapsing would also destroy the part of the loss pushing negatives apart. BYOL and SimSiam use no negatives at all; instead they rely on an asymmetry between two branches — a predictor network on one branch plus a stop-gradient (and, for BYOL, a momentum-averaged target network) on the other — that empirically prevents collapse. The exact theoretical reason this specific asymmetry works is still an active research question; what's well-established is the empirical result, replicated widely, that it reliably works.
DeepWhy is evaluating an unsupervised or self-supervised model harder than evaluating a supervised one, and what do you actually do about it?
Supervised evaluation just checks predictions against ground-truth labels. There's no equivalent ground truth for a clustering or a learned embedding, so you're reduced to proxy metrics. For clustering, something like silhouette score — but it assumes roughly convex clusters, so it can penalize a correct non-convex (DBSCAN-style) result. For self-supervised representations, the standard proxy is a linear probe: freeze the encoder, train only a small linear classifier on top using labeled data, and report that accuracy — this is the exact protocol SimCLR-style papers use to claim their representations are actually useful.
09Go deeper
●Now write it yourself
Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.
Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.