Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 4 · Theme 1: Representational spaces

Nonlinear dimensionality reduction

Thursday, September 24 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Your questions from Tuesday

  • "If there are 50 principle components, why is it/how can it be represented by just 2 axes as on the ppt slides?": it cannot. A PCA plot keeps the two directions along which the points vary most over the whole data set and drops the other 48. Points that differ only along the dropped directions land on top of each other. That is where today's nonlinear methods come in: instead of keeping a few directions, they keep relationships between points, above all who is whose neighbor.
  • "Does this mean neural signals are operating in 2 dimensional space": no. For the Bao et al. objects in AlexNet fc6 (Assignment 1), the first two components hold only 22% of the variance, and 85% takes 87 (Bao et al. report about 50). Counting components to a threshold measures the embedding dimension, a linear count, and treats the rest as noise. If the data lie on a curved surface, the intrinsic dimension, the number of dimensions needed to locate a point on it, can be far smaller.

Next: what every 2-D picture of the data gives up.

Every low-dimensional picture is an approximation

A three-dimensional set of points lying close to, but not on, a plane spanned by two principal directions v1 and v2, with axes x-tilde 1, 2 and 3. Dashed red residuals join several points to the plane; one point, x-tilde, is joined to its projection on the plane PCA: the same 100 points drawn in two dimensions, each at its coordinates z1 and z2 along v1 and v2; the projected point is marked in green, with dotted lines to z1 and z2 on the axes MDS on the Euclidean distances between the same points: a two-dimensional map that looks the same as the PCA map, with the same point in green
  • 100 points in three measured coordinates, x~1,x~2,x~3\tilde{x}_1,\tilde{x}_2,\tilde{x}_3: they lie close to a plane, not on it.
  • How do we draw them on a 2-D page, and what do we give up?

Every low-dimensional picture is an approximation

A three-dimensional set of points lying close to, but not on, a plane spanned by two principal directions v1 and v2, with axes x-tilde 1, 2 and 3. Dashed red residuals join several points to the plane; one point, x-tilde, is joined to its projection on the plane PCA: the same 100 points drawn in two dimensions, each at its coordinates z1 and z2 along v1 and v2; the projected point is marked in green, with dotted lines to z1 and z2 on the axes MDS on the Euclidean distances between the same points: a two-dimensional map that looks the same as the PCA map, with the same point in green
  • PCA keeps each point's projection onto the plane of largest variance; MDS keeps the distances between points.
  • PCA drops each point's distance from the plane, the dashed red residual, assuming it is noise.
  • MDS cannot drop it: to keep every distance, it spreads the residual into the map, which distorts it.

Green point, far above the plane: PCA drops it among the others; MDS pushes it to the edge.

Real images curve: turning one photograph

Left, eight frames of one photograph of a marlin turned from 0 to 310 degrees, each drawn whole on a padded square. Right, all 36 turned images treated as pixel vectors and projected onto their first two principal components: the points trace a closed loop, with five of the frames drawn next to their points

  • Why look? Recognition needs one object's images to stay together. In pixels, turning alone sends the fish on a loop that needs 24 PCs for 90%, yet one dimension, the angle, locates every image.
36 images, 9,216 pixels each: fewer samples than dimensions, so PCA is computed with the SVD of the centered data (Assignment 2)

Shifting traces a curve too

One photograph shifted from 14 pixels left to 14 pixels right: eight frames and, on the first two principal components, an arch-shaped curve with five of the frames drawn next to their points

  • Slide the photograph left to right: one knob, one curve. It bends, so it is not a straight line in pixel space.

Enlarging traces another curve

The same photograph enlarged from half size to 1.45 times: eight frames and, on the first two principal components, another arch-shaped curve with four of the frames drawn next to their points

  • Each knob gives a one-dimensional curve. Turn all three at once and the images fill a three-dimensional curved surface.

All three changes in one PCA

All 90 images of the marlin, turned, shifted and enlarged, projected together onto the first two principal components of the whole set, on the same axes and scale as the previous three slides: the turning loop in red, the shifting arc in blue and the enlarging path in green, each distorted and crossing the others

  • The same 90 images as the last three slides, now in one PCA.
  • Alone, each change was a clean curve. Together they are distorted and cross: from this map you cannot read how far the fish was turned, shifted or enlarged.
  • Each image is set by one knob value, yet PCA needs 36 components to keep 90% of the variance.

In pixels, one fish's images spread over many curved directions. Does a trained network pull them together?

Turn the knobs yourself

What is a manifold?

The Earth photographed from space: a curved surface that looks flat to someone standing on it

A curved surface drawn as a mesh shaped like a bowl: red rings and green spokes. It is made of 216 images of the marlin, turned in 15-degree steps at nine sizes, placed on their first three principal components. Each red ring is one size turned all the way round; each green spoke is one angle enlarged. One small mesh cell is shaded gold. Thumbnails show the smallest marlin at the bottom of the bowl and two large ones at the rim

  • Turn and enlarge one photograph: red rings are one size turned round; green spokes are one angle enlarged.
  • The images fill a curved surface: a manifold. Up close (the gold cell) it is flat.
  • Intrinsic dimension: how many dimensions are needed to locate a point on it? Here 2, the angle and the size.
  • Ambient space: where it sits, 9,216 pixel values.

Next: many objects, each with its own manifold, in a trained network.

Objects seen from many viewpoints

A grid of renderings, one row per category (animal, vehicle, face, vegetable or fruit, house, man-made object). Each row shows two different objects of that category, for example a cat and a duck, an airliner and a bicycle, each seen from three of its 24 viewpoints

  • 1,224 images: 51 objects, each rendered from 24 views, in 6 categories.
  • Bao, She, McGill & Tsao showed these images to monkeys while recording from IT cortex, the late stage of the visual system where neurons respond to whole objects. You also use them in Assignment 1.
  • Here each image is a point in the late layer of ResNet-18, the trained image network of Assignment 1: 4,096 numbers per image.
  • Two levels of grouping: views of the same object, and objects of the same category.
  • The question: pixels scatter an object's views. Does the network keep them together, and keep different objects apart?

One object, 24 views: one patch

The 24 views of one excavator as points on two axes, PC1 and PC2, from a PCA of the network's responses to just these 24 images. The points form one compact triangular patch, shaded red. Eight of the views are drawn around the patch, each joined by a line to its own point

  • Each dot is one photograph of the same excavator, placed by the network's response to it. The axes are the two principal directions of those 24 responses.
  • Very different images, one patch: all the ways this object can look.
  • Compare the pixels, where turning alone traced a long loop: the network gathers the views.

Three objects, three patches

The 24 views of each of three objects, an excavator in red, a swan in blue and a teacup in green, on one shared pair of axes from a PCA of all 72 responses. Each object's views form a small separate patch, far from the other two; one photograph of each object is drawn beside its patch

  • Now one PCA of all 72 responses together. Each object keeps to its own small patch, and the patches are far apart.
  • Viewpoint moves a response within a patch; identity moves it between patches.
  • And the brain? Bao et al. recorded monkey IT cortex with these same images: do its neurons also keep each object together? That question drives the recognition lectures.

Drawn in 2007, computed today

DiCarlo and Cox's drawing: one person's face images trace a curved sheet, a manifold, inside image space, with three example heads at different poses on it
drawn · DiCarlo & Cox, 2007
The three-object figure from two slides back: 24 views each of an excavator, a swan and a teacup, as a trained network's late-layer responses on two principal components; each object forms its own compact patch, far from the others
computed · three objects in a trained network

One sheet per object. In a trained network, the sheets of different objects sit apart. In pixels, they do not: next slide.

Two sheets, tangled

In actual pixel space the sheets of two different people are interleaved, and a flat separating plane drawn through them fails to split them
  • In pixel coordinates the sheets of two people interleave.
  • Recognizing who it is means finding a flat (linear) boundary, a hyperplane, with one person's sheet on each side. Here none exists.
  • That boundary is what a linear classifier learns. A good representation untangles the sheets; testing that with a linear readout is the next lecture.

Next: to see whether a representation untangles them, we need to draw manifolds.

Untangled: a flat boundary now works

Left: the views of two objects form two linked rings in pixel space, so no flat plane can put one ring on each side. Right: in a later layer the same two rings are pulled apart, and a flat plane passes between them

  • Same two objects, two representations: in pixels the sheets are linked; in a later layer they are pulled apart
  • Untangling means a flat boundary can now separate the objects

How do we see a manifold?

To ask whether object manifolds (the responses to one object across all its views) are tangled or untangled, we have to look at them, but they live in thousands of dimensions. We need a 2-D map that keeps their shape. First, test the maps on a manifold whose true shape we know.

Test case: can PCA unroll a known manifold?

A chocolate Swiss roll cake with the cream spiral visible at its cut end

Three panels. Left, the Swiss roll in 3-D with axes x1, x2, x3; points A and B lie on neighboring turns. Middle, the true structure: the same points laid out as the flat sheet before rolling, where A and B are far apart. Right, a flat 2-D plot with axes PC1 and PC2: the PCA projection, where the turns overlap, colors mix and A and B land on the same spot

  • 1 → 2: a 2-D sheet rolled up in 3-D. A and B are 11.5 apart in a straight line (Euclidean distance in 3-D), but 59.5 apart walking along the sheet.
  • 3: PCA is linear: it can only rotate the data and project onto a flat plane. It keeps 71% of the variance, yet folds the roll: A and B land on the same spot.

Linear methods fold the roll

Three panels. Left, the Swiss roll in 3-D, colored blue to gold along the sheet, with points A and B on neighboring turns, exactly as on the previous slide. Middle, its PCA projection: the turns overlap, the colors mix, and A and B land together. Right, MDS on the straight-line distances between points: the map is still a spiral, with colors that should be far apart lying next to each other

  • PCA keeps the directions of largest variance; MDS keeps the straight-line distances. Both see the data through a flat lens.
  • Points on neighboring turns are close through the air, so both methods put them close. Neither recovers position along the sheet. The fix: better distances, not a better projection.

Perich's origami crane

A white paper crane folded from one uncut square of paper, photographed on a plain blue background

  • A crane is folded from one flat square: nothing is cut or glued.
  • It sits in 3-D, yet it is still a 2-D sheet: unfold it and you get the square back.
  • If the square had text printed on it, the crane would hide it. To read it, you must unfold, not just look from a better angle.
Photo: Andreas Bauer · CC BY-SA 2.5 · Wikimedia Commons · analogy from Perich 2026, The Transmitter

Squashing is not unfolding

Three photographs of white paper on blue. Left, a sheet folded like an accordion, standing in 3-D. Middle, the same accordion pressed flat into a thin stack, all its panels on top of each other. Right, the sheet opened out flat, the fold lines visible between five panels
the data: a sheet folded in 3-D
PCA: squashed flat, the folds stacked
what we hope for: the sheet, unfolded
  • PCA picks a flat view, so the panels land on top of each other. A nonlinear method aims to unfold the sheet.

Three ways to count dimensions

Swiss roll Turned photo Bao objects (fc6)
Ambient: coordinates per point 3 9,216 pixels 4,096 units
Embedding: flat (linear) subspace PCA needs 3 24 PCs (90%) 149 PCs (90%)
Intrinsic: dimensions to locate a point on it 2 1 (angle) unknown
  • The three numbers answer different questions; none is the dimension.
  • PCA measures the embedding dimension, a linear count that needs a variance threshold (90% here). It overestimates the intrinsic one whenever the manifold curves.
  • For real data such as the network's responses to the Bao objects, the intrinsic dimension is unknown: that is what the nonlinear methods coming next try to reveal.
Perich 2026 · The Transmitter

How could we fix this?

Keep the MDS idea, but change the distances. Instead of measuring straight through the air, measure along the data: only step from a point to its nearest neighbors, and add up the steps.

Isomap: measure "along the sheet", then run MDS

Left, the Swiss roll with two probe points one turn apart: a short dashed red segment joins them straight through the roll, and a long black path joins them by hopping between nearest neighbors around the spiral. Right, the Isomap layout: a flat strip ordered blue to gold, with the two probes at opposite ends

  • Connect each point to its 10 nearest neighbors. The shortest path through that graph approximates the "along-the-sheet" distance: 71.2 for the two probes, 6.1 in a straight line.
  • Run MDS, from the second lecture, on those path lengths. The roll unrolls if the points cover the sheet densely and no edge jumps between turns.

Folded paper: you cannot read it flat

Four panels. 1, a page printed with the word MANIFOLD. 2, the page rolled into a Swiss roll in 3-D, the letters wrapped around it. 3, the PCA projection: the turns land on top of each other and the letters pile up into an unreadable block. 4, the Isomap layout: the page unrolled, and MANIFOLD reads again

  • The crane, computed: a page printed with text, rolled into a Swiss roll.
  • PCA gives the best flat view, but the letters pile up. To read the text, you must unfold the sheet: a nonlinear method such as Isomap (panel 4).
Perich 2026 · The Transmitter

Isomap in practice

Isomap on the Swiss roll with 10 neighbors: a clean flat strip ordered blue to gold. With 40 neighbors: the neighbor graph jumps between turns and the map curls back into a spiral with mixed colors

  • One wrong edge between turns (a short circuit) and the roll folds again.
  • It also needs densely sampled data without holes, and becomes slow for large data sets.
  • So Isomap is rarely used today. The popular modern methods are t-SNE and UMAP; we focus on t-SNE, which you use in Assignment 1.

t-SNE

Convert distances into neighbor probabilities, then find a low-dimensional layout, usually 2-D, that approximates those neighbor probabilities.

What t-SNE is trying to do

Top, 24 points in three tight groups in a plane (blue, red, green), labelled N points in D dimensions. Bottom, the 1-D t-SNE map of the same points: a single line on which each group occupies its own stretch, red, then green, then blue

  • Like MDS, t-SNE places the points in a low-dimensional map, usually 2-D (here 1-D, a line, so the example is easy to see), so that relationships in the map match relationships in the data.
  • MDS tries to match all distances. t-SNE matches neighbors: close in the data stays close in the map; large distances are not kept.
  • t-SNE does not look for clusters: clusters survive because their members are each other's neighbors.
  • Three steps: (1) data: distances → neighbor probabilities PP; (2) map: similarities QQ; (3) move the map points until QQ matches PP.

Step 1 — in the data: distances → PP

Top, the three-group data in the same positions as on the setup slide, with one point, x-tilde-i, starred; the other points are drawn larger the more neighbor probability the starred point gives them, so only its own blue group is visibly large. Bottom, a Gaussian curve of similarity against distance from the starred point, with every other point placed on it (bottom): the blue points sit on the steep part, the red and green points at almost zero

  • Center a Gaussian on each point x(i)\mathbf{x}^{(i)} and normalize:

pj∣i  ∝  exp⁡ ⁣(−∥x(i)−x(j)∥22σi2)p_{j\mid i}\;\propto\;\exp\!\left(-\frac{\lVert\mathbf{x}^{(i)}-\mathbf{x}^{(j)}\rVert^{2}}{2\sigma_i^{2}}\right)

  • Read it as: "if I am point ii, how likely am I to pick jj as my neighbor?" In MDS terms, PP replaces the dissimilarities, warped: near pairs count a lot, far ones almost nothing.
  • Each point gets its own width σi\sigma_i, a free parameter: narrow in dense regions, wide in sparse ones.
  • Make it symmetric: pij=pj∣i+pi∣j2Np_{ij}=\dfrac{p_{j\mid i}+p_{i\mid j}}{2N}, one probability per pair.

Perplexity: how many neighbors each point listens to

  • You do not set each point's width σi\sigma_i by hand. You choose one number, the perplexity: roughly, how many neighbors each point should pay attention to.
  • t-SNE then widens or narrows each point's Gaussian until it has about that many neighbors: narrow in dense regions, wide in sparse ones.

The toy data with the starred point x-i at three perplexities. At perplexity 2, sigma-i is 0.11 and almost all probability goes to one neighbor; at 5, sigma-i is 0.30 and the probability spreads over its own blue group; at 12, sigma-i is 1.30, the dashed circle marking the Gaussian's width is large, and 16 percent of the probability reaches the other groups

Perplexity changes the map

  • Two clusters of 50 points each, 100 in all. Perplexity is a number of neighbors, so it must stay below the number of points: the last panel, perplexity 100, is not valid.
  • Typical range 5–50. Run a sweep and trust only what holds across it.

The same two clusters embedded at perplexity 2, 5, 30, 50 and 100: fragmented shards at low values, two clean clusters near 30 to 50, and mixed points at 100

Wattenberg, Viégas & Johnson 2016 · Distill · CC BY 4.0

Perplexity changes the map

  • Two clusters of 50 points each, 100 in all. Perplexity is a number of neighbors, so it must stay below the number of points: the last panel, perplexity 100, is not valid.
  • Typical range 5–50. Run a sweep and trust only what holds across it.

The same strip with a red box around the perplexity 5, 30 and 50 panels: in all three the two groups come out separate and never mix. The same two clusters embedded at perplexity 2, 5, 30, 50 and 100: fragmented shards at low values, two clean clusters near 30 to 50, and mixed points at 100

Stable across 5–50: two groups that never mix. That, and only that, is worth interpreting.

Wattenberg, Viégas & Johnson 2016 · Distill · CC BY 4.0

Step 2 — in the map: similarities QQ

  • In the map (in MDS, the map distances): similarity of map points with a Student-tt kernel, not a Gaussian:   qij  ∝  (1+∥y(i)−y(j)∥2)−1\;q_{ij}\;\propto\;\left(1+\lVert\mathbf{y}^{(i)}-\mathbf{y}^{(j)}\rVert^{2}\right)^{-1}
  • Why the t? A point can have many equally-near neighbors in the data, but a 2-D page fits only a few around it (crowding). The heavy tail lets the rest sit farther out.

Left, a Gaussian and a Student-t similarity curve against distance; the t curve stays higher at large distances, shaded as the heavy tail. Middle and right, 24-by-24 matrices of pair probabilities for the three-group data: P in the data and Q in the 1-D map, both with three bright diagonal blocks, one per group

Step 3 — move the map until QQ matches PP

The 1-D t-SNE run on the toy data at four moments on a common scale. At the start the points are scattered at random and the colors are mixed, mismatch 1.52; after 25 steps the three groups have formed, mismatch 0.39; after 100 steps, 0.36; after 1,000 steps the groups sit apart, mismatch 0.11

  • As in MDS: drop the points on the line at random, measure how badly QQ matches PP, nudge every point a little to reduce the mismatch, repeat.
  • The mismatch is the KL divergence, ∑i≠jpijlog⁡pijqij\sum_{i\neq j} p_{ij}\log\frac{p_{ij}}{q_{ij}}; the nudging is gradient descent. Both return when we train neural networks.
  • The random start is set by a random seed: another seed gives another map. Starting from the PCA map (PCA initialization) is more repeatable.

What the mismatch punishes

The finished 1-D map of the toy data, mismatch 0.11. Below it, the same map with one point pulled away from its group to the far end: mismatch rises by 0.38. Below that, the same map with one whole group slid up against another: mismatch rises by only 0.18

  • Tearing one point from its neighbors costs more than pushing a whole group of eight against another.
  • Close pairs (large pijp_{ij}) dominate the mismatch; far pairs barely count.
  • So t-SNE keeps neighbors, not distances. A t-SNE map is not a projection: distances, sizes and gaps in it are not distances in the data.

UMAP: a related way to build a neighborhood map

  • Connect each point to nearby points in the original space, giving stronger connections to closer neighbors.
  • Lay out the map to favor those connections, while keeping unrelated points from collapsing together.
  • n_neighbors sets the neighborhood scale; min_dist how tightly points pack.
  • Same data, same lesson: both maps keep local neighbors and tear the sheet in places.

t-SNE at perplexity 30 and UMAP with 15 neighbors applied to the same Swiss roll. Both keep neighboring colors together, and both break the sheet into pieces rather than unrolling it into one strip

UMAP's two settings

UMAP maps of the Swiss roll in a grid: columns n_neighbors 5, 15 and 50; rows min_dist 0.0 and 0.8. With few neighbors the sheet shatters into many small pieces; with more neighbors the pieces join into longer strands. With min_dist 0.0 points pack into thin tight threads; with 0.8 they spread into broad bands

  • n_neighbors plays the role of perplexity: how many neighbors each point listens to. Few: many small pieces. Many: more global shape.
  • min_dist sets how tightly points may pack: 0 gives thin, tight clumps; larger values spread them out. It changes the look, not the neighbors.
  • Same rule as t-SNE: sweep, and trust what holds across settings.

Match the method to the relationship

Method Main aim What to check
PCA Keep as much variance as a flat projection can Variance kept; did the dropped directions matter?
MDS Reproduce the dissimilarities you give it as map distances Stress (fit error); which dissimilarity you chose
Isomap MDS on distances measured along the data Shortcut edges in the neighbor graph
t-SNE Keep each point's nearest neighbors close Are neighbors kept? Stable across settings and seeds?
UMAP Keep a graph of nearest neighbors Are neighbors kept? Stable across settings and seeds?

Each map keeps only what its method keeps. Before trusting a pattern, check that the method keeps that kind of relationship.

Which one should I use?

  • Start with PCA: linear, fast, applies to new data, and tells you how much variance the 2-D view keeps.
  • To look at neighborhoods and clusters, use t-SNE or UMAP. They are close cousins; Assignment 1 uses t-SNE only to keep it short, and UMAP would teach the same lessons.
  • Whichever you use: PCA initialization, sweep perplexity (t-SNE) or n_neighbors (UMAP) and the random seed, and keep what holds across the sweep.
  • Never pick the prettiest run. If a map shows two groups, check that the images really group that way in the network's own responses, not just in the picture.

Next: back to our own data. Which features of a map can you believe?

Kobak & Linderman 2021 · Nature Biotechnology

Back to our data: same responses, two maps

Suppose we plot the same network responses twice: in the PCA map the objects overlap, in the t-SNE map they form separate islands.

  • Did the network's responses change? No: only the plotting method did.
  • Did the PCA map hide groups that are really there? Possibly: groups that differ only along the dropped directions overlap in two PCs.
  • Or did t-SNE exaggerate the groups? Possibly too: it can make clumps, and its gaps are not distances.
  • Only the original responses can tell: check each image's nearest neighbors there, or train a classifier on some images and test it on others.

A map that looks cleaner does not mean the responses carry more information.

The Bao objects, four ways — live

Which claims survive a change of map?

Wattenberg, Viégas & Johnson (2016) built toy data with known structure and swept every setting. Their examples show what a visualization can change.

What a t-SNE map can fake

Two clusters of very different spread come out the same size in t-SNE
Sizes: a tight and a spread-out cluster look the same (original, then perplexity 2 → 100)
Clusters at different true distances come out at misleading distances
Gaps: distances between clusters are not kept (original, then 2 → 100)
Random noise at low perplexity comes out as clumps
Clumps from noise at low perplexity (2 → 100, left to right)
Early in a run, shapes are pinched and change as optimization proceeds
Too few steps: shapes still changing (10 → 1,000 steps)

Assignment 1 has you test each of these on real network responses: perplexity sweeps, seeds, and what two dimensions lose.

Wattenberg, Viégas & Johnson 2016 · Distill · CC BY 4.0

A map can be stable and still distort the data

Suppose two runs give the same ten nearest neighbors for an image, but only four are among its ten nearest neighbors in the original space.

Check Result What it supports
Map versus map 10/10=1.010/10=1.0 Stable neighborhoods across these runs
Map versus original space 4/10=0.44/10=0.4 Limited neighborhood fidelity here
Six neighbors show the same object 6/10=0.66/10=0.6 purity Object grouping within this map

Repeatability, fidelity and object grouping are different questions. Assignment 1 has you compute each one.

Next: if a map cannot settle a question, what can?

From a map to a test of the representation

Question Evidence to seek
Do views of the same object stay together? Neighbors and object purity in the response space
Can object identity be read out? Accuracy on held-out images of a readout: a weighted sum of the responses, then a threshold, trained on other images
Does a model resemble human judgments or brain responses? Compare responses or dissimilarities for matched stimuli
Does learning change the representation? Compare before and after learning under the same tests

A map suggests the question; the response space answers it.

What have we learned about a representation?

  • t-SNE and UMAP emphasize local neighborhoods; distant pairs are only weakly constrained. Sizes and gaps in the map are not measurements.
  • Dimension is ambient, embedding (what PCA sees) or intrinsic (the surface's own). Isomap keeps "along-the-sheet" distances and can recover the intrinsic shape when the neighbor graph stays on the sheet.
  • A map can reveal what a two-component PCA projection hides, and it can invent structure. Stability does not establish fidelity.
  • To interpret grouping, check neighborhoods and labels in the response space. To interpret task information, test a readout on new data.

A useful map helps us formulate a claim that we can test.

Further reading

The paper you select for your research trail is required; the other research papers are optional.

Minute paper

Submit on Canvas → Minute papers → Minute Paper 5 (access code read out in class).

Write three brief points in your own words:

  1. A t-SNE map shows the objects as separate islands. Name one way to check whether those groups are really there in the network's responses.
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Appendix: extra examples

Published examples of each method (PCA, MDS, Isomap, t-SNE, UMAP), two brain examples (a ring and a torus), the Swiss-roll demo with PCA, Isomap and UMAP, a second application of UMAP, and the precise definition of perplexity.

In the wild: PCA

Premotor population activity of monkeys tapping to a beat, on the first principal components: the activity traces loops, larger for slower tempos
Monkeys, medial premotor cortex, single-neuron recordings: tapping to a beat traces loops; slower tempo, bigger loop
Neural trajectories of two monkeys making the same eight reaches, on the top principal components, before and after alignment
Monkeys, motor cortex, multi-electrode arrays: the same eight reaches in two animals match after a rotation
Gámez et al. 2019 · PLoS Biol · Safaie et al. 2023 · Nature · CC BY 4.0

In the wild: MDS

The same 92 object images arranged by MDS of human IT fMRI responses and of human similarity judgments
Humans, IT cortex (fMRI) vs similarity judgments: 92 objects
MDS of MEG response patterns to object images over time after image onset
Humans, whole-brain MEG: animals and food separate near 120 ms
Mur et al. 2013 · Front Psychol · CC BY 3.0 · Hebart et al. 2023 · eLife · CC0

In the wild: Isomap

Head-direction population activity embedded with Isomap: a ring colored by heading, with activity during sleep travelling around it
Mice, head-direction cells (anterodorsal thalamus), electrode recordings: a ring, travelled during sleep too
Isomap rings of head-direction activity in a blind control mouse and after olfactory ablation
Blind mice, head-direction cells: the ring survives; without smell it drifts from the true heading
Viejo & Peyrache 2020 · Asumbisa et al. 2022 · Nat Commun · CC BY 4.0

In the wild: t-SNE

t-SNE of 23,822 mouse cortical cells by gene expression, forming labelled islands of cell types
Mouse cortex, gene expression of 23,822 single cells: islands of cell types
A t-SNE map of fruit-fly postures and movements divided into behavioural regions such as walking, grooming and wing movements
Fruit flies, video of freely moving animals: a map of behaviours
Kobak & Berens 2019 · Nat Commun · Berman et al. 2014 · J R Soc Interface · CC BY 4.0

In the wild: UMAP

UMAP of 625 spike waveforms from monkey premotor cortex, in eight clusters with their average waveforms
Monkeys, premotor cortex, spike waveforms of 625 neurons: eight putative cell types
Activation Atlas: a grid of feature visualizations placed by UMAP of an image network layer, with animals, landscapes and objects in separate regions
No animal: one layer of an image network (InceptionV1), what its units respond to
Lee et al. 2021 · eLife · Carter et al. 2019 · Distill · CC BY 4.0

Sometimes the curved shape is the answer

Population activity of head-direction cells embedded nonlinearly: the states close into a ring, colored by the animal's heading running once around the loop
  • Head-direction cells in the mouse anterodorsal thalamus: one point per time bin, one dimension per neuron.
  • Isomap closes those points into a ring: one curved dimension, the shape a circular variable should have. A topological test (persistent homology) confirms the loop without trusting the picture.
  • Position around the ring tracks the animal's head direction, and the ring persists in REM sleep: a circuit that maintains an internal direction.
  • Language models do it too: days of the week and months sit on circles inside GPT-2 and Mistral.

PCA says six; a nonlinear map says torus

A sugared ring doughnut: a torus

Grid-cell population activity embedded in 3-D, seen from two angles: the points form a torus, a ring-shaped surface with a hole, with a five-second stretch of the rat's trajectory drawn on it in red and yellow

Persistence barcodes for the same data: one long bar in the H0 row, two long bars in the H1 row and one long bar in the H2 row, marked with arrows

  • Grid cells are neurons in the rat brain that fire whenever the animal is at the corners of a triangular grid laid over the floor. PCA puts the activity of hundreds of them in about 6 dimensions.
  • UMAP shows a torus: an intrinsically 2-D surface, as theory predicted. Persistent homology counts one piece, two loops, one cavity: a torus, tested, not eyeballed.
Gardner et al. 2022 · Nature · Fig. 1b, e · CC BY 4.0

Unroll the Swiss roll

Are vocal repertoires discrete or continuous?

UMAP maps of vocal elements from twelve species, with different degrees of apparent grouping

  • Each point represents a vocal element, described by its spectrogram.
  • The maps suggest differences in how strongly vocal elements form groups.
  • The paper also measures clusterability in UMAP space. A number makes the description precise, but still depends on the embedding.
  • What would support a claim about the animals? Check acoustic relationships, labels or behavior as well.
Sainburg, Thielk & Gentner 2020 · Fig. 8 · CC BY 4.0

Perplexity, precisely

  • For point ii, collect its neighbor probabilities from Step 1 into one distribution, Pi={pj∣i}j≠iP_i=\{p_{j\mid i}\}_{j\neq i} (one row of the conditional probabilities).
  • Its entropy, H(Pi)=−∑jpj∣ilog⁡2pj∣iH(P_i)=-\sum_j p_{j\mid i}\log_2 p_{j\mid i}, measures how evenly point ii spreads its probability: low for a narrow Gaussian, high for a wide one.
  • The perplexity is 2H(Pi)2^{H(P_i)}. t-SNE finds, for each point separately, the σi\sigma_i that makes 2H(Pi)2^{H(P_i)} equal to the value you chose (a one-dimensional search, by bisection).
  • Why "number of neighbors": if point ii spread its probability evenly over kk neighbors, H=log⁡2kH=\log_2 k and the perplexity would be exactly kk.

- The reader-friendly version linked above carries every slide in reading order, with a description of each figure - The last three lectures gave you three ways to draw a representation: MDS from dissimilarities, PCA from variance. Both are flat, linear pictures. Today asks what happens when the data do not lie on a flat surface, and what a nonlinear map can and cannot tell you

- These two questions came up again and again in Tuesday's minute papers, and they are a good way into today, so we start with them. The wording is yours, lightly shortened; reconstruction, the other frequent question, is answered on Ed - A plot has two axes because a page has two. A two-component PCA plot is a projection: the distances you see are only the parts of the true distances that lie in that plane - The Bao et al. numbers are from the Assignment 1 data: PC1 of fc6 holds 13.0% of the variance and PC2 9.0%; 85% takes 87 components and 90% takes 149. Bao et al. report 50 with their own pipeline. "Varies most over the whole data set" means variance over all 1,224 images: a direction can carry small, perfectly consistent differences, for example between two views of one object, and still be dropped - A soft count such as the participation ratio (next lecture; about 25 for fc6) is also a linear count. The intrinsic dimension can be much smaller, as the turned marlin will show in a few slides: 24 components for 90% of its variance, but one dimension, the angle - Next: before any new method, a picture of what every 2-D plot gives up

- Both questions come down to the same thing: any 2-D picture of data that live in more dimensions throws something away. To see exactly what, start with the smallest case, three dimensions drawn in two - This is the plane from the PCA lecture: v1 and v2 are its two principal directions. The dashed red lines show how far some points sit off the plane

- Now the two methods you already know draw the same points in 2-D. The middle panel is what PCA hands you: each point at its coordinates z1 and z2 on the plane. The right panel is MDS, given only the distances between the points in 3-D - The MDS map comes out turned and mirrored: distances do not change under a rotation or a reflection, so MDS has no preferred orientation. Turn it in your head and it matches the PCA map; the distances in the two maps correlate 0.99 - The green point sits 1.55 above the plane. PCA drops that height, so the point lands among the others. MDS tries to keep its true distance to every other point, and in 2-D the only way is to push it outward, about 0.6 of the points' mean distance from their center, away from where PCA puts it - Here the dropped residual holds 2.1% of the squared length, so little is lost. Treating the residual as noise is still an assumption: small-variance directions can carry the differences a task needs. MDS stress measures the same kind of loss, as the mismatch between map distances and the dissimilarities you gave it - When the data lie close to a plane and the dissimilarities are Euclidean, the two methods agree; classical MDS on Euclidean distances gives exactly the PCA map. They part ways when data curve away from any plane, which is where real images take us next

- The 100 toy points lay close to a plane. Real images usually do not. Take one photograph and change a single thing about it, its angle: the images do not move along a straight line in pixel space, they trace a curve - Real computation on one image, a marlin from the Assignment 1 stimuli: 36 frames 10 degrees apart, the object drawn at 64 pixels inside a 96-pixel square so no frame is ever cropped. Each point is a pixel vector, not a network response; PC1 and PC2 come from a PCA of these 36 images only - The percentages on the axes are each component's share of the variance, lambda_k over the sum of all eigenvalues. Computed with the SVD: writing the centered data as X = U S V^T, the covariance is V (S^2/N) V^T, so lambda_k = s_k^2 / N, and in the share the N cancels to s_k^2 / sum of s_i^2 - Why the SVD (footnote): 36 images, 9,216 dimensions. The covariance would be 9,216 by 9,216 and only 35 of its eigenvalues can be nonzero, so it is cheaper to take the SVD of the 36 by 9,216 data matrix. Assignment 2 uses the same shortcut - This is the example promised on the questions slide. Counting components to 90% gives the embedding dimension, 24 here; the intrinsic dimension is 1, because one dimension, the angle, locates every image on the loop. Today's nonlinear methods go after the second number - On these two components, frames 180 degrees apart land close together, so the loop is traced about twice

- Turning is not special. Slide the same photograph left to right and the images again trace a curve rather than a line - 29 frames, one pixel apart, with a PCA of these 29 images only, so the axes differ from the previous slide: compare the shape, not the position. Two components keep 62% of the variance; 9 are needed for 90% - Why it bends: a one-pixel shift changes different pixels depending on where the object already is, so equal steps of the knob are not equal steps in one fixed direction

- A third knob, size, gives a third curve - 25 frames from half size to 1.45 times, again with its own PCA. Two components keep 67% of the variance; 7 are needed for 90% - None of the three curves is a straight line, which is why a flat PCA plane captures only part of each. So what happens when all three changes are in the same data set?

- Real data mix many changes at once. Put all 90 images from the last three slides into a single PCA and draw them on the same axes and scale - PCA counts flat directions, not knobs, so a curved surface looks far higher-dimensional than it is. For recognition that is the problem: one object's images are spread out. The next slides ask whether a trained network gathers them - On their own, each change looked like a clean curve. Here one flat plane must serve all three, and it keeps much less of each than that change's own plane did: turning 13% instead of 34%, shifting 50% instead of 62%, enlarging 26% instead of 67%. The best-fitting planes of the three changes are 63 to 87 degrees apart: each change curves through its own, nearly unrelated directions - Curvature is what inflates the count of 36 components: a curve that bends through many directions needs many flat directions to follow it - What we would like instead is a map where each change is its own straight direction. No linear projection can give that; the rest of the lecture looks at methods that try

- Before naming what we have been seeing, try it: the demo does the last four slides for seven objects - Pick an object and a change, then drag the slider or press Play. The line beside the plot gives the variance two components keep and the number needed for 90%. Turning always needs the most, because its curve bends the most; compare the house and the vegetable with the marlin. The demo is also on the course site with the other demos

- Every change so far traced a curve. Turn and enlarge at the same time and the images fill a curved surface. That surface is what we will call a manifold - Real computation: the marlin at 24 angles and nine sizes, 216 images, drawn on the first three principal components, which keep 33% of the variance, so this is one view of the surface. The bottom of the bowl is the smallest fish; a red ring turns it at a fixed size, a green spoke enlarges it at a fixed angle - Locally flat: the gold cell, bounded by four neighboring images, is nearly flat; the surface curves only over larger distances. That is why linear tools such as PCA still work on a small patch - The Earth is the everyday example: curved, yet flat up close, and two dimensions, latitude and longitude, locate any point on it. Photo: NASA / Apollo 17, public domain - One object gives one manifold. With many objects, each gets its own, which brings us to real stimuli and a trained network

- So far, one photograph and changes we applied ourselves. The stimuli in Assignment 1 do this for real: many objects, each photographed from many viewpoints, and this time we look at a trained network's responses rather than pixels - Objects per category: animals 10, vehicles 11, faces 9, vegetables and fruit 7, houses 4, man-made objects 10 - These are also the images in the live t-SNE and UMAP demo later in the lecture

- Start with a single object: does the network keep its 24 views together? - We took the trained network's late-layer response to each of the 24 photographs of one excavator and ran PCA on those 24 vectors. PC1 and PC2 hold 19% and 12% of their variance. They are not named features such as "rotation" or "size", just the directions along which these responses vary most - Photographs from opposite sides share almost no pixels, yet the responses form one connected region: this object's manifold in the network, seen through a flat view. The shaded outline is only the smallest polygon around the dots, drawn to make the patch visible

- One object gives one patch. The real question is how patches of different objects relate - This is a new PCA of all 72 responses, so its axes differ from the previous slide; the excavator's patch looks smaller because the plot is now scaled to show the large differences between objects - Measured on these axes, the gap between the centers of the two closest patches is about eight times the average distance of a view from its own patch's center: in this layer, identity separates the responses far more than viewpoint does

- The picture we just built from a network's responses was drawn by hand fifteen years earlier, as a theory of object recognition - Left: DiCarlo and Cox's schematic of how pose, lighting and other changes move the images of one identity along a sheet. It is a conceptual model, not a surface fitted to data. Right: the three-object figure from two slides back, a trained network's late-layer responses, where each object's views form a compact patch far from the others - That separation is the goal DiCarlo and Cox set for the ventral stream of visual cortex: transform images, stage by stage, until each object's sheet is compact and apart from the others. The network shown here is one system that has partly done it. The next slide shows the starting point, pixels, where it has not been done

- DiCarlo and Cox's answer: in pixel space the sheets of different identities are tangled, and vision is the work of untangling them. Real variation (rotation in depth, lighting, background, clutter) fills pixel space with sheets that wind through each other: two photographs of different objects can be closer in pixels than two photographs of the same object - A hyperplane is the flat boundary a linear classifier draws: a line in 2-D, a plane in 3-D, everything on one side gets one label. The yellow plane fails because the two sheets wind through each other - Their argument: the visual system transforms images stage after stage until views of one person lie on one side of a simple boundary. In the network's late layer, the three objects already sat in separate patches, so a flat boundary could split them - Later in the course we will test this directly. A linear readout is a weighted sum of one layer's responses followed by a threshold, trained to tell the objects apart; used this way, to test what a layer makes available, it is called a probe. If it can tell the objects apart, that layer has untangled them - The maps in this lecture cannot do that test. A picture where the objects look well separated is a hint, not proof that a flat boundary between them exists

- This is DiCarlo and Cox's idea drawn as our own schematic: the views of each object are drawn as a ring, not fitted to data. - Left: the two rings pass through each other, so any flat plane leaves views of both objects on each side. - Right: nothing about the objects changed, only the representation. The rings no longer interlock, and a single plane splits them. - Whether a real layer does this is what a linear readout tests, which comes in the next lecture.

- We now know what we want to look at: object manifolds, and whether a representation has untangled them. But they live in thousands of dimensions, and we can only draw two - Every method in the rest of the lecture draws such a space as a 2-D map. Before trusting any of them on real data, we try them on a toy manifold where we know the right answer, the Swiss roll, whose true shape is a flat sheet. A method that cannot unroll the Swiss roll will not show us the shape of an object manifold either

- The first candidate is the method we already have, PCA - The data are sklearn's make_swiss_roll: 800 points, noise 0.05. Color is the position along the sheet; panel 2 plots each point at that position and its height, the sheet before rolling, which is what we would like a map to recover - The along-the-sheet distance, 59.5, is measured as the shortest path through the graph linking each point to its 10 nearest neighbors, the geodesic distance approximated as it will be for Isomap later. Panel 3 keeps 71% of the variance and is still wrong: variance kept is not neighborhoods kept. Its outline suggests a cylinder only because points from the front and back of the roll are drawn on top of each other - Photo of the Swiss roll cake: Tiia Monto, CC BY-SA 4.0, Wikimedia Commons

- Maybe MDS does better, since it works from distances rather than variance. It does not: the problem is which distances we give it - MDS here gets the straight-line distance between every pair of points in 3-D and looks for 2-D positions that reproduce them. On a rolled sheet those distances say that neighboring turns are close, so MDS keeps them close. The best axis of either map correlates only weakly with position along the sheet (Spearman |rho| 0.19 for PCA, 0.24 for MDS) - Classical MDS on Euclidean distances gives exactly the PCA map. The embedding step is fine; the distances it is fed are the problem, which is the clue to the fix a few slides on

- Matthew Perich gives a picture of the same problem that is easier to remember than a Swiss roll - He uses it in "Dimensionality: neuroscience's red herring?" (The Transmitter, 7 September 2026) to separate where data sit, 3-D, from what they are, a 2-D sheet - "Looking from a better angle" is what PCA does: it chooses the best flat view. The next slide shows why no view is enough

- The simplest version of the crane is an accordion, and it shows the difference between what PCA does and what we want - The folded sheet has intrinsic dimension 2 but sits in 3-D because of its folds. Pressing it flat, which is what the best flat view amounts to, stacks panels that are far apart along the paper on top of each other, as in the Swiss roll. What we want is the right-hand picture, every point where it was on the open sheet - The image is AI-generated (ChatGPT image generation) as an illustration, not a photograph of a real experiment

- The third column is the Assignment 1 data from earlier: AlexNet fc6 responses to the 1,224 Bao et al. images. Each response has 4,096 numbers; PCA needs 149 components for 90% of the variance (87 for 85%). How many dimensions the objects really vary along (identity, viewpoint, size and so on) is not known in advance, and for real data it never is. Estimating structure like this, without assuming it is flat, is the job of the nonlinear methods in the rest of the lecture - The crane and the accordion show that "the dimension" of a data set can mean different things. Perich names three - In neuroscience, ambient is the number of recorded neurons; embedding the number of principal components needed to capture a set fraction of the activity's variance, a linear count; intrinsic the dimension of the structure inside that space - A circle needs one angle to locate a point but two linear coordinates to draw it without overlap: intrinsic 1, embedding 2. The Swiss roll needs all three PCA components, embedding 3, intrinsic 2. The turned photograph needs 24 components for 90% of its variance, intrinsic 1 - Next lecture adds a fourth, effective dimensionality, a soft count of how many PCA directions carry real variance. What we need now is a method that finds the intrinsic structure

- Back to the clue from the MDS slide: the embedding was fine, the distances were wrong. So keep MDS and change how distances are measured - On a curved sheet, two points on neighboring turns are close in a straight line, but an ant walking on the sheet would have to go all the way around. That walking distance is called the geodesic distance - We do not know the sheet, only the points. Nearest neighbors are close enough that the straight line between them stays on the sheet, so a path of many short neighbor-to-neighbor steps approximates the walk. Trusting only local distances is the idea behind Isomap, and in another form behind t-SNE and UMAP

- Isomap (Tenenbaum, de Silva and Langford, 2000) is that idea written down: build the neighbor graph, measure shortest paths through it, run MDS on those path lengths - It changes the dissimilarities, not the embedding: the MDS step is exactly the one from the second lecture. On the Swiss roll it recovers position along the sheet - It works only while the neighbor graph stays on the sheet; one edge that jumps between turns glues them back together, as the slide after next shows

- Now Perich's crane can be computed. Print a word on a page, roll the page into a Swiss roll, and ask each method to show us the page - 9,000 points sampled on a page, dark where it carries ink. PCA to two components stacks the turns, so the letters pile up; Isomap with 10 neighbors unrolls the page, and its first axis follows position on the page with r = 0.99 - Perich's point: the crane is both 3-D (where it sits) and 2-D (what it is), and both descriptions are useful for different questions

- Isomap works beautifully on these examples. On real data it is fragile, which is why you will rarely see it used today - Both maps are Isomap on the same 800 points. With 10 neighbors, each point's neighbors all lie on its own stretch of the sheet, and the map recovers position along the sheet exactly (Spearman |rho| 1.00) - Why more neighbors make it worse: Isomap trusts the straight line between neighbors as a good estimate of the distance along the sheet. That holds only while neighbors are close enough for the sheet to look flat between them. With 40 neighbors, a point's farther neighbors include points on the next turn of the roll, close through the air but far along the sheet. Isomap then draws an edge across the gap between turns, a "short circuit"; 1,478 such edges appear here. Shortest paths use these shortcuts, the along-the-sheet distances collapse, and the map folds again (|rho| 0.16) - t-SNE and UMAP rely on the same local assumption, and too large a neighborhood hurts them too. The difference is that Isomap chains local steps into long distances, so one bad edge corrupts many long paths at once; t-SNE and UMAP never compute long distances, so a bad neighbor only distorts that point's local picture - The right number of neighbors depends on the data and on noise; the data must cover the manifold densely, with no holes; and all shortest paths between N points get expensive fast - Its successors keep the key idea, trust only local neighborhoods, but stop trying to reproduce long distances: t-SNE (2008) and UMAP (2018). We focus on t-SNE because Assignment 1 uses it; UMAP follows

- t-SNE keeps Isomap's lesson, trust only neighbors, and drops its fragile part: it no longer tries to reproduce long distances at all. The next slides build it in three steps

- Before the mechanics, the goal, and how it differs from MDS - Like MDS, t-SNE produces map coordinates for a given set of points. MDS preserves all the dissimilarities you supply as well as it can; t-SNE converts distances into neighbor probabilities that weight close pairs far more than distant ones, so it keeps local structure and gives up on global distances - "Neighborhood" is the precise version of "cluster". t-SNE does not run a clustering algorithm: islands appear because it keeps each point's neighbors close - To count islands in a finished map you need a rule. The simplest is single linkage: two points are on the same island if a chain of points links them with no gap wider than a threshold. The count depends on the threshold, because gaps in a t-SNE map are not reliable distances - Throughout the next slides, distances and P are measured in the data; similarities Q in the map; only the map points move. The toy data: three groups of eight points in 2-D, mapped to one dimension so the result fits on a line. The inputs can be anything with one vector per item: pixels, neurons, voxels or a network layer a^[l] - Unlike PCA, t-SNE gives no fixed projection: the scikit-learn implementation optimizes the map coordinates directly, and new points need a new run

- Step 1 turns the distances in the data into statements about neighbors - For each point x^(i) we want a probability over all other points: how likely it would pick x^(j) as its neighbor. A Gaussian centered on x^(i) gives exactly that shape, high nearby and falling smoothly with distance; dividing by the sum over j makes the values add up to 1. p_(j|i) is conditional, and not symmetric when the two points sit in regions of different density - Compared with MDS, this warps the space: small distances become large probabilities whose differences matter, while all large distances collapse to near zero and become interchangeable. P plays the role of the dissimilarity matrix, but keeps only local detail - Computed from the data: every distance and every p_(j|i). Not computed: the width sigma_i of each point's Gaussian, one per point; the next slide shows how it is really set - In the figure, x^(i) is the starred point and each other point is drawn larger the more probability it receives. With perplexity 5 the width is 0.30 and all of x^(i)'s probability goes to its own group. Symmetrizing gives one probability per pair, p_ij, adding up to 1: the quantity the map will try to match

- The previous slide gave every point its own width sigma_i and left open how it is chosen. With N points, tuning N widths by hand would be hopeless. t-SNE asks for a single number instead, the perplexity, and sets every width from it - Read perplexity as the number of neighbors a point listens to. Perplexity 2: the starred point's Gaussian is so narrow (0.11) that nearly all its probability goes to one neighbor. Perplexity 5: the width grows to 0.30 and the probability spreads over its own group of eight. Perplexity 12: the width is 1.30 and 16% of the probability now reaches the other two groups - Low perplexity tends to break data into small clumps; high perplexity tends to merge groups. Perplexity must be smaller than the number of points. The precise definition, through the entropy of each point's neighbor probabilities, is in the appendix

- What those widths do to a real t-SNE map: Wattenberg, Viégas and Johnson's two clusters of 50 points, embedded at five perplexities, each run for 5,000 steps - At perplexity 2 each point listens to one or two neighbors and the clusters shatter; at 30 and 50 they come out clean; at 100, with only 100 points, every point would need all others as neighbors, and the clusters mix - 5 to 50 is a starting range, not a guarantee: on some data groups appear at one value and dissolve at the next. Then the honest conclusion is that t-SNE shows no stable grouping, not that one of the pictures is right

- So which features of these maps should you believe? Only what holds across the sweep - Across 5, 30 and 50 the two groups never mix: that is the structure to trust. The small clumps at 5, and the sizes and spacing of the groups, change from panel to panel, so they should not be interpreted. Initialization and the random seed change the map too, and the same rule applies to them

- Back to the mechanics. Step 1 gave the target, P, from the data. Step 2 defines the matching quantity in the map, Q, and it does it with a different kernel (the curve that turns a distance into a similarity; Step 1's was a Gaussian) - Crowding, concretely: on a flat page you can place at most 3 points all the same distance apart (a triangle); in 3-D, 4; in 10-D, 11. High-dimensional data give each point many neighbors that are all about equally near, in many different directions - On a 2-D page there is no room for all of them at that distance, so the map has to compromise. With a Gaussian in the map, a pair that is moderately far apart in the data gets almost no similarity unless it is drawn close, so the optimizer pulls those points closer to each other than they really are. Repeated across the whole data set, every point is pulled toward its moderately distant neighbors, the map shrinks toward its center, and the empty space between clusters disappears: different clusters run into each other and merge - The fix, as a worked example. Suppose a pair of points has a small similarity in the data, 0.011, as moderately distant pairs do. The map has to give that pair the same similarity. With a Gaussian in the map, similarity 0.011 is reached at distance 3. With the t kernel, which falls off much more slowly, the same 0.011 is reached only at distance 9.5. So under the t kernel the pair must be drawn about three times farther apart to get the similarity it should. Moderately distant pairs are therefore placed well away from each other, the map has room, and clusters keep the gaps between them. For close pairs the two kernels almost agree (similarity 0.1: distance 2.2 versus 3), so neighborhoods are barely affected - (This ignores the normalization of Q over all pairs, which shifts the numbers but not the comparison.) - The matrices are real: P from the three groups at perplexity 5, Q from the 1-D map. They look almost identical, and they should: the finished map was fitted to make Q match P, and 98% of Q's mass lies within groups, as in P. The effect of the heavy tail lives in the tiny values between groups, too small to see on this color scale. It matters for the layout, not for how the matrices look. The heavy tail relieves crowding; it does not remove all distortion

- With P from the data and Q from the map, Step 3 moves the map points until the two match, the same way MDS reduces its stress - In MDS you throw the points on the board, measure how far the map distances are from the dissimilarities, move each point a little to lower the error, and repeat. t-SNE does exactly this; only the error is measured differently - The figure is a real run on the toy data. Scattered start, colors mixed: mismatch 1.52. After 25 steps the groups have formed (0.39); after 1,000 they sit well apart (0.11) - The KL divergence is zero when Q equals P and grows as they differ; it returns, with the cross-entropy, when we train neural networks. Gradient descent moves each map point a small step in the direction that lowers the mismatch fastest; it is also how every network is trained - A fixed number of steps does not guarantee the map has settled (see stopping, below) - Why the seed matters: the mismatch has many local minima, so where the points start decides which one the run ends in. Two seeds can give maps with the same groups in different places, or even different groups. Starting from the first two principal components gives every run the same start and keeps the global arrangement closer to the data; that is PCA initialization, and it is the default in recent scikit-learn - When does t-SNE stop? Usually after a fixed number of steps: 1,000 by default in scikit-learn (max_iter). It stops earlier only if the gradient becomes essentially zero (min_grad_norm, 1e-7) or if the mismatch has not improved for 300 steps (n_iter_without_progress). None of these guarantees the map has settled, which is why Assignment 1 checks different budgets and seeds - Standard settings: P exaggerated for the first 100 steps, momentum, 1,000 steps. This run starts from a spread-out layout so the first frame is visible; the usual start is a tighter random layout

- The mismatch is not equally strict about everything. Here is what it punishes and what it lets go, and why that matters for reading any t-SNE map - Both changes start from the finished map. Moving one point from its group to the far end raises the mismatch by 0.38; sliding a whole group of eight against the next group, so that many far pairs become close, raises it by only 0.18 - Each pair counts in proportion to p_ij. Neighbors have large p_ij, so pulling them apart is expensive; pairs from different groups have p_ij near zero, so placing them close or far barely matters. van der Maaten and Hinton state this directly. The terms are coupled through one normalization of Q, so the figure shows the whole mismatch, not single terms - In PCA and MDS the map approximates real distances: twice as far in the map is roughly twice as far in the data. A t-SNE map gives that up. How far apart two groups sit, how large a group looks and how much space surrounds it are side effects of the optimization; to compare distances, sizes or spreads, go back to the response space

- UMAP is the other method you will meet in papers. It rests on the same idea, keep neighbors, built differently - UMAP builds a weighted graph of each point's nearest neighbors and lays it out in 2-D; t-SNE matches neighbor probabilities. The results are close: on this roll 87% of each point's 10 nearest neighbors survive in the t-SNE map and 78% in the UMAP map, and both cut the sheet into pieces, good local fidelity, poor global shape - UMAP is often described as better at global structure. Kobak and Linderman (2021) showed that most of that difference comes from how the two are usually initialized; started the same way, they behave much alike. UMAP is faster on very large data sets

- Like t-SNE, UMAP has settings that change the picture, and the same rule applies - All six maps are UMAP on the same 800-point Swiss roll, seed 0. n_neighbors plays the role of perplexity: with 5 the sheet breaks into many pieces, with 50 longer stretches stay connected - min_dist only sets how close points may sit in the map. At 0.0 they collapse into thin threads that look like sharp clusters; at 0.8 the same neighborhoods spread into bands. The share of each point's 10 nearest neighbors kept stays between 71% and 86% across all six: the look changes far more than the neighborhoods. Tight UMAP clusters in a paper may partly reflect a small min_dist

- We now have five methods. What separates them is which relationship each one keeps, and so what you should check before reading its map - PCA axes are variance directions, not automatically meaningful dimensions. MDS axes have no unique orientation. For t-SNE and UMAP, axis values and island spacing mean nothing in the original units. Checks of a map are useful as long as you know what each one establishes

- In practice, which one should you reach for? - PCA first, always: it is deterministic, applies to new data, and its explained variance tells you how much a 2-D picture can show. If two components already keep most of the variance, a nonlinear map may add little - Published comparisons that favor t-SNE or UMAP usually reflect different default settings. UMAP is faster on very large data and has a transform for new points; t-SNE is older and better studied. Assignment 1 uses t-SNE so you learn one method well; nothing in it would change with UMAP - A map is a hypothesis generator. Picking, among many runs, the one with the cleanest groups is the visual version of p-hacking: report the sweep and check claims in the response space - The appendix shows each method at work in published neuroscience and AI, two examples per method

- Back to our own data, and to the question every map raises: which of its features can you believe? The same network responses can give two very different-looking maps - Nothing about the network changed between the two maps. Both are 2-D summaries of the same responses: PCA keeps the directions of largest variance, t-SNE keeps each point's nearest neighbors. A difference between them is a fact about the methods, not about the network - Why these two checks: nearest neighbors in the full responses use every dimension, not just the two a map keeps, so they cannot be hidden by PCA or created by t-SNE. A classifier tested on images it was not trained on tells you whether the groups can actually be read out, rather than only seen. Assignment 1 runs these checks on the maps; Assignment 2 tests what can be read out and how responses compare with human judgments and neural recordings

- See it happen, on the same objects we looked at with PCA: the Assignment 1 images, all 51 objects with 8 of their 24 views each, 408 images. Each point is AlexNet fc6's response to one image (4,096 numbers), reduced to 50 principal components, which keep 81% of the variance - t-SNE runs live in the browser from these 50 numbers per image; UMAP (with n_neighbors 5, 15 or 50), PCA and MDS are precomputed on the same points. Run t-SNE, change the seed and run it again: the islands land in different places and with different sizes. Move the perplexity and the same data become many small islands or a few large ones - Hover over any image: it appears large in the corner, and lines join it to its 10 nearest neighbors in the original 50-dimensional responses. If the lines reach far across the map, the map has not kept that image's neighborhood - The two meters are the checks from Assignment 1. Neighbor agreement: for each image, the share of its 10 nearest neighbors in 50-D that are still among its 10 nearest in the map (66% for t-SNE at perplexity 30, 62% for UMAP). Object purity: the share of its map neighbors that are other views of the same object (32% and 31%). k is the neighborhood size used by both meters, not a number of components

- The demo showed that the picture changes with the seed and the settings. How do we know which of its features to believe? Wattenberg, Viégas and Johnson answered that with toy data whose true structure is known - Their article, How to Use t-SNE Effectively (Distill), is required reading for Assignment 1, which tests its warnings on object representations

- Four ways a t-SNE map can show you something that is not in the data. Each follows from the mechanics we just built - Sizes: the per-point widths sigma_i expand dense groups and shrink sparse ones, so a tight cluster and a diffuse one look the same size. Measure spread in the response space - Gaps: the mismatch barely penalizes getting far pairs wrong, so the distance between two islands, and even their order, can change between runs. Compare group distances in the response space - Clumps from noise: at low perplexity, structureless data break into clumps. Ask what the map of structureless data would look like - Too few steps: early in a run, pinched and stretched shapes are optimization in progress. A new seed gives a new picture; check that a claim survives re-runs

- A natural defense is repeatability: if every run shows the same neighbors, the map must be right. It need not be - The numbers are a constructed example, not assignment results. Use the same images, neighborhood size and distance in each comparison, and exclude the query image itself. Object purity can be high even when the particular neighboring views change, and none of these checks measures recognition on new images

- If a map cannot settle a question, what can? Each question about a representation has a test in the original response space - A simple readout tests whether task information is easy to access in a representation; a more powerful readout may add computation of its own. Similar accuracy does not mean similar geometry, and similar maps establish neither. Assignment 2 is built around these tests

- Putting the lecture together - The next lecture takes the geometry of population responses (of neurons or network units) as the object of study: how many dimensions a representation really uses, and how the shape of object manifolds relates to what a readout can do. Later, with recurrent networks, the question becomes how a circuit moves through its population states

- For going further, the required reading first, then papers where these methods answer a real question - Chaudhuri, Gerçek, Pandey, Peyrache & Fiete, The intrinsic attractor manifold and population dynamics of a canonical cognitive circuit across waking and sleep, Nature Neuroscience 2019: what evidence connects a population map to a circuit mechanism? - Gardner, Hermansen, Pachitariu, Burak, Baas, Dunn, Moser & Moser, Toroidal topology of population activity in grid cells, Nature 2022 - Nieh, Schottdorf, Freeman et al., Geometry of abstract learned knowledge in the hippocampus, Nature 2021: position and accumulated evidence on one low-dimensional manifold - Engels, Michaud, Liao, Gurnee & Tegmark, Not all language model features are one-dimensionally linear, ICLR 2025 - Chari & Pachter, The specious art of single-cell genomics, PLOS Computational Biology 2023: a critique of reading t-SNE and UMAP maps

- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 5)

These slides were not needed in class; they are here for reference while working on the assignment.

- Examples from the literature, one method per slide, starting with PCA; each caption gives the species, the brain area and the kind of recording - Gámez, Mendoza, Prado, Betancourt and Merchant 2019, PLoS Biology: premotor population activity during rhythmic tapping, on its first principal components; each tempo is a loop, and the loop grows with the interval - Safaie, Chang, Park, Miller, Dyer, Glaser and Gallego 2023, Nature: population activity of different monkeys doing the same reaches, each reduced with PCA; after a rotation (canonical correlation analysis) the dynamics match across animals

- MDS is the workhorse for comparing brain and behavior - Mur, Meys, Bodurka, Goebel, Bandettini and Kriegeskorte 2013, Frontiers in Psychology: the same 92 objects arranged by MDS of human IT response patterns and of similarity judgments; the judgments are more sharply categorical - Hebart, Contier, Teichmann, Rockter, Zheng, Kidder, Corriveau, Vaziri-Pashkam and Baker 2023, eLife (THINGS-data): MDS of MEG patterns at successive times after image onset

- Isomap, despite its fragility, found one of the cleanest results in systems neuroscience - Viejo and Peyrache 2020, Nature Communications: head-direction population activity embedded with Isomap forms a ring; during sleep the activity keeps moving around it - Asumbisa, Peyrache and Trenholm 2022, Nature Communications: the ring survives in blind mice; after olfactory ablation it no longer lines up with the actual heading

- t-SNE made cell types and behavior visible at a glance - Kobak and Berens 2019, Nature Communications, "The art of using t-SNE for single-cell transcriptomics": 23,822 mouse cortical cells placed by their gene expression; the paper also compares PCA and MDS of the same data - Berman, Choi, Bialek and Shaevitz 2014, Journal of the Royal Society Interface: every moment of a freely moving fly, described by its posture dynamics, placed with t-SNE; regions correspond to stereotyped behaviors. Only a low-resolution copy is openly available

- UMAP, from spike shapes to the inside of an image network - Lee, Mendoza, Sathish, Kiani, Kanwisher and Chandrasekaran 2021, eLife (WaveMAP): UMAP of extracellular spike waveforms separates putative cell types - Carter, Armstrong, Schubert, Johnson and Olah 2019, Distill, "Activation Atlas": one InceptionV1 layer laid out with UMAP and shown as feature visualizations; a gridded summary of the layout, not a raw scatter - For each of these maps, the checks from the lecture apply: which features hold across settings and seeds, and which are confirmed in the original data

- The study combines manifold characterization, decoding and comparison across waking and sleep. The ring plot alone is not the evidence for a circuit mechanism - Isomap is the method from earlier today: nearest-neighbor graph, path lengths, then MDS. The paper uses it to visualize and to reduce dimension before its topological analysis, which counts loops (a Betti-1 barcode) rather than reading them off a picture - A circle can also be preserved by a two-dimensional PCA projection; nonlinear methods do not automatically preserve loops - Engels, Michaud, Liao, Gurnee and Tegmark, Not All Language Model Features Are One-Dimensionally Linear: a cyclic variable in a trained network ends up on a circle, as it does in the thalamus - The further scientific question is how recurrent interactions generate and maintain this pattern of population activity

- Gardner, Hermansen, Pachitariu, Burak, Baas, Dunn, Moser and Moser, Toroidal topology of population activity in grid cells, Nature 2022. Hundreds of grid cells recorded at once in one module; the activity is first reduced with PCA to six dimensions and then embedded with UMAP - A grid cell fires on a repeating triangular lattice as the animal runs; moving along either lattice direction eventually brings the cell back to the same state, so position on the lattice is two periodic variables. Two periodic variables make a torus, which continuous-attractor models predicted decades before it was seen - The barcodes are the test the picture needs. Each bar is a feature that persists as the neighborhood radius grows: one long H0 bar means one connected piece, two long H1 bars two independent loops, one long H2 bar one enclosed cavity. That count (1, 2, 1) is a torus's - The head-direction ring on the previous slide is the one-dimensional case of the same idea

Drag to rotate the roll, then drag the unroll slider: every point moves in a straight line from its 3-D position to its place on the flat sheet, so you see which points were neighbors along the sheet all along. The two probes start one turn apart: close in 3-D (the dashed straight segment), far along the sheet (the green path, the shortest route through the 10-nearest-neighbor graph). The three panels on the right embed the same 800 points: PCA keeps as much variance as a flat projection can and so folds the sheet; Isomap approximates along-the-sheet distances and here unrolls it; UMAP favors local neighborhoods, which keeps nearby points together but does not guarantee an unrolled sheet. Isomap is not in Assignment 1.

This is a useful exploratory application across species. The reported Hopkins statistic measures clusterability after UMAP, so it is not an independent validation in the original acoustic space. The choice of acoustic representation and the sampling of vocal elements also affect the interpretation.

- The capital P in Step 3 is different: it is the symmetric matrix of all p_ij, formed after this step by averaging p_(j|i) and p_(i|j). P_i here is one row of the conditional probabilities before symmetrizing - Because every point is tuned to the same perplexity, every point ends up with about the same effective number of neighbors, whatever the density around it. This is also why t-SNE equalizes the apparent size of dense and sparse clusters, one of the pitfalls later in the lecture