- The reader-friendly version linked above carries every slide in reading order, with a description of each figure - The last three lectures gave you three ways to draw a representation: MDS from dissimilarities, PCA from variance. Both are flat, linear pictures. Today asks what happens when the data do not lie on a flat surface, and what a nonlinear map can and cannot tell you
- These two questions came up again and again in Tuesday's minute papers, and they are a good way into today, so we start with them. The wording is yours, lightly shortened; reconstruction, the other frequent question, is answered on Ed - A plot has two axes because a page has two. A two-component PCA plot is a projection: the distances you see are only the parts of the true distances that lie in that plane - The Bao et al. numbers are from the Assignment 1 data: PC1 of fc6 holds 13.0% of the variance and PC2 9.0%; 85% takes 87 components and 90% takes 149. Bao et al. report 50 with their own pipeline. "Varies most over the whole data set" means variance over all 1,224 images: a direction can carry small, perfectly consistent differences, for example between two views of one object, and still be dropped - A soft count such as the participation ratio (next lecture; about 25 for fc6) is also a linear count. The intrinsic dimension can be much smaller, as the turned marlin will show in a few slides: 24 components for 90% of its variance, but one dimension, the angle - Next: before any new method, a picture of what every 2-D plot gives up
- Both questions come down to the same thing: any 2-D picture of data that live in more dimensions throws something away. To see exactly what, start with the smallest case, three dimensions drawn in two - This is the plane from the PCA lecture: v1 and v2 are its two principal directions. The dashed red lines show how far some points sit off the plane
- Now the two methods you already know draw the same points in 2-D. The middle panel is what PCA hands you: each point at its coordinates z1 and z2 on the plane. The right panel is MDS, given only the distances between the points in 3-D - The MDS map comes out turned and mirrored: distances do not change under a rotation or a reflection, so MDS has no preferred orientation. Turn it in your head and it matches the PCA map; the distances in the two maps correlate 0.99 - The green point sits 1.55 above the plane. PCA drops that height, so the point lands among the others. MDS tries to keep its true distance to every other point, and in 2-D the only way is to push it outward, about 0.6 of the points' mean distance from their center, away from where PCA puts it - Here the dropped residual holds 2.1% of the squared length, so little is lost. Treating the residual as noise is still an assumption: small-variance directions can carry the differences a task needs. MDS stress measures the same kind of loss, as the mismatch between map distances and the dissimilarities you gave it - When the data lie close to a plane and the dissimilarities are Euclidean, the two methods agree; classical MDS on Euclidean distances gives exactly the PCA map. They part ways when data curve away from any plane, which is where real images take us next
- The 100 toy points lay close to a plane. Real images usually do not. Take one photograph and change a single thing about it, its angle: the images do not move along a straight line in pixel space, they trace a curve - Real computation on one image, a marlin from the Assignment 1 stimuli: 36 frames 10 degrees apart, the object drawn at 64 pixels inside a 96-pixel square so no frame is ever cropped. Each point is a pixel vector, not a network response; PC1 and PC2 come from a PCA of these 36 images only - The percentages on the axes are each component's share of the variance, lambda_k over the sum of all eigenvalues. Computed with the SVD: writing the centered data as X = U S V^T, the covariance is V (S^2/N) V^T, so lambda_k = s_k^2 / N, and in the share the N cancels to s_k^2 / sum of s_i^2 - Why the SVD (footnote): 36 images, 9,216 dimensions. The covariance would be 9,216 by 9,216 and only 35 of its eigenvalues can be nonzero, so it is cheaper to take the SVD of the 36 by 9,216 data matrix. Assignment 2 uses the same shortcut - This is the example promised on the questions slide. Counting components to 90% gives the embedding dimension, 24 here; the intrinsic dimension is 1, because one dimension, the angle, locates every image on the loop. Today's nonlinear methods go after the second number - On these two components, frames 180 degrees apart land close together, so the loop is traced about twice
- Turning is not special. Slide the same photograph left to right and the images again trace a curve rather than a line - 29 frames, one pixel apart, with a PCA of these 29 images only, so the axes differ from the previous slide: compare the shape, not the position. Two components keep 62% of the variance; 9 are needed for 90% - Why it bends: a one-pixel shift changes different pixels depending on where the object already is, so equal steps of the knob are not equal steps in one fixed direction
- A third knob, size, gives a third curve - 25 frames from half size to 1.45 times, again with its own PCA. Two components keep 67% of the variance; 7 are needed for 90% - None of the three curves is a straight line, which is why a flat PCA plane captures only part of each. So what happens when all three changes are in the same data set?
- Real data mix many changes at once. Put all 90 images from the last three slides into a single PCA and draw them on the same axes and scale - PCA counts flat directions, not knobs, so a curved surface looks far higher-dimensional than it is. For recognition that is the problem: one object's images are spread out. The next slides ask whether a trained network gathers them - On their own, each change looked like a clean curve. Here one flat plane must serve all three, and it keeps much less of each than that change's own plane did: turning 13% instead of 34%, shifting 50% instead of 62%, enlarging 26% instead of 67%. The best-fitting planes of the three changes are 63 to 87 degrees apart: each change curves through its own, nearly unrelated directions - Curvature is what inflates the count of 36 components: a curve that bends through many directions needs many flat directions to follow it - What we would like instead is a map where each change is its own straight direction. No linear projection can give that; the rest of the lecture looks at methods that try
- Before naming what we have been seeing, try it: the demo does the last four slides for seven objects - Pick an object and a change, then drag the slider or press Play. The line beside the plot gives the variance two components keep and the number needed for 90%. Turning always needs the most, because its curve bends the most; compare the house and the vegetable with the marlin. The demo is also on the course site with the other demos
- Every change so far traced a curve. Turn and enlarge at the same time and the images fill a curved surface. That surface is what we will call a manifold - Real computation: the marlin at 24 angles and nine sizes, 216 images, drawn on the first three principal components, which keep 33% of the variance, so this is one view of the surface. The bottom of the bowl is the smallest fish; a red ring turns it at a fixed size, a green spoke enlarges it at a fixed angle - Locally flat: the gold cell, bounded by four neighboring images, is nearly flat; the surface curves only over larger distances. That is why linear tools such as PCA still work on a small patch - The Earth is the everyday example: curved, yet flat up close, and two dimensions, latitude and longitude, locate any point on it. Photo: NASA / Apollo 17, public domain - One object gives one manifold. With many objects, each gets its own, which brings us to real stimuli and a trained network
- So far, one photograph and changes we applied ourselves. The stimuli in Assignment 1 do this for real: many objects, each photographed from many viewpoints, and this time we look at a trained network's responses rather than pixels - Objects per category: animals 10, vehicles 11, faces 9, vegetables and fruit 7, houses 4, man-made objects 10 - These are also the images in the live t-SNE and UMAP demo later in the lecture
- Start with a single object: does the network keep its 24 views together? - We took the trained network's late-layer response to each of the 24 photographs of one excavator and ran PCA on those 24 vectors. PC1 and PC2 hold 19% and 12% of their variance. They are not named features such as "rotation" or "size", just the directions along which these responses vary most - Photographs from opposite sides share almost no pixels, yet the responses form one connected region: this object's manifold in the network, seen through a flat view. The shaded outline is only the smallest polygon around the dots, drawn to make the patch visible
- One object gives one patch. The real question is how patches of different objects relate - This is a new PCA of all 72 responses, so its axes differ from the previous slide; the excavator's patch looks smaller because the plot is now scaled to show the large differences between objects - Measured on these axes, the gap between the centers of the two closest patches is about eight times the average distance of a view from its own patch's center: in this layer, identity separates the responses far more than viewpoint does
- The picture we just built from a network's responses was drawn by hand fifteen years earlier, as a theory of object recognition - Left: DiCarlo and Cox's schematic of how pose, lighting and other changes move the images of one identity along a sheet. It is a conceptual model, not a surface fitted to data. Right: the three-object figure from two slides back, a trained network's late-layer responses, where each object's views form a compact patch far from the others - That separation is the goal DiCarlo and Cox set for the ventral stream of visual cortex: transform images, stage by stage, until each object's sheet is compact and apart from the others. The network shown here is one system that has partly done it. The next slide shows the starting point, pixels, where it has not been done
- DiCarlo and Cox's answer: in pixel space the sheets of different identities are tangled, and vision is the work of untangling them. Real variation (rotation in depth, lighting, background, clutter) fills pixel space with sheets that wind through each other: two photographs of different objects can be closer in pixels than two photographs of the same object - A hyperplane is the flat boundary a linear classifier draws: a line in 2-D, a plane in 3-D, everything on one side gets one label. The yellow plane fails because the two sheets wind through each other - Their argument: the visual system transforms images stage after stage until views of one person lie on one side of a simple boundary. In the network's late layer, the three objects already sat in separate patches, so a flat boundary could split them - Later in the course we will test this directly. A linear readout is a weighted sum of one layer's responses followed by a threshold, trained to tell the objects apart; used this way, to test what a layer makes available, it is called a probe. If it can tell the objects apart, that layer has untangled them - The maps in this lecture cannot do that test. A picture where the objects look well separated is a hint, not proof that a flat boundary between them exists
- This is DiCarlo and Cox's idea drawn as our own schematic: the views of each object are drawn as a ring, not fitted to data. - Left: the two rings pass through each other, so any flat plane leaves views of both objects on each side. - Right: nothing about the objects changed, only the representation. The rings no longer interlock, and a single plane splits them. - Whether a real layer does this is what a linear readout tests, which comes in the next lecture.
- We now know what we want to look at: object manifolds, and whether a representation has untangled them. But they live in thousands of dimensions, and we can only draw two - Every method in the rest of the lecture draws such a space as a 2-D map. Before trusting any of them on real data, we try them on a toy manifold where we know the right answer, the Swiss roll, whose true shape is a flat sheet. A method that cannot unroll the Swiss roll will not show us the shape of an object manifold either
- The first candidate is the method we already have, PCA - The data are sklearn's make_swiss_roll: 800 points, noise 0.05. Color is the position along the sheet; panel 2 plots each point at that position and its height, the sheet before rolling, which is what we would like a map to recover - The along-the-sheet distance, 59.5, is measured as the shortest path through the graph linking each point to its 10 nearest neighbors, the geodesic distance approximated as it will be for Isomap later. Panel 3 keeps 71% of the variance and is still wrong: variance kept is not neighborhoods kept. Its outline suggests a cylinder only because points from the front and back of the roll are drawn on top of each other - Photo of the Swiss roll cake: Tiia Monto, CC BY-SA 4.0, Wikimedia Commons
- Maybe MDS does better, since it works from distances rather than variance. It does not: the problem is which distances we give it - MDS here gets the straight-line distance between every pair of points in 3-D and looks for 2-D positions that reproduce them. On a rolled sheet those distances say that neighboring turns are close, so MDS keeps them close. The best axis of either map correlates only weakly with position along the sheet (Spearman |rho| 0.19 for PCA, 0.24 for MDS) - Classical MDS on Euclidean distances gives exactly the PCA map. The embedding step is fine; the distances it is fed are the problem, which is the clue to the fix a few slides on
- Matthew Perich gives a picture of the same problem that is easier to remember than a Swiss roll - He uses it in "Dimensionality: neuroscience's red herring?" (The Transmitter, 7 September 2026) to separate where data sit, 3-D, from what they are, a 2-D sheet - "Looking from a better angle" is what PCA does: it chooses the best flat view. The next slide shows why no view is enough
- The simplest version of the crane is an accordion, and it shows the difference between what PCA does and what we want - The folded sheet has intrinsic dimension 2 but sits in 3-D because of its folds. Pressing it flat, which is what the best flat view amounts to, stacks panels that are far apart along the paper on top of each other, as in the Swiss roll. What we want is the right-hand picture, every point where it was on the open sheet - The image is AI-generated (ChatGPT image generation) as an illustration, not a photograph of a real experiment
- The third column is the Assignment 1 data from earlier: AlexNet fc6 responses to the 1,224 Bao et al. images. Each response has 4,096 numbers; PCA needs 149 components for 90% of the variance (87 for 85%). How many dimensions the objects really vary along (identity, viewpoint, size and so on) is not known in advance, and for real data it never is. Estimating structure like this, without assuming it is flat, is the job of the nonlinear methods in the rest of the lecture - The crane and the accordion show that "the dimension" of a data set can mean different things. Perich names three - In neuroscience, ambient is the number of recorded neurons; embedding the number of principal components needed to capture a set fraction of the activity's variance, a linear count; intrinsic the dimension of the structure inside that space - A circle needs one angle to locate a point but two linear coordinates to draw it without overlap: intrinsic 1, embedding 2. The Swiss roll needs all three PCA components, embedding 3, intrinsic 2. The turned photograph needs 24 components for 90% of its variance, intrinsic 1 - Next lecture adds a fourth, effective dimensionality, a soft count of how many PCA directions carry real variance. What we need now is a method that finds the intrinsic structure
- Back to the clue from the MDS slide: the embedding was fine, the distances were wrong. So keep MDS and change how distances are measured - On a curved sheet, two points on neighboring turns are close in a straight line, but an ant walking on the sheet would have to go all the way around. That walking distance is called the geodesic distance - We do not know the sheet, only the points. Nearest neighbors are close enough that the straight line between them stays on the sheet, so a path of many short neighbor-to-neighbor steps approximates the walk. Trusting only local distances is the idea behind Isomap, and in another form behind t-SNE and UMAP
- Isomap (Tenenbaum, de Silva and Langford, 2000) is that idea written down: build the neighbor graph, measure shortest paths through it, run MDS on those path lengths - It changes the dissimilarities, not the embedding: the MDS step is exactly the one from the second lecture. On the Swiss roll it recovers position along the sheet - It works only while the neighbor graph stays on the sheet; one edge that jumps between turns glues them back together, as the slide after next shows
- Now Perich's crane can be computed. Print a word on a page, roll the page into a Swiss roll, and ask each method to show us the page - 9,000 points sampled on a page, dark where it carries ink. PCA to two components stacks the turns, so the letters pile up; Isomap with 10 neighbors unrolls the page, and its first axis follows position on the page with r = 0.99 - Perich's point: the crane is both 3-D (where it sits) and 2-D (what it is), and both descriptions are useful for different questions
- Isomap works beautifully on these examples. On real data it is fragile, which is why you will rarely see it used today - Both maps are Isomap on the same 800 points. With 10 neighbors, each point's neighbors all lie on its own stretch of the sheet, and the map recovers position along the sheet exactly (Spearman |rho| 1.00) - Why more neighbors make it worse: Isomap trusts the straight line between neighbors as a good estimate of the distance along the sheet. That holds only while neighbors are close enough for the sheet to look flat between them. With 40 neighbors, a point's farther neighbors include points on the next turn of the roll, close through the air but far along the sheet. Isomap then draws an edge across the gap between turns, a "short circuit"; 1,478 such edges appear here. Shortest paths use these shortcuts, the along-the-sheet distances collapse, and the map folds again (|rho| 0.16) - t-SNE and UMAP rely on the same local assumption, and too large a neighborhood hurts them too. The difference is that Isomap chains local steps into long distances, so one bad edge corrupts many long paths at once; t-SNE and UMAP never compute long distances, so a bad neighbor only distorts that point's local picture - The right number of neighbors depends on the data and on noise; the data must cover the manifold densely, with no holes; and all shortest paths between N points get expensive fast - Its successors keep the key idea, trust only local neighborhoods, but stop trying to reproduce long distances: t-SNE (2008) and UMAP (2018). We focus on t-SNE because Assignment 1 uses it; UMAP follows
- t-SNE keeps Isomap's lesson, trust only neighbors, and drops its fragile part: it no longer tries to reproduce long distances at all. The next slides build it in three steps
- Before the mechanics, the goal, and how it differs from MDS - Like MDS, t-SNE produces map coordinates for a given set of points. MDS preserves all the dissimilarities you supply as well as it can; t-SNE converts distances into neighbor probabilities that weight close pairs far more than distant ones, so it keeps local structure and gives up on global distances - "Neighborhood" is the precise version of "cluster". t-SNE does not run a clustering algorithm: islands appear because it keeps each point's neighbors close - To count islands in a finished map you need a rule. The simplest is single linkage: two points are on the same island if a chain of points links them with no gap wider than a threshold. The count depends on the threshold, because gaps in a t-SNE map are not reliable distances - Throughout the next slides, distances and P are measured in the data; similarities Q in the map; only the map points move. The toy data: three groups of eight points in 2-D, mapped to one dimension so the result fits on a line. The inputs can be anything with one vector per item: pixels, neurons, voxels or a network layer a^[l] - Unlike PCA, t-SNE gives no fixed projection: the scikit-learn implementation optimizes the map coordinates directly, and new points need a new run
- Step 1 turns the distances in the data into statements about neighbors - For each point x^(i) we want a probability over all other points: how likely it would pick x^(j) as its neighbor. A Gaussian centered on x^(i) gives exactly that shape, high nearby and falling smoothly with distance; dividing by the sum over j makes the values add up to 1. p_(j|i) is conditional, and not symmetric when the two points sit in regions of different density - Compared with MDS, this warps the space: small distances become large probabilities whose differences matter, while all large distances collapse to near zero and become interchangeable. P plays the role of the dissimilarity matrix, but keeps only local detail - Computed from the data: every distance and every p_(j|i). Not computed: the width sigma_i of each point's Gaussian, one per point; the next slide shows how it is really set - In the figure, x^(i) is the starred point and each other point is drawn larger the more probability it receives. With perplexity 5 the width is 0.30 and all of x^(i)'s probability goes to its own group. Symmetrizing gives one probability per pair, p_ij, adding up to 1: the quantity the map will try to match
- The previous slide gave every point its own width sigma_i and left open how it is chosen. With N points, tuning N widths by hand would be hopeless. t-SNE asks for a single number instead, the perplexity, and sets every width from it - Read perplexity as the number of neighbors a point listens to. Perplexity 2: the starred point's Gaussian is so narrow (0.11) that nearly all its probability goes to one neighbor. Perplexity 5: the width grows to 0.30 and the probability spreads over its own group of eight. Perplexity 12: the width is 1.30 and 16% of the probability now reaches the other two groups - Low perplexity tends to break data into small clumps; high perplexity tends to merge groups. Perplexity must be smaller than the number of points. The precise definition, through the entropy of each point's neighbor probabilities, is in the appendix
- What those widths do to a real t-SNE map: Wattenberg, Viégas and Johnson's two clusters of 50 points, embedded at five perplexities, each run for 5,000 steps - At perplexity 2 each point listens to one or two neighbors and the clusters shatter; at 30 and 50 they come out clean; at 100, with only 100 points, every point would need all others as neighbors, and the clusters mix - 5 to 50 is a starting range, not a guarantee: on some data groups appear at one value and dissolve at the next. Then the honest conclusion is that t-SNE shows no stable grouping, not that one of the pictures is right
- So which features of these maps should you believe? Only what holds across the sweep - Across 5, 30 and 50 the two groups never mix: that is the structure to trust. The small clumps at 5, and the sizes and spacing of the groups, change from panel to panel, so they should not be interpreted. Initialization and the random seed change the map too, and the same rule applies to them
- Back to the mechanics. Step 1 gave the target, P, from the data. Step 2 defines the matching quantity in the map, Q, and it does it with a different kernel (the curve that turns a distance into a similarity; Step 1's was a Gaussian) - Crowding, concretely: on a flat page you can place at most 3 points all the same distance apart (a triangle); in 3-D, 4; in 10-D, 11. High-dimensional data give each point many neighbors that are all about equally near, in many different directions - On a 2-D page there is no room for all of them at that distance, so the map has to compromise. With a Gaussian in the map, a pair that is moderately far apart in the data gets almost no similarity unless it is drawn close, so the optimizer pulls those points closer to each other than they really are. Repeated across the whole data set, every point is pulled toward its moderately distant neighbors, the map shrinks toward its center, and the empty space between clusters disappears: different clusters run into each other and merge - The fix, as a worked example. Suppose a pair of points has a small similarity in the data, 0.011, as moderately distant pairs do. The map has to give that pair the same similarity. With a Gaussian in the map, similarity 0.011 is reached at distance 3. With the t kernel, which falls off much more slowly, the same 0.011 is reached only at distance 9.5. So under the t kernel the pair must be drawn about three times farther apart to get the similarity it should. Moderately distant pairs are therefore placed well away from each other, the map has room, and clusters keep the gaps between them. For close pairs the two kernels almost agree (similarity 0.1: distance 2.2 versus 3), so neighborhoods are barely affected - (This ignores the normalization of Q over all pairs, which shifts the numbers but not the comparison.) - The matrices are real: P from the three groups at perplexity 5, Q from the 1-D map. They look almost identical, and they should: the finished map was fitted to make Q match P, and 98% of Q's mass lies within groups, as in P. The effect of the heavy tail lives in the tiny values between groups, too small to see on this color scale. It matters for the layout, not for how the matrices look. The heavy tail relieves crowding; it does not remove all distortion
- With P from the data and Q from the map, Step 3 moves the map points until the two match, the same way MDS reduces its stress - In MDS you throw the points on the board, measure how far the map distances are from the dissimilarities, move each point a little to lower the error, and repeat. t-SNE does exactly this; only the error is measured differently - The figure is a real run on the toy data. Scattered start, colors mixed: mismatch 1.52. After 25 steps the groups have formed (0.39); after 1,000 they sit well apart (0.11) - The KL divergence is zero when Q equals P and grows as they differ; it returns, with the cross-entropy, when we train neural networks. Gradient descent moves each map point a small step in the direction that lowers the mismatch fastest; it is also how every network is trained - A fixed number of steps does not guarantee the map has settled (see stopping, below) - Why the seed matters: the mismatch has many local minima, so where the points start decides which one the run ends in. Two seeds can give maps with the same groups in different places, or even different groups. Starting from the first two principal components gives every run the same start and keeps the global arrangement closer to the data; that is PCA initialization, and it is the default in recent scikit-learn - When does t-SNE stop? Usually after a fixed number of steps: 1,000 by default in scikit-learn (max_iter). It stops earlier only if the gradient becomes essentially zero (min_grad_norm, 1e-7) or if the mismatch has not improved for 300 steps (n_iter_without_progress). None of these guarantees the map has settled, which is why Assignment 1 checks different budgets and seeds - Standard settings: P exaggerated for the first 100 steps, momentum, 1,000 steps. This run starts from a spread-out layout so the first frame is visible; the usual start is a tighter random layout
- The mismatch is not equally strict about everything. Here is what it punishes and what it lets go, and why that matters for reading any t-SNE map - Both changes start from the finished map. Moving one point from its group to the far end raises the mismatch by 0.38; sliding a whole group of eight against the next group, so that many far pairs become close, raises it by only 0.18 - Each pair counts in proportion to p_ij. Neighbors have large p_ij, so pulling them apart is expensive; pairs from different groups have p_ij near zero, so placing them close or far barely matters. van der Maaten and Hinton state this directly. The terms are coupled through one normalization of Q, so the figure shows the whole mismatch, not single terms - In PCA and MDS the map approximates real distances: twice as far in the map is roughly twice as far in the data. A t-SNE map gives that up. How far apart two groups sit, how large a group looks and how much space surrounds it are side effects of the optimization; to compare distances, sizes or spreads, go back to the response space
- UMAP is the other method you will meet in papers. It rests on the same idea, keep neighbors, built differently - UMAP builds a weighted graph of each point's nearest neighbors and lays it out in 2-D; t-SNE matches neighbor probabilities. The results are close: on this roll 87% of each point's 10 nearest neighbors survive in the t-SNE map and 78% in the UMAP map, and both cut the sheet into pieces, good local fidelity, poor global shape - UMAP is often described as better at global structure. Kobak and Linderman (2021) showed that most of that difference comes from how the two are usually initialized; started the same way, they behave much alike. UMAP is faster on very large data sets
- Like t-SNE, UMAP has settings that change the picture, and the same rule applies - All six maps are UMAP on the same 800-point Swiss roll, seed 0. n_neighbors plays the role of perplexity: with 5 the sheet breaks into many pieces, with 50 longer stretches stay connected - min_dist only sets how close points may sit in the map. At 0.0 they collapse into thin threads that look like sharp clusters; at 0.8 the same neighborhoods spread into bands. The share of each point's 10 nearest neighbors kept stays between 71% and 86% across all six: the look changes far more than the neighborhoods. Tight UMAP clusters in a paper may partly reflect a small min_dist
- We now have five methods. What separates them is which relationship each one keeps, and so what you should check before reading its map - PCA axes are variance directions, not automatically meaningful dimensions. MDS axes have no unique orientation. For t-SNE and UMAP, axis values and island spacing mean nothing in the original units. Checks of a map are useful as long as you know what each one establishes
- In practice, which one should you reach for? - PCA first, always: it is deterministic, applies to new data, and its explained variance tells you how much a 2-D picture can show. If two components already keep most of the variance, a nonlinear map may add little - Published comparisons that favor t-SNE or UMAP usually reflect different default settings. UMAP is faster on very large data and has a transform for new points; t-SNE is older and better studied. Assignment 1 uses t-SNE so you learn one method well; nothing in it would change with UMAP - A map is a hypothesis generator. Picking, among many runs, the one with the cleanest groups is the visual version of p-hacking: report the sweep and check claims in the response space - The appendix shows each method at work in published neuroscience and AI, two examples per method
- Back to our own data, and to the question every map raises: which of its features can you believe? The same network responses can give two very different-looking maps - Nothing about the network changed between the two maps. Both are 2-D summaries of the same responses: PCA keeps the directions of largest variance, t-SNE keeps each point's nearest neighbors. A difference between them is a fact about the methods, not about the network - Why these two checks: nearest neighbors in the full responses use every dimension, not just the two a map keeps, so they cannot be hidden by PCA or created by t-SNE. A classifier tested on images it was not trained on tells you whether the groups can actually be read out, rather than only seen. Assignment 1 runs these checks on the maps; Assignment 2 tests what can be read out and how responses compare with human judgments and neural recordings
- See it happen, on the same objects we looked at with PCA: the Assignment 1 images, all 51 objects with 8 of their 24 views each, 408 images. Each point is AlexNet fc6's response to one image (4,096 numbers), reduced to 50 principal components, which keep 81% of the variance - t-SNE runs live in the browser from these 50 numbers per image; UMAP (with n_neighbors 5, 15 or 50), PCA and MDS are precomputed on the same points. Run t-SNE, change the seed and run it again: the islands land in different places and with different sizes. Move the perplexity and the same data become many small islands or a few large ones - Hover over any image: it appears large in the corner, and lines join it to its 10 nearest neighbors in the original 50-dimensional responses. If the lines reach far across the map, the map has not kept that image's neighborhood - The two meters are the checks from Assignment 1. Neighbor agreement: for each image, the share of its 10 nearest neighbors in 50-D that are still among its 10 nearest in the map (66% for t-SNE at perplexity 30, 62% for UMAP). Object purity: the share of its map neighbors that are other views of the same object (32% and 31%). k is the neighborhood size used by both meters, not a number of components
- The demo showed that the picture changes with the seed and the settings. How do we know which of its features to believe? Wattenberg, Viégas and Johnson answered that with toy data whose true structure is known - Their article, How to Use t-SNE Effectively (Distill), is required reading for Assignment 1, which tests its warnings on object representations
- Four ways a t-SNE map can show you something that is not in the data. Each follows from the mechanics we just built - Sizes: the per-point widths sigma_i expand dense groups and shrink sparse ones, so a tight cluster and a diffuse one look the same size. Measure spread in the response space - Gaps: the mismatch barely penalizes getting far pairs wrong, so the distance between two islands, and even their order, can change between runs. Compare group distances in the response space - Clumps from noise: at low perplexity, structureless data break into clumps. Ask what the map of structureless data would look like - Too few steps: early in a run, pinched and stretched shapes are optimization in progress. A new seed gives a new picture; check that a claim survives re-runs
- A natural defense is repeatability: if every run shows the same neighbors, the map must be right. It need not be - The numbers are a constructed example, not assignment results. Use the same images, neighborhood size and distance in each comparison, and exclude the query image itself. Object purity can be high even when the particular neighboring views change, and none of these checks measures recognition on new images
- If a map cannot settle a question, what can? Each question about a representation has a test in the original response space - A simple readout tests whether task information is easy to access in a representation; a more powerful readout may add computation of its own. Similar accuracy does not mean similar geometry, and similar maps establish neither. Assignment 2 is built around these tests
- Putting the lecture together - The next lecture takes the geometry of population responses (of neurons or network units) as the object of study: how many dimensions a representation really uses, and how the shape of object manifolds relates to what a readout can do. Later, with recurrent networks, the question becomes how a circuit moves through its population states
- For going further, the required reading first, then papers where these methods answer a real question - Chaudhuri, Gerçek, Pandey, Peyrache & Fiete, The intrinsic attractor manifold and population dynamics of a canonical cognitive circuit across waking and sleep, Nature Neuroscience 2019: what evidence connects a population map to a circuit mechanism? - Gardner, Hermansen, Pachitariu, Burak, Baas, Dunn, Moser & Moser, Toroidal topology of population activity in grid cells, Nature 2022 - Nieh, Schottdorf, Freeman et al., Geometry of abstract learned knowledge in the hippocampus, Nature 2021: position and accumulated evidence on one low-dimensional manifold - Engels, Michaud, Liao, Gurnee & Tegmark, Not all language model features are one-dimensionally linear, ICLR 2025 - Chari & Pachter, The specious art of single-cell genomics, PLOS Computational Biology 2023: a critique of reading t-SNE and UMAP maps
- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 5)
These slides were not needed in class; they are here for reference while working on the assignment.
- Examples from the literature, one method per slide, starting with PCA; each caption gives the species, the brain area and the kind of recording - Gámez, Mendoza, Prado, Betancourt and Merchant 2019, PLoS Biology: premotor population activity during rhythmic tapping, on its first principal components; each tempo is a loop, and the loop grows with the interval - Safaie, Chang, Park, Miller, Dyer, Glaser and Gallego 2023, Nature: population activity of different monkeys doing the same reaches, each reduced with PCA; after a rotation (canonical correlation analysis) the dynamics match across animals
- MDS is the workhorse for comparing brain and behavior - Mur, Meys, Bodurka, Goebel, Bandettini and Kriegeskorte 2013, Frontiers in Psychology: the same 92 objects arranged by MDS of human IT response patterns and of similarity judgments; the judgments are more sharply categorical - Hebart, Contier, Teichmann, Rockter, Zheng, Kidder, Corriveau, Vaziri-Pashkam and Baker 2023, eLife (THINGS-data): MDS of MEG patterns at successive times after image onset
- Isomap, despite its fragility, found one of the cleanest results in systems neuroscience - Viejo and Peyrache 2020, Nature Communications: head-direction population activity embedded with Isomap forms a ring; during sleep the activity keeps moving around it - Asumbisa, Peyrache and Trenholm 2022, Nature Communications: the ring survives in blind mice; after olfactory ablation it no longer lines up with the actual heading
- t-SNE made cell types and behavior visible at a glance - Kobak and Berens 2019, Nature Communications, "The art of using t-SNE for single-cell transcriptomics": 23,822 mouse cortical cells placed by their gene expression; the paper also compares PCA and MDS of the same data - Berman, Choi, Bialek and Shaevitz 2014, Journal of the Royal Society Interface: every moment of a freely moving fly, described by its posture dynamics, placed with t-SNE; regions correspond to stereotyped behaviors. Only a low-resolution copy is openly available
- UMAP, from spike shapes to the inside of an image network - Lee, Mendoza, Sathish, Kiani, Kanwisher and Chandrasekaran 2021, eLife (WaveMAP): UMAP of extracellular spike waveforms separates putative cell types - Carter, Armstrong, Schubert, Johnson and Olah 2019, Distill, "Activation Atlas": one InceptionV1 layer laid out with UMAP and shown as feature visualizations; a gridded summary of the layout, not a raw scatter - For each of these maps, the checks from the lecture apply: which features hold across settings and seeds, and which are confirmed in the original data
- The study combines manifold characterization, decoding and comparison across waking and sleep. The ring plot alone is not the evidence for a circuit mechanism - Isomap is the method from earlier today: nearest-neighbor graph, path lengths, then MDS. The paper uses it to visualize and to reduce dimension before its topological analysis, which counts loops (a Betti-1 barcode) rather than reading them off a picture - A circle can also be preserved by a two-dimensional PCA projection; nonlinear methods do not automatically preserve loops - Engels, Michaud, Liao, Gurnee and Tegmark, Not All Language Model Features Are One-Dimensionally Linear: a cyclic variable in a trained network ends up on a circle, as it does in the thalamus - The further scientific question is how recurrent interactions generate and maintain this pattern of population activity
- Gardner, Hermansen, Pachitariu, Burak, Baas, Dunn, Moser and Moser, Toroidal topology of population activity in grid cells, Nature 2022. Hundreds of grid cells recorded at once in one module; the activity is first reduced with PCA to six dimensions and then embedded with UMAP - A grid cell fires on a repeating triangular lattice as the animal runs; moving along either lattice direction eventually brings the cell back to the same state, so position on the lattice is two periodic variables. Two periodic variables make a torus, which continuous-attractor models predicted decades before it was seen - The barcodes are the test the picture needs. Each bar is a feature that persists as the neighborhood radius grows: one long H0 bar means one connected piece, two long H1 bars two independent loops, one long H2 bar one enclosed cavity. That count (1, 2, 1) is a torus's - The head-direction ring on the previous slide is the one-dimensional case of the same idea
Drag to rotate the roll, then drag the unroll slider: every point moves in a straight line from its 3-D position to its place on the flat sheet, so you see which points were neighbors along the sheet all along. The two probes start one turn apart: close in 3-D (the dashed straight segment), far along the sheet (the green path, the shortest route through the 10-nearest-neighbor graph). The three panels on the right embed the same 800 points: PCA keeps as much variance as a flat projection can and so folds the sheet; Isomap approximates along-the-sheet distances and here unrolls it; UMAP favors local neighborhoods, which keeps nearby points together but does not guarantee an unrolled sheet. Isomap is not in Assignment 1.
This is a useful exploratory application across species. The reported Hopkins statistic measures clusterability after UMAP, so it is not an independent validation in the original acoustic space. The choice of acoustic representation and the sampling of vocal elements also affect the interpretation.
- The capital P in Step 3 is different: it is the symmetric matrix of all p_ij, formed after this step by averaging p_(j|i) and p_(i|j). P_i here is one row of the conditional probabilities before symmetrizing - Because every point is tuned to the same perplexity, every point ends up with about the same effective number of neighbors, whatever the density around it. This is also why t-SNE equalizes the apparent size of dense and sparse clusters, one of the pitfalls later in the lecture