- PCA: the number of directions needed for 90% of the variance is the embedding dimension - A manifold is a curved, lower-dimensional shape inside the space of responses, such as a ring or a sheet
- Question 1 uses the eigenvalues of PCA on recordings of thousands of neurons - Question 2 uses the dot product from the MDS lecture: one weight per unit and the flat (linear) boundary between object manifolds of the last lecture - IT: inferotemporal cortex, the late stage of the primate visual pathway, where neurons respond to whole objects
- Question 1 uses the eigenvalues of PCA on recordings of thousands of neurons - Question 2 uses the dot product from the MDS lecture: one weight per unit and the flat (linear) boundary between object manifolds of the last lecture - IT: inferotemporal cortex, the late stage of the primate visual pathway, where neurons respond to whole objects
- Of all the directions available in the neural state space (one axis per neuron), how many do the responses use?
- Recall the Hubel & Wiesel video: a V1 neuron fires for a bar at one orientation in one place. The population tiles orientations and positions - A few hundred orientations and positions could be covered by many more neurons than that, each close to a copy of another. Copies add neurons but not directions - Mouse V1: Stringer et al. (2019) recorded about 10,000 neurons at once; the answer comes in a few slides - GPT-2: Engels et al. (2025), on the next slide
- Ambient dimension: how many units we recorded (768) - Embedding dimension: how many principal components PCA needs for 90% of the variance (2, a plane) - Intrinsic dimension: the number of dimensions needed to locate a point on the ring (1, the angle) - Perich (2026): asking for "the" dimensionality of a population is the wrong question; each of the three measures answers a different one, and all three describe the same responses - Engels et al. (2025): larger models (Mistral 7B, Llama 3 8B) use such circles to answer "Tuesday plus three days"
- Independent columns: no column is a weighted sum of the others. Any two of the three columns determine the third, so the rank is 2 - Read by rows, the six points lie on the plane n3 = n1 + n2, and PCA gives the third principal component zero variance
- Every recorded response carries noise, and noise is not shared exactly between neurons, so real response matrices have full rank - The useful question is how the variance is spread over the directions: the scree plot of the PCA lecture, next - In papers, an x axis labeled "rank" is the position of a component in the list (component number k), not the rank of a matrix
- f_k = λ_k / Σ_j λ_j: the eigenvalue of component k divided by the sum of all eigenvalues - The scree plot and the eigenspectrum are one plot under two names
- Calcium imaging: neurons express a protein that fluoresces when calcium enters the cell, which happens when it fires; a microscope records thousands of neurons at once, one number per neuron per image - Component 1 carries 3.4% of the variance, the first 10 carry 25%, and 90% of the variance needs K = 744 components - A power law λ_k ∝ k^(−α) would shrink this way; α is the decay exponent. The next slide plots every component on log-log axes
- α is minus the slope of a straight line fitted to log f_k against log k over components 11 to 500, the range Stringer et al. used. Per session 0.99 to 1.13, mean 1.06; the paper reports about 1.04 - A slope of −1: component 100 carries a tenth of the variance of component 10, and component 1,000 a tenth of component 100
- Without an elbow, even late directions carry repeatable variance, and the embedding dimension grows with the number of neurons and images recorded - A power law with α = 1 over 1,000 components: the first 10 carry about 39% of the variance, the first 100 about 69%, and 90% needs about 500 - Two extremes: copies of one signal put all the variance in one component (a sharp elbow at 1); independent signals of equal size give every component the same variance (a flat spectrum). V1 lies between
- Minute paper 6 asks about this slide - Check: k equal eigenvalues λ give (kλ)² / (kλ²) = k - The course calls the participation ratio the effective dimensionality: a soft count of the directions that carry variance, related to the embedding dimension, with no threshold - A direction with a small eigenvalue adds a little, not a whole dimension - The code that computes PR from a response matrix is in the appendix
- 5 hidden signals: five signals, independent of one another across images, and every unit is a different weighted mix of the same five plus small noise. The units are therefore strongly correlated: all 512 columns lie in the same 5-dimensional subspace (rank 5, plus noise), so five directions carry the variance. The mixing weights are random, not orthogonal; PCA finds five orthogonal directions inside that subspace - The intuition for 5 → 32 → 401: PR asks how many directions share the variance evenly. Five signals give five directions; 512 equal signals give nearly 512; 512 signals whose sizes shrink by 3% each put most of the variance on the largest few, so PR is small even though no unit copies another - Independent, equal variances: PR falls short of 512 because 1,854 stimuli are a finite sample and chance correlations spread the eigenvalues - Unequal variances: a low PR can come from correlated units or from independent units with very unequal variances, so a low number does not by itself mean redundancy - ResNet-18's last layer (512 units), 1,854 THINGS images, one per concept: PR 99. On other stimuli the value differs: PR describes the responses to one stimulus set - PR does not measure information: a population with a low PR can still let a readout decode a variable
- Computed from the cross-validated spectra of the seven sessions (positive part), 2,800 images: PR 75, 117, 112, 94, 145, 89, 140; mean 110 - Far below the 744 components needed for 90% of the variance, and far below the ~10,000 neurons recorded: most directions carry a little variance each, and PR weights them by how much they carry - With plain PCA (noise included) the PR is higher, 180 to 500: noise spreads variance over many directions
- Arbitrary scales: a voxel's signal depends on its distance from the scanner coil; a neuron's fluorescence on how much indicator it expresses. These differences are about the measurement, not the stimuli - Firing rates in spikes/s, or a layer's activations as the next layer receives them, are sizes that matter - If a conclusion changes between the two choices, the conclusion is about the scaling, not about the population
- Counting dimensions says nothing about what is in them: which information about the images another neuron could use
- Each person's images form a curved sheet, an object manifold; in pixels the two sheets are tangled. In a later stage of the visual pathway they are pulled apart and a flat boundary separates them - Information here means: the responses differ between two groups of images in a way another neuron could use. "Untangled" means a linear readout suffices - The last lecture promised this test
- x₁ and x₂ are the responses of the two units, the two axes of the figure; w₁ and w₂ are their weights - The bias shifts the threshold. The next lecture treats the same weighted sum as a model neuron
- (wᵀx + b)/‖w‖ is the signed distance of a point to the boundary: positive on the "yes" side - With 480 neurons the boundary is a hyperplane we cannot draw; the next slides draw two directions of it - How the weights are fitted (logistic regression) is a Theme 2 topic; here a routine fits them
- 482 units recorded, minus one broken channel and one duplicate unit: 480 neurons, the data of the assignment - Only the first two principal components are used so the whole readout can be drawn. With all 480 neurons a readout does better - The direction that best separates the groups is not a principal direction: most variance need not mean most useful - Kriegeskorte et al. (2008) found the animate/inanimate split in monkey and human IT (MDS lecture)
- The accuracy is the share of views on the correct side of the boundary
- Animate = animals and faces, 37% of the test views, so the majority class is inanimate - The majority-class baseline depends on the split: here 63% of the test views are inanimate - A third comparison: a readout fitted on shuffled labels sees the same responses with the wrong answers; its test accuracy is what "no information" looks like
- Training accuracy 93%, test accuracy 93%, majority-class baseline 63% on the test views - Why hold views out: a readout with enough weights can fit anything it is trained on, including noise. Its accuracy on views it never saw is the fair measure - The test set is never used to fit anything, the z-scoring included, so no information leaks from it - A readout used this way, to test what a representation makes available, is a probe. With several classes, a probe computes one weighted sum per class and answers with the largest (the assignment's version)
- The same IT points as the previous slide (480 neurons, first two principal components). Place the boundary by hand on the training views first, then reveal the test views - The fitted readout: 93% of training views, 93% of test views, against a majority-class baseline of 63% - The demo is on the course site with the other demos
- Animate against inanimate was one two-class task, and a boundary worked. Not every two-class task is like this
- XOR (exclusive or) answers "yes" when exactly one of two inputs is on; the figure colors its complement, and the geometry is the same - 75% was found by trying every boundary
- The third unit responds to the product x₁x₂: positive at the green corners, negative at the blue ones. Its response is not a weighted sum of the two inputs - One task needed an extra dimension. What about all tasks? Next - A hidden layer of a network learns such units, a later lecture
- Gauthaman et al. (2025): when one person's voxels are recombined to match another's, far more dimensions are shared than anatomy alone aligns - Kong et al. (2022): networks trained to resist adversarial attacks (changes to an image we cannot see) have steeper spectra, closer to V1's, and predict V1 responses better - Pavuluri & Kohn (2026): in V2, different textures line up more consistently, so a readout tells apart textures it never saw; standard networks did not reproduce this - Bernardi et al. (2020): read the Introduction and Figures 1–2; skip the Methods - Before posting, skim the existing entries on your paper; your post must add something not already said
- Submitted on Canvas (Minute papers → Minute Paper 6); credit for a thoughtful attempt
- Code and extra examples for the lecture's measures; each is pointed to from a main slide
- Not needed for the assignment. Why the variance falls as a power law, whether the same law holds across species, and what it may have to do with robustness in deep networks
- Monkey V1: Papale et al. (2025), Utah arrays in two macaques, 512 V1 sites per monkey, 100 THINGS photographs each shown 30 times (the dataset of the assignment) - Fitted over components 2 to 30, because 100 images give at most 99 components - The two exponents are not measured under the same conditions, so compare the shapes, not the second decimal. On the same mouse data, 100 random images give α ≈ 0.57 - Two species, two recording methods: the same shape, no elbow
- The pixels: every image as a vector of 4,590 pixel values, ordinary PCA across the 2,800 images, as for the faces of the PCA lecture - Natural images have far more contrast at coarse scales than at fine ones, so their pixel spectrum is also close to a power law with exponent 1
- Whitening divides each spatial frequency by its average amplitude across natural images, so every scale has the same contrast; the images look sharpened and grainy - 7 recording sessions saw natural images and 4 other sessions saw the 2,800 images whitened; in the 4 whitened sessions α ranges 0.96 to 1.20, mean 1.09 - Pixels and neurons are both vectors, and the geometry of the neural responses is a property of the neurons
- The figure: one estimator (cross-validated PCA) on the 100 THINGS test images, monkey V1, V4 and IT (Papale et al. 2025, two monkeys averaged) and human visual cortex (THINGS-fMRI, three people averaged). Fitted over components 2 to 30: 1.35, 1.25, 1.24, 1.49 - What holds across datasets is the shape; the exponent does not. On the same mouse V1 data the exponent moves with the estimator and the number of images (the number of images changes it too) - Gauthaman et al. (2025) title their paper "Universal scale-free representations in human visual cortex"; they also find the power law shared between people once their responses are aligned - Shepard (1987) called his exponential a universal law because the same shape held across stimuli, senses and species
- The dotted line is the 1/k spectrum: component k carries a variance proportional to 1/k, slope −1 on log-log axes - The theory, for stimuli that vary along d dimensions: a smooth code needs α ≥ 1 + 2/d. Natural images have many dimensions, so the bound is close to 1. For gratings that vary only in orientation (d = 1) the bound is 3, and mouse V1's spectrum for 32 gratings is much steeper (about 2.5 in a rough fit), as the theory predicts - The bound concerns the tail of the spectrum with very many neurons and stimuli; a fitted exponent does not by itself prove smoothness. Pospisil & Pillow (2024) argue the V1 spectrum is a broken power law
- How the change was found: 40 times in a row, nudge every pixel by a quarter of a level in the direction that makes the network more confident in "ostrich", never letting a pixel move by more than 2 of 255 levels. That direction is the gradient, which Theme 2 explains - Random noise of the same size leaves the answer unchanged; the change is aimed at the network's weaknesses - Szegedy et al. (2014) found such images; Madry et al. (2018) showed that training on attacked images (adversarial training) makes networks harder to fool, at some cost in accuracy on clean images - Veerabadran et al. (2023): changes this small can bias which of two labels people choose when forced to guess, but people do not see an ostrich
- How Nassar et al. forced the power law: on each batch of training images they computed the covariance of a hidden layer's activations, took its eigenvalues, and added to the loss a penalty for how far those eigenvalues departed from a 1/k spectrum. The network learns its task and a V1-like spectrum at once - Kong et al. (2022) compared about 40 networks with mouse and macaque V1: the adversarially trained ones had steeper spectra and predicted V1 responses better - The network exponents are fitted over components 10 to 200 on 1,854 images, V1's over 11 to 500 on 2,800: compare the slopes loosely - Both results hold for the networks and attacks those papers tested; a steep spectrum by itself does not prove robustness - The six networks differ in how they were trained: from labels (ResNet-18, ResNet-50), from two views of one image (DINO, DINOv2), from image–caption pairs (CLIP). Exponents fitted over components 10 to 200
- Backs the participation-ratio slide; the assignment computes the participation ratio the same way - The fit range (components 11 to 500) avoids the first component and the noisy tail, which bend the line; with fewer stimuli or units, use a shorter range
- The 2-D drawing only shows that directions must overlap (seven directions in two dimensions are 51 degrees apart); the claim is about high dimensions - This packing is called superposition (Elhage et al. 2022); it returns in the lecture on explaining networks