Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 5 · Theme 1: Representational spaces

Geometry of neural representations

Tuesday, September 29 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Where we are

Lecture Takes in Recovers Why it matters
MDS an RDM, e.g. people's similarity ratings a map: the psychological space a mental space measured from judgments alone
PCA the response matrix (images × neurons) the directions of largest variance, the principal components how many directions hold most of the variance
t-SNE, UMAP, Isomap the response matrix a 2-D map that keeps neighbors on a curved shape (a manifold) whether one object's views stay together

Two inputs: an RDM (one number per pair of images) or a response matrix (one per image and neuron).

Today: two questions

  • The data: the response matrix (images × neurons) of a population of neurons
  • Two questions about a recorded population:
    1. How many dimensions does a population of neurons use?
    2. What can another neuron read from the responses?

Claims such as "V1 is high-dimensional" or "IT encodes animacy" rest on these computations.

Today: two questions

  • The data: the response matrix (images × neurons) of a population of neurons
  • Two questions about a recorded population:
    1. How many dimensions does a population of neurons use?
    2. What can another neuron read from the responses?
  • How: (1) count directions with PCA's eigenvalues, summed up by the participation ratio; (2) fit a linear readout, a boundary between groups of images, scored on new images
  • Next lecture: we finish question 2 (shattering, CCGP), then compare two representations

Claims such as "V1 is high-dimensional" or "IT encodes animacy" rest on these computations.

How many dimensions does a population of neurons use?

Why count dimensions?

  • Each image's response is one point in the neural state space: one axis per recorded neuron
  • Hubel & Wiesel (course introduction): V1 neurons prefer different orientations at different positions. There are only so many orientations, so many neurons may repeat one another
  • Mouse V1, ~10,000 neurons: how many independent directions do the responses use?
  • GPT-2, 768 units of layer 7: how many directions do the days of the week use?

With heavy redundancy, the population uses far fewer dimensions than neurons recorded.

Three measures of dimension: GPT-2

Two panels from GPT-2-small, layer 7, projected on principal components 2 and 3, with a wide gap between them. Left, the representations of the seven days of the week form a ring of colored clusters in calendar order, Monday to Sunday. Right, the twelve months form a ring in calendar order, January to December

Reproduced from Engels et al. (2025), Fig. 1 (CC BY 4.0): GPT-2, layer 7

  • Ambient 768 units · embedding 2: 90% of the variance fits in a plane · intrinsic 1
  • t-SNE or UMAP on these responses would draw the ring: they estimate the manifold
  • Each measure answers a different question, so none is the dimensionality (Perich 2026)

Rank: how many directions the points span

Three panels computed from a toy matrix of 6 stimuli by 3 neurons. Left, the matrix with its numbers, in which every value in neuron 3's column is the sum of neurons 1 and 2. Middle, the six stimuli as points in a 3-D space with axes n1, n2 and n3, all lying exactly on a tilted gold plane. Right, bars of the variance fraction of principal components 1, 2 and 3, the third at zero

Toy data, computed

  • Rank = number of directions with nonzero variance = number of principal components (PCA) with nonzero variance. Neuron 3 = neuron 1 + neuron 2, so here rank 2: a plane

Rank: how many directions the points span

The same three panels with one stimulus moved off the plane: its value for neuron 3 no longer equals the sum of neurons 1 and 2, the point sits off the gold plane in the 3-D view, and the third principal component now has a small nonzero bar

Toy data, computed

  • Rank = number of directions with nonzero variance = number of principal components (PCA) with nonzero variance. Neuron 3 = neuron 1 + neuron 2, so here rank 2: a plane
  • One stimulus off the plane: rank 3. Noise does this to every real recording (480 IT neurons: rank 480). So we look at how the variance spreads

From the PCA lecture: the scree plot

The scree plot from the PCA lecture for the 400 Olivetti faces: the variance fraction of the first 30 principal components falls steeply over the first few and then flattens, an elbow

From the PCA lecture: 400 Olivetti faces

  • Variance fraction fk=λk/∑jλjf_k = \lambda_k / \sum_j \lambda_j, largest first. An elbow = a natural place to stop

Mouse V1: variance fraction per component

Two panels for mouse V1 responses to 2,800 natural images, averaged over seven sessions. Left: bars of the variance fraction of the first 30 principal components, falling slowly from about 0.035 to about 0.006 with no elbow. Right: cumulative variance fraction against components retained, on a log axis out to 2,366; a dashed line at 0.9 is crossed at K = 744

Computed from Stringer et al.'s data (7 sessions)

  • Two-photon imaging, ~10,000 neurons, each image shown twice. Cross-validated PCA: principal components from the first showing; each one's variance = how much the two showings co-vary along it, so trial noise cancels
  • No elbow: 90% of the variance needs 744 components

On log-log axes, a straight line

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1, seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; the legend gives alpha = 1.06

Computed from Stringer et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)

No elbow: what it means

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

Computed from Stringer et al.'s data

  • With an elbow, a few directions carry most of the variance: the embedding dimension is clear
  • Without one, the embedding dimension depends on the threshold (90%? 95%?)
  • V1 uses many dimensions, but unequally: each further component carries a little less. We want one number with no threshold
  • Why V1 has no elbow, and why it may matter for deep networks: see the appendix

One number: the participation ratio

  • The participation ratio of the eigenvalues λk\lambda_k (variances along the principal components):

PR=(∑jλj)2∑jλj2,1≤PR≤D\mathrm{PR} = \frac{\left(\sum_j \lambda_j\right)^2}{\sum_j \lambda_j^2}, \qquad 1 \le \mathrm{PR} \le D

  • With the variance fractions fjf_j: PR=1/∑jfj2\mathrm{PR} = 1/\sum_j f_j^2. Not to be confused with fkf_k: one number per component; PR is one number for the whole spectrum
  • PR = 1 when one direction carries all the variance; PR = DD when all DD share it equally. No threshold; report PR/D\mathrm{PR}/D

The participation ratio counts how many directions the responses use: if k directions share the variance equally, PR = k. (minute paper)

Participation ratio, four populations

Log-log plot of variance fraction against principal component k for four populations of 512 units on 1,854 stimuli. 512 units that each mix the same five hidden signals: the first five components hold almost all the variance, PR about 5. Independent signals of equal variance: a nearly flat line, PR about 401. Independent signals of unequal variances: a line that falls steeply, PR about 32, although no unit is redundant. ResNet-18's last layer on 1,854 THINGS images: a line that falls steadily, PR 99

Simulated: 512 units, 1,854 stimuli, Gaussian signals + noise

  • 5 hidden signals, each unit a different mix: units correlated, rank ≈ 5, PR ≈ 5
  • 512 independent signals, one per unit, equal size: variance everywhere, PR ≈ 401
  • Same, but unequal sizes: the strongest few dominate, PR ≈ 32. ResNet-18: 99
  • PR counts how evenly the variance is spread, not how much information there is

Back to mouse V1

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

Computed from Stringer et al.'s data

  • ~10,000 neurons recorded; 90% of the variance needs 744 components
  • Participation ratio: about 110 (75 to 145 per session)
  • One number for a spectrum with no elbow: many dimensions, used unequally

Scale matters: z-score or not

  • The eigenvalues are variances, so a unit measured on a large scale dominates them, and the participation ratio with them
  • z-score each unit (subtract its mean, divide by its standard deviation) when scales are arbitrary: fMRI voxels, calcium fluorescence. Keep the raw scale when sizes matter: spikes/s
  • z-scoring gives every unit variance 1: PR then counts patterns of co-activation, not loud units. Independent units of unequal size: PR 32 raw, 402 z-scored
  • ResNet-18 on THINGS: PR 99 raw, 107 z-scored. Report which you used

What can another neuron read from the responses?

What could another neuron use?

In pixel space the sheets of face images of two different people are interleaved, and a flat separating plane drawn through them fails to split them

Redrawn in the last lecture from DiCarlo & Cox (2007)

  • In pixels, two people's face images interleave: no flat (linear) plane separates them
  • "IT encodes animacy" is a claim about information another neuron could use
  • The simplest such neuron: a weighted sum and a threshold, a linear readout

A linear readout draws a hyperplane

A 2-D plane with axes x1 and x2 holding two groups of dots, green "yes" points upper right and blue "no" points lower left. A fitted straight line runs between them and the two sides are lightly shaded. A magenta arrow labelled w starts on the line at a right angle and points into the green side. A dotted segment from a far point to the line is labelled distance

Toy data, computed

  • Response vector x\mathbf{x}, weights w\mathbf{w}, bias bb: the readout answers "yes" when w⊤x+b>0\mathbf{w}^{\top}\mathbf{x} + b > 0
  • w⊤x=w1x1+w2x2\mathbf{w}^{\top}\mathbf{x} = w_1 x_1 + w_2 x_2, the dot product from the MDS lecture: one weight per unit

x=(3,1)\mathbf{x} = (3, 1), w=(2,−1)\mathbf{w} = (2, -1), b=−4b = -4
2⋅3−1⋅1−4=1>02 \cdot 3 - 1 \cdot 1 - 4 = 1 > 0, so the answer is "yes"

A linear readout draws a hyperplane

A 2-D plane with axes x1 and x2 holding two groups of dots, green "yes" points upper right and blue "no" points lower left. A fitted straight line runs between them and the two sides are lightly shaded. A magenta arrow labelled w starts on the line at a right angle and points into the green side. A dotted segment from a far point to the line is labelled distance

Toy data, computed

  • Response vector x\mathbf{x}, weights w\mathbf{w}, bias bb: the readout answers "yes" when w⊤x+b>0\mathbf{w}^{\top}\mathbf{x} + b > 0
  • w⊤x=w1x1+w2x2\mathbf{w}^{\top}\mathbf{x} = w_1 x_1 + w_2 x_2, the dot product from the MDS lecture: one weight per unit
  • The boundary w⊤x+b=0\mathbf{w}^{\top}\mathbf{x} + b = 0: a line in 2-D, a plane in 3-D, a hyperplane in DD dimensions
  • w\mathbf{w} is perpendicular to it. Each point has a distance to the boundary

Monkey IT: reading animacy

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Green dots are animate images (animals and faces), blue dots inanimate ones. A straight tilted line fitted to the points separates most animate images from inanimate ones

Computed: monkey IT, 480 neurons, 2 principal components

  • Bao objects: 51 objects × 24 views, recorded in monkey IT, 480 neurons
  • Axes: IT's first two principal components. Green: animals and faces; blue: inanimate
  • Neither component alone separates them: the best boundary is tilted, a weighted sum of both
  • IT separates animate from inanimate objects (Kriegeskorte et al., 2008)

Measuring with a readout

Monkey IT responses to the 1,224 Bao object images on their first two principal components, animate views in green and inanimate in blue, with the fitted line and the share of views it classifies correctly

Computed: monkey IT, 480 neurons, 2 principal components

  • Result: the readout reads animacy correctly for 93% of the views. Is that a lot?

Measuring with a readout

Monkey IT responses to the 1,224 Bao object images on their first two principal components, animate views in green and inanimate in blue, with the fitted line and the share of views it classifies correctly

Computed: monkey IT, 480 neurons, 2 principal components

  • Result: the readout reads animacy correctly for 93% of the views. Is that a lot?
  • Baseline: guessing, 50%; always "inanimate", 63%: the majority-class baseline

Measuring with a readout

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Faint dots are the 18 training views per object used to fit the line; solid dots are the 6 test views per object, held out, on which the line is scored

Computed: monkey IT, 480 neurons, 2 principal components

  • Result: the readout reads animacy correctly for 93% of the views. Is that a lot?
  • Baseline: guessing, 50%; always "inanimate", 63%: the majority-class baseline
  • Split: fit on 18 of each object's 24 views (training set); score on the other 6 (test set). Both 93%: the readout did not overfit

Draw the boundary yourself

  • Each point has a distance to the boundary; points near it flip when the boundary moves

Linearly separable, or not

Toy data, two panels side by side at the same size. Two groups of points in a plane, green and blue, that one straight boundary separates, beside the XOR layout of four clusters at the corners of a square

Toy data, computed: inputs coded −1 and +1

  • Two classes are linearly separable when one hyperplane puts them on opposite sides

Linearly separable, or not

Toy data, two panels side by side at the same size. The XOR layout: four clusters at the corners of a square, green at the upper right and lower left, blue at the upper left and lower right; the best straight boundary classifies 75 percent

Toy data, computed: inputs coded −1 and +1

  • Two classes are linearly separable when one hyperplane puts them on opposite sides
  • XOR: green when both inputs have the same sign. No boundary works: 75% at best

Linearly separable, or not

Toy data, two panels side by side at the same size. Left, the XOR clusters in the plane of x1 and x2. Right, the same clusters lifted by a third unit x1 times x2: the green clusters rise above a translucent plane and the blue ones drop below it, so one plane separates them

Toy data, computed: inputs coded −1 and +1

  • Two classes are linearly separable when one hyperplane puts them on opposite sides
  • XOR: green when both inputs have the same sign. No boundary works: 75% at best
  • A third unit x1x2x_1 x_2 lifts the points into 3-D: one plane separates them

Further reading

Optional: a few research-trail entries over the semester. Any paper mentioned in any lecture qualifies, but papers from this block are preferred.

Minute paper

Submit on Canvas → Minute papers → Minute Paper 6 (access code read out in class).

Write three brief points in your own words:

  1. What does the participation ratio measure?
    If k directions share the variance equally, what is its value?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Appendix: derivations and extra examples

Material not covered in class, for reference

Going further: what does the power law mean?

Monkey V1 falls the same way

The same log-log plot with monkey V1 added: two green lines, one per monkey, for principal components 1 to 60, starting higher and falling a little more steeply than the mouse line, with alpha 1.32 and 1.39

Computed from Stringer et al.'s and Papale et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)
  • Monkey V1: electrode arrays; 100 THINGS images (photos of everyday objects). A little steeper, α=1.32\alpha = 1.32 and 1.391.39

Pixels fall the same way

The same log-log plot with a red line added for the pixels of the 2,800 images the mice saw: it falls almost exactly on top of the mouse V1 line, alpha 1.09

Computed from Stringer et al.'s and Papale et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)
  • Monkey V1, electrode arrays, 100 THINGS images: a little steeper, α=1.32\alpha = 1.32 and 1.391.39
  • Pixels of the 2,800 images: α=1.09\alpha = 1.09, the same slope. Is V1 just copying the images?

Neurons are not pixels

Log-log plot of variance fraction against principal component k for natural and whitened images. For natural images the pixel spectrum, alpha 1.09, and the mouse V1 spectrum, alpha 1.06, lie almost on top of each other. For whitened images the pixel spectrum is much flatter, alpha 0.41, while the mouse V1 spectrum still follows the dotted line of slope minus 1, alpha 1.09

Stringer et al.'s data: 7 natural, 4 whitened sessions

  • Whitened images: all scales, equal contrast
  • Their pixels flatten: α\alpha from 1.09 to 0.41. Mouse V1 does not: 1.06 to 1.09
  • V1's power law is not inherited from the images
  • It is a property of how V1 encodes images (why a slope near 1 matters: two slides on)

A power law across species and methods

Log-log plot of variance fraction against principal component k from 1 to 60 for the same 100 THINGS images in monkey V1, V4 and IT (electrodes) and human visual cortex (fMRI), all with the same cross-validated estimator. All four fall along nearly parallel straight lines, with alpha 1.35, 1.25, 1.24 and 1.49

Computed here: one estimator, 100 THINGS images

  • Variance falls as a power law: mouse V1; monkey V1, V4, IT; human cortex
  • Calcium imaging, electrodes, fMRI: same shape; the exponent depends on the estimator
  • Shepard's law of generalization (MDS lecture) was such a law for behavior
  • A candidate law of neural coding

Why a slope near 1? A smooth code

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1 responses to 2,800 natural images: seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; alpha 1.06

Computed from Stringer et al.'s data

  • Smooth code: similar images give similar responses, as small changes should
  • Slower than 1/k1/k: fine directions dominate; near-identical images land far apart
  • Stringer et al.'s theory: a slope near −1-1 is about the flattest a smooth code allows. V1 uses as many dimensions as smoothness permits
  • Why that matters: networks are not smooth

Deep networks are fooled by invisible changes

Three images in a row. Left, a public-domain photograph of an alligator lying on grass, which ImageNet-trained ResNet-18 labels alligator with 99 percent probability. Middle, the change added to it, magnified 50 times: colored noise; no pixel moves by more than 2 of 255 intensity levels. Right, the photograph plus the change, which looks identical to the original, labeled ostrich with 100 percent probability

Computed here: ResNet-18 trained on ImageNet, a public-domain THINGS photograph

  • Adversarial attack: a change of ≤ 2 of 255 levels per pixel flips the network's answer; we still see the same alligator in both photographs
  • It matters for cars, medical imaging and face recognition, and no defense removes the problem

Could a V1-like spectrum make networks robust?

Log-log plot of the variance fraction of each principal component against its number k, for six networks' last layers on 1,854 THINGS images, with a dotted line of slope minus 1. Five trained networks fall more gently than the dotted line, with exponents between 0.67 and 0.96; the untrained ResNet-18, dashed, falls more steeply, exponent 1.43

  • Six networks, computed here: trained ones are flatter than V1 (0.67–0.96 vs 1.06)
  • Kong et al. (2022): networks trained to resist attacks have steeper, more V1-like spectra
  • Nassar et al. (2020): a training penalty pulls a layer's spectrum toward 1/k1/k: harder to fool
  • Open: is robustness why V1 has a power law?

Effective dimensionality from an eigenspectrum

import numpy as np
# X: (N stimuli, D units) response matrix -- neurons, voxels or model units
Xc = X - X.mean(axis=0, keepdims=True)          # center: PCA needs it (the PCA lecture)

C = Xc.T @ Xc / (Xc.shape[0] - 1)               # the D x D covariance matrix
lam = np.sort(np.linalg.eigvalsh(C))[::-1]      # its eigenvalues, largest first
lam = lam[lam > 0]

pr = lam.sum()**2 / (lam**2).sum()              # participation ratio, 1 <= PR <= D
print(f"ambient D = {X.shape[1]},  effective dimensionality = {pr:.1f}")

f = lam / lam.sum()                             # variance fraction f_k of component k
k = np.arange(1, len(f) + 1)
alpha = -np.polyfit(np.log(k[10:500]), np.log(f[10:500]), 1)[0]
print(f"variance fraction decays as k^-{alpha:.2f}")

More features than units

Two circles of directions: two perpendicular feature directions where each unit means one thing, versus seven directions packed into the same plane, forced to overlap

  • DD perpendicular directions hold at most DD features. Nearly perpendicular directions, plentiful in high dimensions, hold far more, and each unit sits on several
  • A population can code more features than it has units: no unit stands for one thing

- PCA: the number of directions needed for 90% of the variance is the embedding dimension - A manifold is a curved, lower-dimensional shape inside the space of responses, such as a ring or a sheet

- Question 1 uses the eigenvalues of PCA on recordings of thousands of neurons - Question 2 uses the dot product from the MDS lecture: one weight per unit and the flat (linear) boundary between object manifolds of the last lecture - IT: inferotemporal cortex, the late stage of the primate visual pathway, where neurons respond to whole objects

- Question 1 uses the eigenvalues of PCA on recordings of thousands of neurons - Question 2 uses the dot product from the MDS lecture: one weight per unit and the flat (linear) boundary between object manifolds of the last lecture - IT: inferotemporal cortex, the late stage of the primate visual pathway, where neurons respond to whole objects

- Of all the directions available in the neural state space (one axis per neuron), how many do the responses use?

- Recall the Hubel & Wiesel video: a V1 neuron fires for a bar at one orientation in one place. The population tiles orientations and positions - A few hundred orientations and positions could be covered by many more neurons than that, each close to a copy of another. Copies add neurons but not directions - Mouse V1: Stringer et al. (2019) recorded about 10,000 neurons at once; the answer comes in a few slides - GPT-2: Engels et al. (2025), on the next slide

- Ambient dimension: how many units we recorded (768) - Embedding dimension: how many principal components PCA needs for 90% of the variance (2, a plane) - Intrinsic dimension: the number of dimensions needed to locate a point on the ring (1, the angle) - Perich (2026): asking for "the" dimensionality of a population is the wrong question; each of the three measures answers a different one, and all three describe the same responses - Engels et al. (2025): larger models (Mistral 7B, Llama 3 8B) use such circles to answer "Tuesday plus three days"

- Independent columns: no column is a weighted sum of the others. Any two of the three columns determine the third, so the rank is 2 - Read by rows, the six points lie on the plane n3 = n1 + n2, and PCA gives the third principal component zero variance

- Every recorded response carries noise, and noise is not shared exactly between neurons, so real response matrices have full rank - The useful question is how the variance is spread over the directions: the scree plot of the PCA lecture, next - In papers, an x axis labeled "rank" is the position of a component in the list (component number k), not the rank of a matrix

- f_k = λ_k / Σ_j λ_j: the eigenvalue of component k divided by the sum of all eigenvalues - The scree plot and the eigenspectrum are one plot under two names

- Calcium imaging: neurons express a protein that fluoresces when calcium enters the cell, which happens when it fires; a microscope records thousands of neurons at once, one number per neuron per image - Component 1 carries 3.4% of the variance, the first 10 carry 25%, and 90% of the variance needs K = 744 components - A power law λ_k ∝ k^(−α) would shrink this way; α is the decay exponent. The next slide plots every component on log-log axes

- α is minus the slope of a straight line fitted to log f_k against log k over components 11 to 500, the range Stringer et al. used. Per session 0.99 to 1.13, mean 1.06; the paper reports about 1.04 - A slope of −1: component 100 carries a tenth of the variance of component 10, and component 1,000 a tenth of component 100

- Without an elbow, even late directions carry repeatable variance, and the embedding dimension grows with the number of neurons and images recorded - A power law with α = 1 over 1,000 components: the first 10 carry about 39% of the variance, the first 100 about 69%, and 90% needs about 500 - Two extremes: copies of one signal put all the variance in one component (a sharp elbow at 1); independent signals of equal size give every component the same variance (a flat spectrum). V1 lies between

- Minute paper 6 asks about this slide - Check: k equal eigenvalues λ give (kλ)² / (kλ²) = k - The course calls the participation ratio the effective dimensionality: a soft count of the directions that carry variance, related to the embedding dimension, with no threshold - A direction with a small eigenvalue adds a little, not a whole dimension - The code that computes PR from a response matrix is in the appendix

- 5 hidden signals: five signals, independent of one another across images, and every unit is a different weighted mix of the same five plus small noise. The units are therefore strongly correlated: all 512 columns lie in the same 5-dimensional subspace (rank 5, plus noise), so five directions carry the variance. The mixing weights are random, not orthogonal; PCA finds five orthogonal directions inside that subspace - The intuition for 5 → 32 → 401: PR asks how many directions share the variance evenly. Five signals give five directions; 512 equal signals give nearly 512; 512 signals whose sizes shrink by 3% each put most of the variance on the largest few, so PR is small even though no unit copies another - Independent, equal variances: PR falls short of 512 because 1,854 stimuli are a finite sample and chance correlations spread the eigenvalues - Unequal variances: a low PR can come from correlated units or from independent units with very unequal variances, so a low number does not by itself mean redundancy - ResNet-18's last layer (512 units), 1,854 THINGS images, one per concept: PR 99. On other stimuli the value differs: PR describes the responses to one stimulus set - PR does not measure information: a population with a low PR can still let a readout decode a variable

- Computed from the cross-validated spectra of the seven sessions (positive part), 2,800 images: PR 75, 117, 112, 94, 145, 89, 140; mean 110 - Far below the 744 components needed for 90% of the variance, and far below the ~10,000 neurons recorded: most directions carry a little variance each, and PR weights them by how much they carry - With plain PCA (noise included) the PR is higher, 180 to 500: noise spreads variance over many directions

- Arbitrary scales: a voxel's signal depends on its distance from the scanner coil; a neuron's fluorescence on how much indicator it expresses. These differences are about the measurement, not the stimuli - Firing rates in spikes/s, or a layer's activations as the next layer receives them, are sizes that matter - If a conclusion changes between the two choices, the conclusion is about the scaling, not about the population

- Counting dimensions says nothing about what is in them: which information about the images another neuron could use

- Each person's images form a curved sheet, an object manifold; in pixels the two sheets are tangled. In a later stage of the visual pathway they are pulled apart and a flat boundary separates them - Information here means: the responses differ between two groups of images in a way another neuron could use. "Untangled" means a linear readout suffices - The last lecture promised this test

- x₁ and x₂ are the responses of the two units, the two axes of the figure; w₁ and w₂ are their weights - The bias shifts the threshold. The next lecture treats the same weighted sum as a model neuron

- (wᵀx + b)/‖w‖ is the signed distance of a point to the boundary: positive on the "yes" side - With 480 neurons the boundary is a hyperplane we cannot draw; the next slides draw two directions of it - How the weights are fitted (logistic regression) is a Theme 2 topic; here a routine fits them

- 482 units recorded, minus one broken channel and one duplicate unit: 480 neurons, the data of the assignment - Only the first two principal components are used so the whole readout can be drawn. With all 480 neurons a readout does better - The direction that best separates the groups is not a principal direction: most variance need not mean most useful - Kriegeskorte et al. (2008) found the animate/inanimate split in monkey and human IT (MDS lecture)

- The accuracy is the share of views on the correct side of the boundary

- Animate = animals and faces, 37% of the test views, so the majority class is inanimate - The majority-class baseline depends on the split: here 63% of the test views are inanimate - A third comparison: a readout fitted on shuffled labels sees the same responses with the wrong answers; its test accuracy is what "no information" looks like

- Training accuracy 93%, test accuracy 93%, majority-class baseline 63% on the test views - Why hold views out: a readout with enough weights can fit anything it is trained on, including noise. Its accuracy on views it never saw is the fair measure - The test set is never used to fit anything, the z-scoring included, so no information leaks from it - A readout used this way, to test what a representation makes available, is a probe. With several classes, a probe computes one weighted sum per class and answers with the largest (the assignment's version)

- The same IT points as the previous slide (480 neurons, first two principal components). Place the boundary by hand on the training views first, then reveal the test views - The fitted readout: 93% of training views, 93% of test views, against a majority-class baseline of 63% - The demo is on the course site with the other demos

- Animate against inanimate was one two-class task, and a boundary worked. Not every two-class task is like this

- XOR (exclusive or) answers "yes" when exactly one of two inputs is on; the figure colors its complement, and the geometry is the same - 75% was found by trying every boundary

- The third unit responds to the product x₁x₂: positive at the green corners, negative at the blue ones. Its response is not a weighted sum of the two inputs - One task needed an extra dimension. What about all tasks? Next - A hidden layer of a network learns such units, a later lecture

- Gauthaman et al. (2025): when one person's voxels are recombined to match another's, far more dimensions are shared than anatomy alone aligns - Kong et al. (2022): networks trained to resist adversarial attacks (changes to an image we cannot see) have steeper spectra, closer to V1's, and predict V1 responses better - Pavuluri & Kohn (2026): in V2, different textures line up more consistently, so a readout tells apart textures it never saw; standard networks did not reproduce this - Bernardi et al. (2020): read the Introduction and Figures 1–2; skip the Methods - Before posting, skim the existing entries on your paper; your post must add something not already said

- Submitted on Canvas (Minute papers → Minute Paper 6); credit for a thoughtful attempt

- Code and extra examples for the lecture's measures; each is pointed to from a main slide

- Not needed for the assignment. Why the variance falls as a power law, whether the same law holds across species, and what it may have to do with robustness in deep networks

- Monkey V1: Papale et al. (2025), Utah arrays in two macaques, 512 V1 sites per monkey, 100 THINGS photographs each shown 30 times (the dataset of the assignment) - Fitted over components 2 to 30, because 100 images give at most 99 components - The two exponents are not measured under the same conditions, so compare the shapes, not the second decimal. On the same mouse data, 100 random images give α ≈ 0.57 - Two species, two recording methods: the same shape, no elbow

- The pixels: every image as a vector of 4,590 pixel values, ordinary PCA across the 2,800 images, as for the faces of the PCA lecture - Natural images have far more contrast at coarse scales than at fine ones, so their pixel spectrum is also close to a power law with exponent 1

- Whitening divides each spatial frequency by its average amplitude across natural images, so every scale has the same contrast; the images look sharpened and grainy - 7 recording sessions saw natural images and 4 other sessions saw the 2,800 images whitened; in the 4 whitened sessions α ranges 0.96 to 1.20, mean 1.09 - Pixels and neurons are both vectors, and the geometry of the neural responses is a property of the neurons

- The figure: one estimator (cross-validated PCA) on the 100 THINGS test images, monkey V1, V4 and IT (Papale et al. 2025, two monkeys averaged) and human visual cortex (THINGS-fMRI, three people averaged). Fitted over components 2 to 30: 1.35, 1.25, 1.24, 1.49 - What holds across datasets is the shape; the exponent does not. On the same mouse V1 data the exponent moves with the estimator and the number of images (the number of images changes it too) - Gauthaman et al. (2025) title their paper "Universal scale-free representations in human visual cortex"; they also find the power law shared between people once their responses are aligned - Shepard (1987) called his exponential a universal law because the same shape held across stimuli, senses and species

- The dotted line is the 1/k spectrum: component k carries a variance proportional to 1/k, slope −1 on log-log axes - The theory, for stimuli that vary along d dimensions: a smooth code needs α ≥ 1 + 2/d. Natural images have many dimensions, so the bound is close to 1. For gratings that vary only in orientation (d = 1) the bound is 3, and mouse V1's spectrum for 32 gratings is much steeper (about 2.5 in a rough fit), as the theory predicts - The bound concerns the tail of the spectrum with very many neurons and stimuli; a fitted exponent does not by itself prove smoothness. Pospisil & Pillow (2024) argue the V1 spectrum is a broken power law

- How the change was found: 40 times in a row, nudge every pixel by a quarter of a level in the direction that makes the network more confident in "ostrich", never letting a pixel move by more than 2 of 255 levels. That direction is the gradient, which Theme 2 explains - Random noise of the same size leaves the answer unchanged; the change is aimed at the network's weaknesses - Szegedy et al. (2014) found such images; Madry et al. (2018) showed that training on attacked images (adversarial training) makes networks harder to fool, at some cost in accuracy on clean images - Veerabadran et al. (2023): changes this small can bias which of two labels people choose when forced to guess, but people do not see an ostrich

- How Nassar et al. forced the power law: on each batch of training images they computed the covariance of a hidden layer's activations, took its eigenvalues, and added to the loss a penalty for how far those eigenvalues departed from a 1/k spectrum. The network learns its task and a V1-like spectrum at once - Kong et al. (2022) compared about 40 networks with mouse and macaque V1: the adversarially trained ones had steeper spectra and predicted V1 responses better - The network exponents are fitted over components 10 to 200 on 1,854 images, V1's over 11 to 500 on 2,800: compare the slopes loosely - Both results hold for the networks and attacks those papers tested; a steep spectrum by itself does not prove robustness - The six networks differ in how they were trained: from labels (ResNet-18, ResNet-50), from two views of one image (DINO, DINOv2), from image–caption pairs (CLIP). Exponents fitted over components 10 to 200

- Backs the participation-ratio slide; the assignment computes the participation ratio the same way - The fit range (components 11 to 500) avoids the first component and the noisy tail, which bend the line; with fewer stimuli or units, use a shorter range

- The 2-D drawing only shows that directions must overlap (seven directions in two dimensions are 51 degrees apart); the claim is about high dimensions - This packing is called superposition (Elhage et al. 2022); it returns in the lecture on explaining networks