Geometry of neural representations

Reader-friendly version of the 42 lecture slides: the lecture content in normal flow, with figure descriptions and data tables where the figure carries data. Open the presentation.

Slide 1

Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 5 · Theme 1: Representational spaces

Geometry of neural representations

Tuesday, September 29 · Fall 2026

Return to contents

Slide 2

Where we are

Where we are — table
Lecture Takes in Recovers Why it matters
MDS an RDM, e.g. people's similarity ratings a map: the psychological space a mental space measured from judgments alone
PCA the response matrix (images × neurons) the directions of largest variance, the principal components how many directions hold most of the variance
t-SNE, UMAP, Isomap the response matrix a 2-D map that keeps neighbors on a curved shape (a manifold) whether one object's views stay together

Two inputs: an RDM (one number per pair of images) or a response matrix (one per image and neuron).

Return to contents

Slide 4

Today: two questions

Claims such as "V1 is high-dimensional" or "IT encodes animacy" rest on these computations.

Return to contents

Slide 5

How many dimensions does a population of neurons use?

Return to contents

Slide 6

Why count dimensions?

With heavy redundancy, the population uses far fewer dimensions than neurons recorded.

Return to contents

Slide 7

Three measures of dimension: GPT-2

Two panels from GPT-2-small, layer 7, projected on principal components 2 and 3, with a wide gap between them. Left, the representations of the seven days of the week form a ring of colored clusters in calendar order, Monday to Sunday. Right, the twelve months form a ring in calendar order, January to December

Two panels from GPT-2-small, layer 7, projected on principal components 2 and 3, with a wide gap between them. Left, the representations of the seven days of the week form a ring of colored clusters in calendar order, Monday to Sunday. Right, the twelve months form a ring in calendar order, January to December

Open full-size figure

Reproduced from Engels et al. (2025), Fig. 1 (CC BY 4.0): GPT-2, layer 7

Return to contents

Slide 9

Rank: how many directions the points span

The same three panels with one stimulus moved off the plane: its value for neuron 3 no longer equals the sum of neurons 1 and 2, the point sits off the gold plane in the 3-D view, and the third principal component now has a small nonzero bar

The same three panels with one stimulus moved off the plane: its value for neuron 3 no longer equals the sum of neurons 1 and 2, the point sits off the gold plane in the 3-D view, and the third principal component now has a small nonzero bar

Open full-size figure

Toy data, computed

Return to contents

Slide 10

From the PCA lecture: the scree plot

The scree plot from the PCA lecture for the 400 Olivetti faces: the variance fraction of the first 30 principal components falls steeply over the first few and then flattens, an elbow

The scree plot from the PCA lecture for the 400 Olivetti faces: the variance fraction of the first 30 principal components falls steeply over the first few and then flattens, an elbow

Open full-size figure

From the PCA lecture: 400 Olivetti faces

Return to contents

Slide 11

Mouse V1: variance fraction per component

Two panels for mouse V1 responses to 2,800 natural images, averaged over seven sessions. Left: bars of the variance fraction of the first 30 principal components, falling slowly from about 0.035 to about 0.006 with no elbow. Right: cumulative variance fraction against components retained, on a log axis out to 2,366; a dashed line at 0.9 is crossed at K = 744

Two panels for mouse V1 responses to 2,800 natural images, averaged over seven sessions. Left: bars of the variance fraction of the first 30 principal components, falling slowly from about 0.035 to about 0.006 with no elbow. Right: cumulative variance fraction against components retained, on a log axis out to 2,366; a dashed line at 0.9 is crossed at K = 744

Open full-size figure

Computed from Stringer et al.'s data (7 sessions)

Return to contents

Slide 12

On log-log axes, a straight line

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1, seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; the legend gives alpha = 1.06

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1, seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; the legend gives alpha = 1.06

Open full-size figure

Computed from Stringer et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)

Return to contents

Slide 13

No elbow: what it means

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

Open full-size figure

Computed from Stringer et al.'s data

  • With an elbow, a few directions carry most of the variance: the embedding dimension is clear
  • Without one, the embedding dimension depends on the threshold (90%? 95%?)
  • V1 uses many dimensions, but unequally: each further component carries a little less. We want one number with no threshold
  • Why V1 has no elbow, and why it may matter for deep networks: see the appendix

Return to contents

Slide 14

One number: the participation ratio

PR=(∑jλj)2∑jλj2,1≤PR≤D\mathrm{PR} = \frac{\left(\sum_j \lambda_j\right)^2}{\sum_j \lambda_j^2}, \qquad 1 \le \mathrm{PR} \le D

The participation ratio counts how many directions the responses use: if k directions share the variance equally, PR = k. (minute paper)

Return to contents

Slide 15

Participation ratio, four populations

Log-log plot of variance fraction against principal component k for four populations of 512 units on 1,854 stimuli. 512 units that each mix the same five hidden signals: the first five components hold almost all the variance, PR about 5. Independent signals of equal variance: a nearly flat line, PR about 401. Independent signals of unequal variances: a line that falls steeply, PR about 32, although no unit is redundant. ResNet-18's last layer on 1,854 THINGS images: a line that falls steadily, PR 99

Log-log plot of variance fraction against principal component k for four populations of 512 units on 1,854 stimuli. 512 units that each mix the same five hidden signals: the first five components hold almost all the variance, PR about 5. Independent signals of equal variance: a nearly flat line, PR about 401. Independent signals of unequal variances: a line that falls steeply, PR about 32, although no unit is redundant. ResNet-18's last layer on 1,854 THINGS images: a line that falls steadily, PR 99

Open full-size figure

Simulated: 512 units, 1,854 stimuli, Gaussian signals + noise

  • 5 hidden signals, each unit a different mix: units correlated, rank ≈ 5, PR ≈ 5
  • 512 independent signals, one per unit, equal size: variance everywhere, PR ≈ 401
  • Same, but unequal sizes: the strongest few dominate, PR ≈ 32. ResNet-18: 99
  • PR counts how evenly the variance is spread, not how much information there is

Return to contents

Slide 16

Back to mouse V1

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

The log-log plot of variance fraction against principal component k for mouse V1 (2,800 images, alpha 1.06), a nearly straight line with no bend, beside a dotted guide of slope minus 1

Open full-size figure

Computed from Stringer et al.'s data

  • ~10,000 neurons recorded; 90% of the variance needs 744 components
  • Participation ratio: about 110 (75 to 145 per session)
  • One number for a spectrum with no elbow: many dimensions, used unequally

Return to contents

Slide 17

Scale matters: z-score or not

Return to contents

Slide 18

What can another neuron read from the responses?

Return to contents

Slide 19

What could another neuron use?

In pixel space the sheets of face images of two different people are interleaved, and a flat separating plane drawn through them fails to split them

In pixel space the sheets of face images of two different people are interleaved, and a flat separating plane drawn through them fails to split them

Open full-size figure

Redrawn in the last lecture from DiCarlo & Cox (2007)

  • In pixels, two people's face images interleave: no flat (linear) plane separates them
  • "IT encodes animacy" is a claim about information another neuron could use
  • The simplest such neuron: a weighted sum and a threshold, a linear readout

Return to contents

Slide 21

A linear readout draws a hyperplane

A 2-D plane with axes x1 and x2 holding two groups of dots, green "yes" points upper right and blue "no" points lower left. A fitted straight line runs between them and the two sides are lightly shaded. A magenta arrow labelled w starts on the line at a right angle and points into the green side. A dotted segment from a far point to the line is labelled distance

A 2-D plane with axes x1 and x2 holding two groups of dots, green "yes" points upper right and blue "no" points lower left. A fitted straight line runs between them and the two sides are lightly shaded. A magenta arrow labelled w starts on the line at a right angle and points into the green side. A dotted segment from a far point to the line is labelled distance

Open full-size figure

Toy data, computed

  • Response vector 𝐱\mathbf{x}, weights 𝐰\mathbf{w}, bias bb: the readout answers "yes" when 𝐰⊤𝐱+b>0\mathbf{w}^{\top}\mathbf{x} + b > 0
  • 𝐰⊤𝐱=w1x1+w2x2\mathbf{w}^{\top}\mathbf{x} = w_1 x_1 + w_2 x_2, the dot product from the MDS lecture: one weight per unit
  • The boundary 𝐰⊤𝐱+b=0\mathbf{w}^{\top}\mathbf{x} + b = 0: a line in 2-D, a plane in 3-D, a hyperplane in DD dimensions
  • 𝐰\mathbf{w} is perpendicular to it. Each point has a distance to the boundary

Return to contents

Slide 22

Monkey IT: reading animacy

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Green dots are animate images (animals and faces), blue dots inanimate ones. A straight tilted line fitted to the points separates most animate images from inanimate ones

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Green dots are animate images (animals and faces), blue dots inanimate ones. A straight tilted line fitted to the points separates most animate images from inanimate ones

Open full-size figure

Computed: monkey IT, 480 neurons, 2 principal components

  • Bao objects: 51 objects × 24 views, recorded in monkey IT, 480 neurons
  • Axes: IT's first two principal components. Green: animals and faces; blue: inanimate
  • Neither component alone separates them: the best boundary is tilted, a weighted sum of both
  • IT separates animate from inanimate objects (Kriegeskorte et al., 2008)

Return to contents

Slide 25

Measuring with a readout

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Faint dots are the 18 training views per object used to fit the line; solid dots are the 6 test views per object, held out, on which the line is scored

Monkey IT responses to the 1,224 Bao object images on their first two principal components. Faint dots are the 18 training views per object used to fit the line; solid dots are the 6 test views per object, held out, on which the line is scored

Open full-size figure

Computed: monkey IT, 480 neurons, 2 principal components

  • Result: the readout reads animacy correctly for 93% of the views. Is that a lot?
  • Baseline: guessing, 50%; always "inanimate", 63%: the majority-class baseline
  • Split: fit on 18 of each object's 24 views (training set); score on the other 6 (test set). Both 93%: the readout did not overfit

Return to contents

Slide 26

Draw the boundary yourself

Return to contents

Slide 29

Linearly separable, or not

Toy data, two panels side by side at the same size. Left, the XOR clusters in the plane of x1 and x2. Right, the same clusters lifted by a third unit x1 times x2: the green clusters rise above a translucent plane and the blue ones drop below it, so one plane separates them

Toy data, two panels side by side at the same size. Left, the XOR clusters in the plane of x1 and x2. Right, the same clusters lifted by a third unit x1 times x2: the green clusters rise above a translucent plane and the blue ones drop below it, so one plane separates them

Open full-size figure

Toy data, computed: inputs coded −1 and +1

Return to contents

Slide 30

Further reading

Optional: a few research-trail entries over the semester. Any paper mentioned in any lecture qualifies, but papers from this block are preferred.

Return to contents

Slide 31

Minute paper

Submit on Canvas → Minute papers → Minute Paper 6 (access code read out in class).

Write three brief points in your own words:

  1. What does the participation ratio measure?
    If k directions share the variance equally, what is its value?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Return to contents

Slide 32

Appendix: derivations and extra examples

Material not covered in class, for reference

Return to contents

Slide 33

Going further: what does the power law mean?

Return to contents

Slide 34

Monkey V1 falls the same way

The same log-log plot with monkey V1 added: two green lines, one per monkey, for principal components 1 to 60, starting higher and falling a little more steeply than the mouse line, with alpha 1.32 and 1.39

The same log-log plot with monkey V1 added: two green lines, one per monkey, for principal components 1 to 60, starting higher and falling a little more steeply than the mouse line, with alpha 1.32 and 1.39

Open full-size figure

Computed from Stringer et al.'s and Papale et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)
  • Monkey V1: electrode arrays; 100 THINGS images (photos of everyday objects). A little steeper, α=1.32\alpha = 1.32 and 1.391.39

Return to contents

Slide 35

Pixels fall the same way

The same log-log plot with a red line added for the pixels of the 2,800 images the mice saw: it falls almost exactly on top of the mouse V1 line, alpha 1.09

The same log-log plot with a red line added for the pixels of the 2,800 images the mice saw: it falls almost exactly on top of the mouse V1 line, alpha 1.09

Open full-size figure

Computed from Stringer et al.'s and Papale et al.'s data

  • Log-log axes: each tick is ×10, so a power law fk∝k−αf_k \propto k^{-\alpha} is a line of slope −α-\alpha
  • Mouse V1: straight over the measured range, α=1.06\alpha = 1.06 (dotted: slope −1-1)
  • Monkey V1, electrode arrays, 100 THINGS images: a little steeper, α=1.32\alpha = 1.32 and 1.391.39
  • Pixels of the 2,800 images: α=1.09\alpha = 1.09, the same slope. Is V1 just copying the images?

Return to contents

Slide 36

Neurons are not pixels

Log-log plot of variance fraction against principal component k for natural and whitened images. For natural images the pixel spectrum, alpha 1.09, and the mouse V1 spectrum, alpha 1.06, lie almost on top of each other. For whitened images the pixel spectrum is much flatter, alpha 0.41, while the mouse V1 spectrum still follows the dotted line of slope minus 1, alpha 1.09

Log-log plot of variance fraction against principal component k for natural and whitened images. For natural images the pixel spectrum, alpha 1.09, and the mouse V1 spectrum, alpha 1.06, lie almost on top of each other. For whitened images the pixel spectrum is much flatter, alpha 0.41, while the mouse V1 spectrum still follows the dotted line of slope minus 1, alpha 1.09

Open full-size figure

Stringer et al.'s data: 7 natural, 4 whitened sessions

  • Whitened images: all scales, equal contrast
  • Their pixels flatten: α\alpha from 1.09 to 0.41. Mouse V1 does not: 1.06 to 1.09
  • V1's power law is not inherited from the images
  • It is a property of how V1 encodes images (why a slope near 1 matters: two slides on)

Return to contents

Slide 37

A power law across species and methods

Log-log plot of variance fraction against principal component k from 1 to 60 for the same 100 THINGS images in monkey V1, V4 and IT (electrodes) and human visual cortex (fMRI), all with the same cross-validated estimator. All four fall along nearly parallel straight lines, with alpha 1.35, 1.25, 1.24 and 1.49

Log-log plot of variance fraction against principal component k from 1 to 60 for the same 100 THINGS images in monkey V1, V4 and IT (electrodes) and human visual cortex (fMRI), all with the same cross-validated estimator. All four fall along nearly parallel straight lines, with alpha 1.35, 1.25, 1.24 and 1.49

Open full-size figure

Computed here: one estimator, 100 THINGS images

  • Variance falls as a power law: mouse V1; monkey V1, V4, IT; human cortex
  • Calcium imaging, electrodes, fMRI: same shape; the exponent depends on the estimator
  • Shepard's law of generalization (MDS lecture) was such a law for behavior
  • A candidate law of neural coding

Return to contents

Slide 38

Why a slope near 1? A smooth code

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1 responses to 2,800 natural images: seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; alpha 1.06

Log-log plot of variance fraction against principal component k from 1 to 1,000 for mouse V1 responses to 2,800 natural images: seven thin lines for the recording sessions and one thick line for their mean, falling along a nearly straight line close to a dotted guide of slope minus 1; alpha 1.06

Open full-size figure

Computed from Stringer et al.'s data

  • Smooth code: similar images give similar responses, as small changes should
  • Slower than 1/k1/k: fine directions dominate; near-identical images land far apart
  • Stringer et al.'s theory: a slope near −1-1 is about the flattest a smooth code allows. V1 uses as many dimensions as smoothness permits
  • Why that matters: networks are not smooth

Return to contents

Slide 39

Deep networks are fooled by invisible changes

Three images in a row. Left, a public-domain photograph of an alligator lying on grass, which ImageNet-trained ResNet-18 labels alligator with 99 percent probability. Middle, the change added to it, magnified 50 times: colored noise; no pixel moves by more than 2 of 255 intensity levels. Right, the photograph plus the change, which looks identical to the original, labeled ostrich with 100 percent probability

Three images in a row. Left, a public-domain photograph of an alligator lying on grass, which ImageNet-trained ResNet-18 labels alligator with 99 percent probability. Middle, the change added to it, magnified 50 times: colored noise; no pixel moves by more than 2 of 255 intensity levels. Right, the photograph plus the change, which looks identical to the original, labeled ostrich with 100 percent probability

Open full-size figure

Computed here: ResNet-18 trained on ImageNet, a public-domain THINGS photograph

Return to contents

Slide 40

Could a V1-like spectrum make networks robust?

Log-log plot of the variance fraction of each principal component against its number k, for six networks' last layers on 1,854 THINGS images, with a dotted line of slope minus 1. Five trained networks fall more gently than the dotted line, with exponents between 0.67 and 0.96; the untrained ResNet-18, dashed, falls more steeply, exponent 1.43

Log-log plot of the variance fraction of each principal component against its number k, for six networks' last layers on 1,854 THINGS images, with a dotted line of slope minus 1. Five trained networks fall more gently than the dotted line, with exponents between 0.67 and 0.96; the untrained ResNet-18, dashed, falls more steeply, exponent 1.43

Open full-size figure
  • Six networks, computed here: trained ones are flatter than V1 (0.67–0.96 vs 1.06)
  • Kong et al. (2022): networks trained to resist attacks have steeper, more V1-like spectra
  • Nassar et al. (2020): a training penalty pulls a layer's spectrum toward 1/k1/k: harder to fool
  • Open: is robustness why V1 has a power law?

Return to contents

Slide 41

Effective dimensionality from an eigenspectrum

import numpy as np # X: (N stimuli, D units) response matrix -- neurons, voxels or model units Xc = X - X.mean(axis=0, keepdims=True) # center: PCA needs it (the PCA lecture) C = Xc.T @ Xc / (Xc.shape[0] - 1) # the D x D covariance matrix lam = np.sort(np.linalg.eigvalsh(C))[::-1] # its eigenvalues, largest first lam = lam[lam > 0] pr = lam.sum()**2 / (lam**2).sum() # participation ratio, 1 <= PR <= D print(f"ambient D = {X.shape[1]}, effective dimensionality = {pr:.1f}") f = lam / lam.sum() # variance fraction f_k of component k k = np.arange(1, len(f) + 1) alpha = -np.polyfit(np.log(k[10:500]), np.log(f[10:500]), 1)[0] print(f"variance fraction decays as k^-{alpha:.2f}")

Return to contents

Slide 42

More features than units

Two circles of directions: two perpendicular feature directions where each unit means one thing, versus seven directions packed into the same plane, forced to overlap

Two circles of directions: two perpendicular feature directions where each unit means one thing, versus seven directions packed into the same plane, forced to overlap

Open full-size figure

Return to contents