Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 7 · Theme 1: Representational spaces

Geometry, baselines & tests

Tuesday, October 6 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Flexible or abstract: can one population be both?

The geometry of a population: how it arranges its conditions decides whether it is flexible (shattering) or abstract (CCGP)

Reminder: shattering measures flexibility

The shattering figure from the last lecture: two blocks of eight small panels, one per way to color three stimuli blue or red. Top, three stimuli in general position: every panel has a line separating blue from red, 8 of 8. Bottom, three stimuli on one line: 6 of 8
  • Flexible: a downstream neuron can learn many groupings of the same stimuli
  • Shattering accuracy: fit a readout for each split, score it on held-out trials, and average. 0.5 = guessing
  • 3 stimuli in 2 units, in general position: every split is separable, shattering 1.00. On one line, 2 of the 8 splits get at most 2 of 3 right: 0.92
  • In general: NN stimuli in general position in D=N−1D = N - 1 units, every split is separable
  • More dimensions, more splits solved. But it says nothing about stimuli never seen

Reminder: CCGP measures abstraction

A readout for animate against inanimate is fitted on dogs and chairs, then applied to cats and tables. Left, shattering refits a readout for every split of the same objects; right, CCGP fits once and keeps the weights fixed for objects it never saw

  • CCGP (cross-condition generalization performance): fit an "animate or not" readout on dogs and chairs, then test it, weights fixed, on cats and tables
  • A high CCGP means the code for animacy is abstract: not tied to the training examples (dogs, chairs), so it carries over to new ones (cats, tables)

The task: a hidden context

Schematic of one trial: inter-trial interval 1750 ms, hold the button and fixate 400 ms, one of four images A to D for 500 ms, hold or release within 900 ms, a 500 ms wait, then the outcome, juice or nothing. Below, two tables, each row tagged with its condition. Context 1 (red): A release, juice, A plus; B hold, juice, B plus; C release, none, C minus; D hold, none, D minus. Context 2 (blue): A hold, none, A minus; B release, juice, B plus; C hold, juice, C plus; D release, none, D minus. Between them: switch every 50 to 70 trials, no cue, inferred from outcomes

Schematic of the task of Bernardi et al. (2020)

  • Memorized associations are not enough: the monkey must hold context
  • An uncued context flips the rules. Neurons are read before the response: action and reward are from the previous trial

The monkeys infer the hidden context

Bar chart redrawn from Bernardi et al. (2020) Fig. 1c: accuracy in percent, 0 to 100, on the first showing of each image around an uncued context switch, marked by a vertical line. Last image before the switch: 95.0% (95% confidence interval 92.6 to 96.9). After the switch, 1st image, old rule, wrong: 8.7% (6.3 to 11.6). Images 2 to 4, not yet seen in the new context: 56.0% (51.4 to 60.6), 74.2% (69.9 to 78.1) and 87.1% (83.7 to 90.0). A dashed line marks chance, 50%
  • After one surprise, the monkeys switch every rule
  • Accuracy on the last image before a switch, then on the first showing of each image after it
  • 1st image after the switch: about 9% correct. Nothing warned the monkey
  • 2nd to 4th images, not yet seen in the new context: above chance (dashed line, 50%)
  • Bernardi et al. call this inference: it needs a context variable coded apart from any one image
  • Which geometry of the conditions allows this? Shattering and CCGP measure it

Toy geometry: clustered by context

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Abstract, not flexible
  • Red/blue: context 1/2. Gold plane: the decision boundary of a readout that decodes the trial's context, fitted on A+ and C+ (filled) and tested, weights frozen, on C− and A− (hollow)
  • The conditions cluster by context
  • CCGP 1.0: context transfers. Shattering 0.73: conditions within a context cannot be told apart. PR 1.01: almost one direction, context

Toy geometry: points in general position

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Flexible, not abstract
  • Gold plane: the context readout's decision boundary (filled: fitted on; hollow: never seen)
  • N=4N = 4 conditions at random in D=3D = 3 units: general position, so every split is separable (D=N−1D = N - 1). PR 2.77
  • Context can be decoded (shattering 1.0), but CCGP is 0.50: the plane fitted on A+, C+ cuts right through C− and A−

Toy geometry: a square

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context, CCGP reward

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Abstract, but not fully flexible
  • Gold plane: the context readout's decision boundary (filled: fitted on; hollow: never seen)
  • Corners of a square (PR 2.00): one side is context ({A+, C−} vs {A−, C+}), the other reward ({A+, C+} vs {A−, C−})
  • CCGP 0.78 (context), 0.72 (reward). Shattering 0.82: the diagonal split, image A (A+, A−) vs image C (C+, C−), fails. It is context XOR reward

Toy geometry: the square, bent

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context, CCGP reward

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Flexible and abstract at once
  • Gold plane: the context readout's decision boundary (filled: fitted on; hollow: never seen)
  • Bending the square adds a dimension: the A and C conditions move apart along a third direction, so the four no longer fit in a plane (PR 2.00 → 2.45)
  • Shattering 1.0: the diagonal split now separates. CCGP: 0.78 → 0.77. PR helps shattering, not generalization

Try it: flexible or abstract?

Small reminder sketch: a cube whose eight corners are the eight conditions, red for context 1 and blue for context 2, with its edges labelled context, action and reward
  • Mix: cube (0) to general position (1). Shattering rises, CCGP falls, decoding stays high

Hippocampus: context transfers, action does not

Bernardi et al. (2020) Fig. 3e. For hippocampus (HPC), dorsolateral prefrontal cortex (DLPFC) and anterior cingulate cortex (ACC), a vertical axis from 0.4 to 1 labelled shattering dimensionality and CCGP, with a dashed line at 0.5. Short black lines mark the measured CCGP of three variables, each on a set of coloured dots from 100 runs of a perfect-cube model: context in red (HPC 0.96, DLPFC 0.74, ACC 0.81), reward of the previous trial (the paper’s ‘value’) in purple (0.79, 0.88, 0.94) and action of the previous trial in orange (0.44, 0.75, 0.90). Short grey lines mark the measured shattering dimensionality (0.70, 0.75, 0.74). Clusters of white dots, the cube model's shattering dimensionality, sit lower, at about 0.60, 0.62 and 0.65
  • In all three areas context transfers and shattering is high, like the demo's compromise
  • Both scores are average accuracies of linear readouts; chance is 0.5 (dashed)
  • Black: CCGP (red context, orange previous action, purple reward, the paper’s ‘value’). Grey: shattering. Dots: the authors' model
  • Context CCGP: 0.96 HPC, 0.74 DLPFC, 0.81 ACC. Shattering: 0.70, 0.75, 0.74
  • Previous action in HPC: 0.44: below the dashed line, but within chance range (0.41–0.59). Not abstract, yet decodable

Errors: the geometry breaks, the information does not

Bernardi et al. (2020) Fig. 6a and 6b. A: context CCGP on correct trials (red) and error trials (magenta) in HPC, DLPFC and ACC; on error trials it is lower in all three areas, about 0.74 to 0.64, 0.72 to 0.65 and 0.67 to 0.60, each marked significant. B: decoding accuracy for context on correct and error trials in the same areas, about 0.70 to 0.63, 0.77 to 0.71 and 0.72 to 0.67, none significant
  • When the monkeys err, context loses its abstract format but can still be decoded
  • A: context CCGP drops on error trials in all three areas. B: decoding of context, with all conditions in training, does not drop significantly

Recap: measures of one population

Measure What it says about the representation Compared with
Participation ratio How many dimensions it uses the number of units DD
Readout accuracy Whether a variable is explicit majority-class baseline
Shattering accuracy How flexible it is: how many groupings it allows guessing, 0.5
CCGP How abstract it is: one format across conditions chance, 0.5

Why it matters: both read the population's geometry, how its conditions are arranged. Shattering says how many new rules a downstream neuron could learn; CCGP says whether what it learned carries over to new situations, and it drops when the monkeys err.

Each score is one number about one population. When does such a number mean anything?

Baselines and tests

When does a score mean anything?

Two 51 by 51 dissimilarity matrices for the same Bao objects in the same category order, with object thumbnails along the edges, from monkey IT (480 neurons, left) and ResNet-18's late layer (4,096 units, right). Both show a light block for faces and lighter blocks within animals and within vehicles

From the last lecture: correlation-distance RDMs of 51 object means, monkey IT and ResNet-18's late layer

  • IT vs ResNet-18: RSA = 0.71. Is that high? A quantitative claim needs three things: is it statistically significant, is it reliable across samples of images, how close is it to its ceiling?

1. Is the score above chance?

What a meaningless pairing scores: shuffles and random baselines

Correlation by chance: the null distribution

Two overlaid histograms of the correlation between two unrelated lists of random ratings, from 20,000 simulated pairs. With 10 movies (blue) the histogram is wide, spanning about minus 0.6 to plus 0.6, with its 95th percentile at 0.55. With 100 movies (red) it is narrow, about minus 0.2 to plus 0.2, with its 95th percentile at 0.16
  • Unrelated data correlate a little by luck
  • Null hypothesis: the two lists are unrelated. Two people rate 100 movies without watching them
  • Their ratings still correlate by chance, up to about ±0.2; with 10 movies, ±0.6
  • Null distribution: the spread of the score over many unrelated pairs
  • A score is above chance if it beats the null distribution's 95th percentile: only 5% of unrelated pairs score higher

Above chance? Shuffle the images

Horizontal bars of RSA between six networks and one person's IT measured with fMRI on 100 THINGS images, all trained networks between 0.16 and 0.19 and the untrained one near 0.06. A grey band near 0.02 marks the 95th percentile of the shuffled-image scores; a dashed line marks another person's IT and a dotted line the split-half noise ceiling, both near 0.24

Computed: THINGS-fMRI, one person's IT (3,720 voxels), 100 images; six networks; 1,000 shuffles

  • Only the part of a score above its shuffled baseline comes from matching the images
  • Permutation (shuffle) test: shuffle one system's image order, recompute the score, repeat 1,000 times. This builds the null distribution from the data themselves
  • A score passes if it lies above the 95th percentile of the shuffled scores. Monkey IT vs ResNet-18 passes: 0.71 against 0.055. So does every trained network in the figure
  • Dashed and dotted lines: another person's IT and the noise ceiling (later today)

Readouts: shuffle the labels

Two panels for a readout of animate versus inanimate from 480 monkey IT neurons, trained on 102 images and tested on 306. Left: accuracy bars; with true labels, training 1.00 and test 0.96; with shuffled training labels, training 1.00 and test 0.56 on average, below a dashed majority-class line at 0.63. Right: histogram of the test accuracy over 1,000 label shuffles, centred near 0.55, with its 95th percentile marked near 0.67 and the true-label test accuracy, 0.96, far to the right

  • Bao et al. (2020) IT, 480 neurons. Animate vs inanimate, fit on 102 images, tested on 306
  • Shuffled labels: 1.00 on training, chance on test (0.56). So we need separate test images
  • Curse of dimensionality: more neurons (480) than training images (102), so any labels fit. Deep networks too (Zhang et al. 2017)

Random directions in many dimensions

Histograms of the cosine between two random vectors for 2, 10, 100 and 1,000 dimensions, and a curve of the mean absolute cosine against the number of dimensions: in 2 dimensions the cosines spread from minus 1 to 1; as the dimension grows they pile up around 0, and the mean absolute cosine falls roughly as one over the square root of the dimension

Computed: random Gaussian vectors, 4,000 pairs per dimension

  • Within one system: each stimulus is a vector of DD unit responses. In 1,000-D, two random stimulus vectors are nearly perpendicular: mean ∣cos⁡ϕ∣|\cos\phi| 0.026 (2-D: 0.64)
  • More units than stimuli: all nearly perpendicular, so 102 images in 480 neurons fit any labels

Few stimuli: unrelated systems can score high

Linear CKA between two matrices of random Gaussian numbers, with more units than stimuli, against the number of stimuli on a log axis: near 1 with ten stimuli and falling as stimuli are added, although the two matrices share nothing

  • Linear CKA compares two systems' Gram matrices (dot products between stimuli). Two unrelated random systems, 4,096 units: near 1 with 10 stimuli
  • Few stimuli, many units: each Gram matrix is mostly its diagonal, alike in any two systems

2. Would it hold on new images?

Two kinds of new images: images the fit never saw (cross-validation), and another sample of images (bootstrap)

Hold out every image once: cross-validation

Left, a schematic of 5-fold cross-validation: five rows labelled fit 1 to fit 5, each a strip of five folds; in each row a different fold is red and marked test and the other four are grey, so every fold is tested exactly once. Right, for one monkey IT site predicted from a network's features, each of the 1,224 images' out-of-fold prediction against its recorded response, scattered along the diagonal, R-squared 0.60 over all images

  • 5-fold cross-validation (left): 5 folds; fit on four, test on the fifth, rotate. Every image is tested once, by a fit that never saw it
  • Right: one IT site predicted from ResNet-18 (fit next lecture). Score R²: here the squared correlation of prediction and response (r=0.77r = 0.77, R2=0.60R^2 = 0.60)
  • One split can be lucky: one fold alone gives R² 0.53 to 0.64; all five together, 0.60

Leakage: keep the test images out of every step

  • Leakage: the test images influence some step of the fit. The test score is then too optimistic
  • Every step is fit on the training images only: the z-scoring mean and standard deviation, the PCA directions, any setting (such as a penalty strength), and the weights
  • Small leak: z-scoring with the mean of all images. Large leak: choosing units or settings by their test score; then even pure noise can look predictive
  • Split first, then fit every step on the training images

Note: in our assignments the IT recordings come already cleaned, with missing values filled and each unit centred and scaled, using all 1,224 images before any split. That is a small leak, accepted to keep the assignments simple. To publish, do these steps on the training images only.

Would the ranking hold on a new sample of images?

Top: overlapping distributions of RSA with human IT for six networks over resampled image sets, the trained ones spanning roughly 0.11 to 0.25 and the untrained one near 0.06. Bottom: across 1,000 resamples of the 100 images, ResNet-50 ranks first in 55 percent, CLIP in 18 percent, DINOv2 in 16 percent, ResNet-18 in 8 percent, ResNet-50 DINO in 4 percent and the untrained network in 0 percent

Computed: the same THINGS-fMRI data, 1,000 resamples

  • Cross-validation tests a fit on new images. But any score, even RSA with no fit, depends on which images the study used
  • Bootstrap: draw 100 of our 100 images with replacement (some twice, some never), and rescore every model; repeat 1,000 times
  • Criterion: compare two networks resample by resample. One clearly wins if it scores higher in at least 95% of resamples. ResNet-50 vs CLIP: 77%; vs the untrained network: 99%
  • Verdict: all trained networks beat the untrained one; none clearly beats another trained one

3. How high could the score be?

Noise in the recordings sets a ceiling

Brains are noisy, networks are not

Responses of one monkey IT electrode on each of 30 repeats of three THINGS images, a chipmunk, a stalagmite and a pan, with a dashed line at each image's mean. The chipmunk responses sit far above the others. The stalagmite's mean is slightly above the pan's, but the two sets of single responses overlap, and the pan comes out above the stalagmite on 9 of the 30 presentations. Beside it, one ResNet-18 unit whose response to each image is identical on every repeat, in the same order

  • 30 repeats of one image: an IT electrode responds differently each time; a network unit does not
  • Pan's mean (dashed) is below stalagmite's, yet pan comes out above it on 9 of 30 presentations
  • Noise caps every score. The split-half reliability, how well the means of two halves of the repeats agree across images, says how much of the recording repeats

How high could a score be? The noise ceiling

Line plot of the CKA between two separate halves of a monkey's IT recordings, as more repeats are averaged in each half, 1, 2, 4 and 8, for two monkeys. Agreement rises toward 1 as repeats are added, to 0.95 for monkey N and 0.92 for monkey F with 8 repeats per half, and the spread across random splits narrows

Computed: IT electrodes of two monkeys, 100 THINGS images, 20 random splits per number of repeats

  • A model's score: CKA between the model and the IT recording
  • Its ceiling: the same CKA between two halves of the recording, how well the brain agrees with itself. More repeats per half, closer to 1
  • ResNet-18 vs IT, all repeats: 0.19 (monkey N), 0.36 (F). The halves agree at 0.95 and 0.92: the model is far from the ceiling
  • Read a score against its ceiling: 0.3 is near the top if the ceiling is 0.35

The same area in another brain

A 6 by 6 design grid: V1, V4 and IT of monkey F and of monkey N along both axes. Cells are labeled by what they compare: the same area in the same animal (the diagonal), the same area in the other animal, and different areas; no values are shown

Design only: 100 THINGS images × 30 repeats in V1, V4 and IT of two monkeys (Papale et al. 2025)

  • Score all pairs of recordings. Reference for "brain-like": the same area in another animal
  • Report a score between its references: the shuffled floor, the noise ceiling, another brain

Change the measure, and the winner changes

Slope plot ranking seven networks by similarity to monkey IT under two measures. CKA ranks ResNet-18 first, ResNet-50 second and CLIP third; RSA puts CLIP first, ResNet-18 second and ResNet-50 fourth. The untrained ResNet-18 is last under both

Computed: monkey IT (Bao et al.) vs seven networks, 51 object means, units z-scored

  • Seven networks ranked by similarity to monkey IT on the same 51 objects. Only the measure changes
  • RSA puts CLIP first; CKA puts ResNet-18 first
  • The top networks nearly tie, so changing the measure reorders them

Recap

  • Before trusting a score: its shuffled floor, a test on held-out images, its noise ceiling and the same area in another brain; then resample the images and try more than one measure
  • Next lecture: one model neuron, and how its weights are learned

Further reading for the research trail

Trail entry 2, due Fri 10/9: pick any paper cited in a lecture so far, read it, and post four to six sentences on Ed under Research trail: what it did, what surprised you or the question it left you with, and what you would test next. Reply once to a classmate. Best three of four entries count.

Minute paper

Submit on Canvas → Minute papers → Minute Paper 8 (access code read out in class).

Write three brief points in your own words:

  1. A network's RSA with monkey IT is 0.20. Name three reference scores you would compute
    before calling it high or low, and say what each one tells you.
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

- The reader-friendly version linked above carries every slide in reading order, with a description of each figure

- Thursday stopped after the two CCGP slides, then jumped to the comparison scores (RSA, CKA). Today opens with a reminder of both scores and the shortest route to the trade-off between them - The path today: the task and the monkeys' behavior, four toy geometries (clustered, general position, square, bent square), the cube demo, then the recordings and what happens on error trials (Bernardi et al. 2020)

- This repeats the last lecture's slide on shattering three stimuli. The rule generalizes: N stimuli in general position in D = N − 1 units can be split every way by a hyperplane - The splits are arbitrary: the score counts the groupings a downstream readout could learn from the population, whatever the population is used for - Bernardi et al. (2020) count only splits into two equal halves, so that guessing is 0.5 for every split - In the assignment the splits are random halves of the 51 objects, and each readout is scored on held-out views - On a line, the split that puts the middle stimulus against the two ends cannot be drawn: the best line gets 2 of the 3 right. Two splits score 2/3, six score 1: (6 + 2 × 2/3) / 8 = 0.92 - "Dimensions used" on the toy figures later today is the participation ratio (PR): 1 when the conditions lie along one direction, 2 when they spread evenly over a plane, up to 3 for four conditions in three units

- Shattering counts how many splits of the same stimuli a readout can separate. CCGP tests whether a readout trained on some objects works on new objects - On the left of the figure, the same points sit in the plane of two units, with three boundaries, one per split. Shattering fits a new readout for every split - The figure's score is the shattering accuracy of the previous slide: the readout accuracy for each split, on held-out trials, averaged over splits - On the right, one readout is fitted on the filled points, then frozen and scored on the hollow points, conditions never used to fit it. That score is CCGP, cross-condition generalization performance - Train an animacy readout on dogs and chairs. Then test it on cats and tables without changing its weights - A high score means animacy is coded along the same direction for all these objects - In the assignment, the split holds out half of the animate and half of the inanimate objects, and the score is averaged over random halvings - The assignment compares CCGP with the majority-class baseline instead of 0.5, because its two classes have different numbers of objects

- Bernardi et al. (2020) trained two monkeys on this task - On each trial the monkey holds a button and fixates. A fractal image appears for 500 ms. The monkey releases the button within 900 ms (R) or keeps holding (H), then gets juice (+) or nothing (−) - Call the images A to D from top to bottom. In context 1, A means release for juice, B hold for juice, C release for nothing, D hold for nothing - In context 2 every image flips its action, and two images also flip their reward. So action and reward are not tied together - A correct response to an unrewarded image avoids a timeout and a repeat of the trial, so the monkey still has a reason to get it right - Context is hidden. The rules flip every 50–70 trials without warning, and nothing on the screen says which context is active - There are eight conditions, 2 contexts × 4 images. Each image fixes the action and the reward, so the conditions are context × action × reward - The neurons are analysed from 800 ms before the next image to 100 ms after it appears. They still hold the previous trial's context, action and reward - So A+ and C− happen only in context 1, and A− and C+ only in context 2 - Each variable (context, action, reward) splits the 8 conditions into two classes, 4 against 4

- Figure: redrawn from Bernardi et al. (2020), Fig. 1c; values extracted from the published figure - The bars count only the first time each image appears after a switch - The first image after a switch gets the old response, so it is almost always wrong. An unexpected outcome is the only sign that the context changed - Images 2 to 4 have not been seen in the new context. A monkey that relearned each image by trial and error would be at chance on them. These monkeys are above chance, so one surprise changes their response to the other images too - Bernardi et al. (2020) call this inference. It needs a variable for context that is separate from any one image

- The gold plane is the decision boundary of our analysis readout, not the monkey's decision. It decodes the context, 1 or 2, from the neurons: a variable the monkey is never shown - CCGP: fit on one condition per context (here A+ and C+), score with weights frozen on the other two (C− and A−). A+ and C+ is one of four such choices - The plane fitted on A+ and C+ classifies C− and A− perfectly, because each lies next to its context partner - The four choices all score 1.00, so CCGP is 1.0 - Splits that separate conditions within a context fail, which pulls shattering down to 0.73 - Dimensions used (PR) 1.01: almost all of the variance lies along one direction, context (99.3%); the small differences within each context add the 0.01. One direction supports one split, which is why only the context split works - How the toy numbers are computed. Each of the four conditions is a mean point in 3 units plus 15 simulated trials: the mean plus Gaussian noise (s.d. 0.10 here). Every readout is a logistic regression - Shattering: for each of the three balanced splits, fit on half of each condition's trials and test on the other half, then swap; repeat over 15 random halves and average. CCGP: fit on all trials of one condition per context, test on all trials of the other two; average over the four choices

- Four conditions at random positions in 3 neurons, with noisy trials. Random placement puts them in general position with probability 1 - Decoding and transfer are different questions. Context is decodable: trained on trials of all four conditions, a plane separates red from blue perfectly, as for every other split (shattering 1.0). CCGP asks something else: does a plane fitted on only two conditions work on the other two? - Here it does not. The drawn plane is fitted on A+ and C+, and C− and A− happen to lie on it, so about half of each one's trials fall on each side: 53% and 47% correct, 50% in all. The readout separates A+ from C+ but does not transfer to C− and A− - The four choices (fitted on → tested on: accuracy) - A+, C+ → C−, A−: 0.50 (the plane drawn) - A+, A− → C−, C+: 0.50 - C−, A− → A+, C+: 0.50 - C−, C+ → A+, A−: 0.50 - Mean: CCGP 0.50, chance - Dimensions used (PR) 2.77, close to the most four points can use (3). With every condition in its own direction, every split is separable, and no direction is shared between them - Same simulation as the clustered toy (four conditions, 15 trials each, logistic-regression readouts), with noise s.d. 0.04

- The gold plane is the decision boundary of the context readout fitted on A+ and C+ only. Any plane between the trials of those two conditions would do, so the fit lands slightly tilted from vertical (about 14°). It still classifies C− and A− correctly; the 3-D view exaggerates the tilt - Shattering here is the mean of three balanced splits: context 1.0, reward 1.0, diagonal 0.46. For images A and C, action is tied to context (release in context 1, hold in context 2), so it is not a separate split here - Between the two extremes: give each variable its own direction. With two variables, the four conditions sit at the corners of a square. This is the XOR layout of the last lecture. Context and reward are two perpendicular directions, so a readout for either transfers. Of the three balanced splits, the diagonal one (A+ and A− against C+ and C−) cannot be separated by a plane - Dimensions used (PR) 2.00, exactly: the four corners lie in a plane. Two directions give the two variables, and nothing is left for the diagonal split - Same simulation as the clustered toy, with noise s.d. 0.02

- Bending adds one dimension to the population, as the lifting unit $x_1x_2$ did for XOR: the embedding dimension goes from 2 to 3. The PR rises only to 2.45, not 3, because the bend is smaller than the sides of the square: the third direction carries less variance than the other two. In neurons, that extra direction comes from nonlinear mixed selectivity: units that respond to a combination of context and reward, here image A vs image C - PR helps shattering but says little about generalization. Generalization needs a coding direction shared across conditions, whatever the number of dimensions: the clustered toy has PR 1.01 and CCGP 1.0, the square 2.00 and 0.78, the bent square 2.45 and 0.77, general position 2.77 and 0.50. The demo shows the same: the cube has PR 3.00 and CCGP 0.95, the compromise 4.14 and 0.80 - PR and shattering go together across the four toys: PR 1.01 → shattering 0.73 (clustered), 2.00 → 0.82 (square), 2.45 → 1.00 (bent), 2.77 → 1.00 (general position). More dimensions, more splits a plane can separate. They are not the same measure: PR weights directions by their variance and ignores the noise, while shattering also depends on how small the trial noise is along the directions that separate each split. Bernardi et al. (2020) note that shattering is correlated with, but not identical to, dimensionality counted from principal components - The context and reward directions stay nearly parallel across conditions, so the readouts still transfer. This is the geometry Bernardi et al. (2020, Fig. 2d) propose for the recordings. Next: add the third variable, action, and the square becomes a cube - Same simulation as the clustered toy, with noise s.d. 0.022

- From the square to the cube: add the third variable, action. The four conditions of the toys become eight, at the corners of a cube, and the mix slider sets how much the cube is bent - Toy model, not fitted to data: eight conditions, each the mean response of 30 model units, plus trial-to-trial noise. Each condition has 20 training and 20 test trials, with Gaussian noise σ = 0.75 by default (the noise slider). Readouts are ridge regressions onto ±1 labels. Shattering averages all 35 balanced splits of the eight conditions; CCGP averages the 16 ways to hold out one condition per side. They are named as in the task, image and outcome; red is context 1, blue context 2 - Why a cube (Bernardi et al. 2020, Fig. 2c, a factorized geometry): three two-valued variables (context, action, reward) make 2 × 2 × 2 = 8 conditions. If each variable has its own direction in the population, the eight conditions sit at the corners of a cube. This is the abstract geometry; the square of the last lecture is the same idea with two variables - Mix = 0, the cube. Take any condition and its partner that differs from it only in context: same action, same reward, other context - In a cube, the step from a condition to its partner is the same step along the same edge, whichever pair you pick. On the map, the gold segments join these pairs: in the cube setting they are horizontal, parallel and of equal length - The map: horizontally, a context readout fitted on six conditions (filled), with its boundary; the two hollow conditions were held out. Vertically, action and reward, plus the direction in which the gold segments differ once the cube bends. The inset (Dimensions used) shows the variance of the eight means along each direction: three directions for the cube, participation ratio 3.00. It is the same PR as under the toy cubes: there it went 1.01 (clustered), 2.00 (square), 2.45 (bent), 2.77 (general position); here 3.00 (cube), 4.14 (compromise) and 7.00 (general position, the most eight points can use) - So one plane, perpendicular to that step, puts every red corner on one side and every blue corner on the other. A context readout fitted on some conditions works on the ones it never saw (CCGP about 0.95) - The cube's limit: groupings that mix the variables, such as context XOR reward, cannot be split by one plane (shattering about 0.72) - Mix = 1, general position: each condition moves to its own random direction. The steps between partners now point in different directions, so the gold segments are no longer parallel, and a plane that separates some pairs need not separate the others. Almost any grouping can be split (shattering about 0.99), but nothing transfers (CCGP about 0.5) - In between, each condition sits part of the way: (1 − mix) times its cube corner plus mix times its random direction. The random directions are drawn twice as far from the centre as the corners, so a small step already counts: at mix = 0.25 both shattering and CCGP are about 0.8. On the map the gold segments tilt a little, each its own way, and the held-out pair still lands on the correct sides; the inset gains four small bars (participation ratio 4.14). Bending adds dimensions without destroying the shared context direction. Bernardi et al. (2020, Fig. 2d) show the same: distort the cube enough and shattering becomes maximal while CCGP stays high, the pattern they measured in hippocampus and prefrontal cortex - All readouts are linear, fitted on training trials and scored on test trials. Decoding uses all eight conditions in training, so it never has to transfer, and it stays above 0.93 in every setting. Only CCGP tells the settings apart - The demo is on the course site with the other demos

- Paper: Bernardi et al. (2020), The geometry of abstraction in the hippocampus and prefrontal cortex, Cell 183:954–967. Open copy: https://pmc.ncbi.nlm.nih.gov/articles/PMC8451959/ - The window is the 900 ms ending 100 ms after the next image appears, before the response: the current trial's action and reward are not known yet, so the variables are the previous trial's. Context carries over across trials - Panel: Bernardi et al. (2020), Fig. 3e - Only the short lines are data. Black lines are CCGP, grey lines are shattering accuracy (the axis label says shattering dimensionality; it is the same average accuracy as on the earlier slides). The coloured and white dots come from the authors' model, not shown here - CCGP for context: the readout is trained on 3 conditions per context and tested on the held-out pair, one from each context. There are 4 × 4 = 16 ways to pick that pair, and the score is averaged over them - Values read from the figure: context CCGP is 0.96 in HPC, 0.74 in DLPFC and 0.81 in ACC. Reward is 0.79, 0.88 and 0.94. Previous action is 0.44, 0.75 and 0.90. Shattering is given in the paper's legend: 0.70, 0.75 and 0.74 - CCGP is compared with a geometric null model, which scores randomly arranged conditions. Its ±2 s.d. band is about 0.41–0.59. Context is above it in all three areas - Previous action in the hippocampus: CCGP 0.44, the orange cluster under the dashed line. It is below 0.5 but inside the null band, so it is not significantly different from chance: the four conditions on each side share no action direction. Action can still be decoded when all conditions are in the training set (Bernardi et al. 2020, Fig. 3a). The paper: action "was not in an abstract format in HPC despite being decodable using traditional methods" - A CCGP well below the null band would mean something else: the readout fitted on three conditions per side would classify the held-out pair the wrong way round, so the action direction would reverse across conditions

- Panels: Bernardi et al. (2020), Fig. 6a–b - Error trials are rare, so this analysis uses its own sample: 180 neurons per area, resampled. Its numbers are not directly comparable with the previous slide's - Average drop in context CCGP on error trials: 0.107 in HPC, 0.069 in DLPFC, 0.065 in ACC, all significant. Average drop in decoding accuracy: 0.069, 0.063 and 0.053, none significant - Decoding asks whether context can be read out. CCGP asks whether it is coded the same way across conditions, which is what transfer to new situations needs. Only the second one goes with the monkeys' errors - Why would the geometry break? The paper does not say, and it does not mention attention. These are not the errors right after a switch: errors within 5 trials of a switch were left out, so these are slips while the rules are stable. Possible reasons, none tested here: a lapse of attention or engagement on that trial, a change in arousal that changes how strongly neurons respond, or a moment of doubt about which context is in force - It is a correlation: the drop in CCGP could cause the error, or both could follow from a third factor such as a lapse of attention

- The four measures describe one set of responses at a time, such as monkey IT, the pixels, or one network layer - Participation ratio is compared with the number of units as PR/D. A variable is explicit when it is available to a downstream neuron, i.e. a linear readout can recover it. CCGP is compared with chance, 0.5, or with the majority-class baseline when the classes differ in size - A representation can score high on one measure and low on another. Flexibility and abstraction pull against each other; the recordings of Bernardi et al. (2020) score high on both - Last Thursday also covered two scores that compare two populations, RSA and CKA. The rest of today asks what any of these numbers has to beat

- The last lecture gave one number per question: readout accuracy, shattering, CCGP, RSA and CKA. Each is one number. Before reading it as "alike" or "decodable", we need to know what a meaningless pairing scores, whether a new sample of images would give the same answer, and how high the noise in the recordings lets a score go - The figure is the pair of RDMs from the last lecture: correlation distances (1 − r) between the 51 object means, in monkey IT (480 neurons) and in ResNet-18's late layer (4,096 units). Their RSA is 0.71 - The rest of today takes the three questions in order. Significance: shuffles give the null distribution. Reliability: cross-validation for a fitted model, the bootstrap for any score. Ceiling: the noise ceiling

- First question of three: what would the score be if there were nothing to find? Shuffles give that baseline for RSA and for readouts

- Simulation: 20,000 pairs of lists of independent random numbers, 10 or 100 items each. For 100 items the correlation has a standard deviation of about 0.1, so 95% of pairs fall within ±0.2; the 95th percentile is 0.16. For 10 items: ±0.6, 95th percentile 0.55 - Fewer items, more luck: a correlation of 0.4 means nothing with 10 movies and is far above chance with 100 - The 95th-percentile rule is a convention: we accept being fooled by luck 5% of the time - The next slide builds the null distribution from the data themselves, by shuffling

- The shuffle measures the luck for the actual data, with their quirks included. The 4,950 entries of an RDM come from only 100 images, so they are far from 4,950 independent pieces of evidence, and the luck is larger than that count suggests - Shuffling keeps everything about each representation (its units, its spread, its dimensionality) and breaks only the pairing of images. The shuffled scores show what those properties produce on their own, with no match between images - The 95th percentile is the value below which 95% of the shuffled scores fall. If the two systems did not match at all, a score above it would come up less than 5% of the time. This is the permutation test of Nili et al. (2014) - The last lecture's monkey IT vs ResNet-18 RSA, 0.712, against a 95th percentile of 0.055 over 1,000 shuffles of the object order - The figure: RSA between six networks and the IT of one person scanned with fMRI (Hebart et al. 2023), on 100 THINGS images averaged over 12 repeats. The trained networks score 0.16 to 0.19, the untrained one 0.06, and the shuffled scores stay near 0.02. The dashed and dotted lines return later in the lecture - You met the same idea with MDS. In the first assignment you refit MDS to a shuffled dissimilarity matrix, which has no structure to recover: 2-D stress 0.446, against 0.204 for the real one

- Training accuracy alone says nothing: with random labels the readout still reaches 1.00 on the training images. Only accuracy on separate test images, never used in the fit, shows what it learned. With true labels: 0.96 on test - As in the shattering reminder at the start of today: N points in general position in N − 1 dimensions can be split in every way. 102 training images in the space of 480 neurons are such points, so a readout can fit any labels, random ones included - This is one form of the curse of dimensionality: with more dimensions (neurons) than training examples, a readout has enough free weights to fit any labelling. The usual remedies are more training images, fewer dimensions (e.g. a few principal components) or regularization, and in every case a score on held-out images - **Generalization gap** = training accuracy − test accuracy. With true labels: 1.00 − 0.96 = 0.04, so what the readout learned holds on new images. With shuffled labels: 1.00 − 0.56 = 0.44: the readout memorized the 102 training images and learned nothing that carries over - Computed from the Bao et al. (2020) recordings: the geometry lecture's animate vs inanimate readout, now on all 480 IT neurons instead of two principal components, trained on 2 views of each of the 51 objects (102 images) and tested on the 306 test views. True labels: training 1.00, test 0.96. Over 1,000 shuffles of the training labels: training 1.00 every time, test 0.56 on average, 95th percentile 0.67 - Guessing at random scores 0.50 on average, and always answering "inanimate" scores 0.63, the majority-class baseline of the geometry lecture. On test images, a readout fit to shuffled labels does no better than guessing - Shuffling the labels keeps the neural responses and the number of animate and inanimate images, and changes only which label goes with which image. The test accuracy after a shuffle is what a readout scores when responses and labels are unrelated - The previous slide did the same for RSA: shuffling the image order in one RDM keeps the structure of each RDM but breaks the pairing between the two - Zhang et al. (2017) trained deep image networks on labels shuffled at random; they still reached 100% training accuracy

- Mean |cos φ| of two random Gaussian vectors: 0.637 in 2-D, 0.263 in 10-D, 0.081 in 100-D, 0.026 in 1,000-D, 0.013 in 4,096-D, close to √(2/(πD)), where D is the number of dimensions - The cosine of the angle φ between two vectors is their dot product divided by both lengths (the MDS lecture): 1 when they point the same way, 0 when they are perpendicular - Here the comparison is within one system: the cosine is between the response vectors of two stimuli, each with one number per unit. Comparisons between two systems come on the next slide - One form of the curse of dimensionality: with many dimensions and few samples, every sample looks about equally far from every other - The same fact makes high-dimensional spaces roomy: many nearly independent directions fit, so a population of many units can carry many variables at once

- Computed: two independent random matrices of N stimuli × 4,096 units: CKA 0.998 with N = 10, 0.976 with 100, 0.804 with 1,000, 0.506 with 4,000 - The previous slide was within one system; this one compares two. The Gram matrix of a system (last lecture) is the N × N table of dot products between its stimuli's response vectors. When the stimuli are nearly perpendicular and of similar length, the off-diagonal entries are near 0 and the diagonal holds each stimulus's squared length. After centering, any two such systems have nearly the same Gram matrix, so their CKA is close to 1 whatever the two systems are - The high score comes from having few stimuli. Each added stimulus adds a row of off-diagonal entries, which differ between unrelated systems and pull CKA down - The conditions: more units than stimuli, response vectors of similar length, and responses spread over many directions - RSA compares only the off-diagonal entries (the distances between different stimuli), so its shuffled baseline stays near 0 at every number of stimuli. With few stimuli the real RSA score is noisy instead - Huh et al. (2024) argued that large networks converge on one representation; Gröger et al. (2026) showed that much of that convergence disappears once every score is compared with its shuffled baseline

- A fitted model, such as a readout or a regression, must be scored on images it was not fit on. Cross-validation does that, and leakage is how it fails - Any score, fitted or not, also depends on which images the study happened to use. The bootstrap estimates how much it would change with another sample - The two answer different questions: cross-validation scores a fit on images it never saw; the bootstrap measures how much any score, fitted or not, varies with the sample of images

- The label-shuffle slide showed that a readout can fit even random labels, so only images held out of the fit tell us what it learned. Cross-validation holds every image out once - The numbers right of the rows: each fit's R² on its own test fold, what a single train/test split would report (0.53 to 0.64). Cross-validation pools all 1,224 held-out predictions: 0.60 - R² is the fraction of the site's variance across images that the predictions explain: 1 is perfect, 0 is no better than predicting the site's mean response. Here it equals the squared correlation r² between prediction and recorded response (r = 0.77). It can differ from r², and even go negative, when predictions are systematically off - Each view of an object is its own image, so test folds hold new views of known objects, not new objects

- When test information reaches the fit, it is called leakage. Z-scoring with the mean of all images, or fitting PCA on all images, lets the test images shape the features the model sees. The effect is often small. Choosing settings, or the most informative units, on the test images can be large: a readout can then look good on pure noise - With cross-validation, this holds inside each fold: the four training folds set the z-scoring, the PCA and the settings, and the fifth fold is only scored - The note at the bottom: the shipped IT file (responses_cleaned) was prepared once, on all 1,224 images: each unit's missing entries filled and the unit centred and scaled. Every split you make afterwards inherits that small leak. With 1,224 images, a mean and a spread estimated from all of them barely differ from those estimated from 80% of them, so the effect on your scores is tiny; the readouts in the assignment z-score again with the training rows only. In a study you publish, do every step inside the training images

- Computed on the human-IT data of the shuffle slide: ResNet-50 first in 55% of resamples, CLIP 18%, DINOv2 16%, ResNet-18 8%, ResNet-50 DINO 4%, the untrained network 0.3% (drawn as 0%) - The criterion is a paired comparison: in each resample, subtract one network's RSA from the other's. If the difference is above 0 in at least 95% of resamples (equivalently, its 95% interval stays above 0, here as a one-sided test), the first network clearly wins. ResNet-50 minus CLIP: higher in 77% of resamples, difference −0.036 to +0.083; minus DINOv2: 75%, −0.038 to +0.071; minus ResNet-18: 87%; minus ResNet-50 DINO: 93%; minus the untrained network: 99.4%, +0.043 to +0.215 - Why paired: both networks are scored on the same resampled images, so a hard image set lowers both. Comparing their two separate intervals would ignore that and be too strict - How to read it: if ResNet-50 were really the best match, nearly every resample (say 95% or more) would put it first. At 55% it is the most likely winner, but the gaps between the trained networks are within what a different sample of 100 images could change. That difference does hold: in 98–99% of resamples, every trained network scores above the untrained one - An image drawn twice appears twice in the RDM; the pairs of an image with its own copy are left out, since their distance is zero in every RDM - For a fitted model, the two can be combined: cross-validate within each resample - Resampling tests replication with new images of the same kind; it does not test other kinds of images - Schütt et al. (2023) resample images and participants together

- Third question: a brain recorded twice does not agree perfectly with itself, so no model can score 1. The noise ceiling says how high a score could go

- Split-half reliability is the split-half correlation of the first lecture, which made the point with one image: single presentations of the same image correlated at 0.17, averages of 25 at 0.69 - Computed from the Papale et al. (2025) THINGS recordings: 100 images × 30 repeats - The figure: the most reliable IT electrode of monkey N on each of the 30 repeats of three THINGS images, beside one unit of ResNet-18's last layer, identical on every repeat. Dashed lines are the means over the 30 repeats - The chipmunk is the electrode's strongest image. Stalagmite (mean 0.18) and pan (mean −0.09) differ on average, but single presentations swap their order 9 times in 30. One presentation cannot rank two images; averaging over repeats can - Even this electrode, the most reliable one, is this noisy; a typical electrode is noisier - No model can be expected to predict a recording better than the recording predicts itself - Averaging more repeats lowers the noise in the mean, which is why the split-half agreement rises with the number of repeats (next slide)

- Spearman–Brown corrects a split-half score to the full number of repeats - One recording has a different ceiling for each measure: compare a CKA score with a CKA ceiling, an RSA score with an RSA ceiling - Computed: for each number of repeats R per half, two disjoint random sets of R repeats are averaged and compared with CKA, over 20 random splits. ResNet-18 vs IT is the CKA between the network's 512 late-layer features and the IT responses averaged over all 30 repeats, on the same 100 images - The ceiling is a property of the recording, not of any model: the model is not on this plot. It is scored against the recording, and that score is read against the ceiling - A caution for CKA: noise in a recording averaged over few repeats spreads the images' response vectors in random directions, which makes them more nearly perpendicular and can raise a model's CKA a little (the same effect as on the few-stimuli slide). ResNet-18's CKA with monkey N's IT is 0.24 with one repeat and 0.19 with all 30. Score models against well-averaged recordings - A score of 0.3 means something different when the ceiling is 0.35 than when it is 0.95: with a low ceiling, the recording is too noisy to tell models apart - In the human fMRI figure earlier, the dotted line was the noise ceiling (0.24), measured with RSA between two halves of the 12 repeats; every network sat between the shuffled scores and the ceiling

- No neuron is matched across animals; the score compares how each recording arranges the images. The grid shows the design only: which recordings are compared with which, and what each comparison is for. On the diagonal, a recording compared with itself scores 1; how far below 1 a comparison between two halves of the same recording falls is its split-half reliability - Safaie et al. (2023): the structure of motor-cortex population activity is preserved across animals doing the same task; Perich (2025) reviews such cross-animal comparisons - A network scored against one monkey's IT should be read next to the other monkey's IT scored against it. Another animal is a reference, not a ceiling: noise lowers both recordings, so correct both scores for noise before comparing - In the human fMRI figure earlier, the dashed line was this reference: another person's IT

- Computed: monkey IT (Bao et al. 2020) against seven networks on the 51 object means, with RSA and linear CKA. RSA moves CLIP from third to first and ResNet-50 from second to fourth. The untrained network is last under every measure - Report more than one measure, say which changes your measure ignores, and treat a conclusion that holds under only one measure as provisional. Soni et al. (2024) find the same disagreement across measures, datasets and model families

- Scores: a floor (shuffled images or labels), held-out images and cross-validation, a ceiling (split-half noise ceiling, another brain), and robustness (resampling, several measures)

- Bernardi, Benna, Rigotti, Munuera, Fusi & Salzman (2020), The geometry of abstraction in the hippocampus and prefrontal cortex - Schütt, Kipnis, Diedrichsen & Kriegeskorte (2023), Statistical inference on representational geometries - Safaie et al. (2023), Preserved neural dynamics across animals performing similar behaviour - Papale, Wang, Self & Roelfsema (2025), An extensive dataset of spiking activity to reveal the syntax of the ventral stream. The monkey recordings behind the noise-ceiling slides - Zhang, Bengio, Hardt, Recht & Vinyals (2017), Understanding deep learning requires rethinking generalization - Before posting, skim the existing entries on your paper; your post must add something not already said

- Submitted on Canvas (Minute papers → Minute Paper 8); credit for a thoughtful attempt