- The reader-friendly version linked above carries every slide in reading order, with a description of each figure
- Thursday stopped after the two CCGP slides, then jumped to the comparison scores (RSA, CKA). Today opens with a reminder of both scores and the shortest route to the trade-off between them - The path today: the task and the monkeys' behavior, four toy geometries (clustered, general position, square, bent square), the cube demo, then the recordings and what happens on error trials (Bernardi et al. 2020)
- This repeats the last lecture's slide on shattering three stimuli. The rule generalizes: N stimuli in general position in D = N − 1 units can be split every way by a hyperplane - The splits are arbitrary: the score counts the groupings a downstream readout could learn from the population, whatever the population is used for - Bernardi et al. (2020) count only splits into two equal halves, so that guessing is 0.5 for every split - In the assignment the splits are random halves of the 51 objects, and each readout is scored on held-out views - On a line, the split that puts the middle stimulus against the two ends cannot be drawn: the best line gets 2 of the 3 right. Two splits score 2/3, six score 1: (6 + 2 × 2/3) / 8 = 0.92 - "Dimensions used" on the toy figures later today is the participation ratio (PR): 1 when the conditions lie along one direction, 2 when they spread evenly over a plane, up to 3 for four conditions in three units
- Shattering counts how many splits of the same stimuli a readout can separate. CCGP tests whether a readout trained on some objects works on new objects - On the left of the figure, the same points sit in the plane of two units, with three boundaries, one per split. Shattering fits a new readout for every split - The figure's score is the shattering accuracy of the previous slide: the readout accuracy for each split, on held-out trials, averaged over splits - On the right, one readout is fitted on the filled points, then frozen and scored on the hollow points, conditions never used to fit it. That score is CCGP, cross-condition generalization performance - Train an animacy readout on dogs and chairs. Then test it on cats and tables without changing its weights - A high score means animacy is coded along the same direction for all these objects - In the assignment, the split holds out half of the animate and half of the inanimate objects, and the score is averaged over random halvings - The assignment compares CCGP with the majority-class baseline instead of 0.5, because its two classes have different numbers of objects
- Bernardi et al. (2020) trained two monkeys on this task - On each trial the monkey holds a button and fixates. A fractal image appears for 500 ms. The monkey releases the button within 900 ms (R) or keeps holding (H), then gets juice (+) or nothing (−) - Call the images A to D from top to bottom. In context 1, A means release for juice, B hold for juice, C release for nothing, D hold for nothing - In context 2 every image flips its action, and two images also flip their reward. So action and reward are not tied together - A correct response to an unrewarded image avoids a timeout and a repeat of the trial, so the monkey still has a reason to get it right - Context is hidden. The rules flip every 50–70 trials without warning, and nothing on the screen says which context is active - There are eight conditions, 2 contexts × 4 images. Each image fixes the action and the reward, so the conditions are context × action × reward - The neurons are analysed from 800 ms before the next image to 100 ms after it appears. They still hold the previous trial's context, action and reward - So A+ and C− happen only in context 1, and A− and C+ only in context 2 - Each variable (context, action, reward) splits the 8 conditions into two classes, 4 against 4
- Figure: redrawn from Bernardi et al. (2020), Fig. 1c; values extracted from the published figure - The bars count only the first time each image appears after a switch - The first image after a switch gets the old response, so it is almost always wrong. An unexpected outcome is the only sign that the context changed - Images 2 to 4 have not been seen in the new context. A monkey that relearned each image by trial and error would be at chance on them. These monkeys are above chance, so one surprise changes their response to the other images too - Bernardi et al. (2020) call this inference. It needs a variable for context that is separate from any one image
- The gold plane is the decision boundary of our analysis readout, not the monkey's decision. It decodes the context, 1 or 2, from the neurons: a variable the monkey is never shown - CCGP: fit on one condition per context (here A+ and C+), score with weights frozen on the other two (C− and A−). A+ and C+ is one of four such choices - The plane fitted on A+ and C+ classifies C− and A− perfectly, because each lies next to its context partner - The four choices all score 1.00, so CCGP is 1.0 - Splits that separate conditions within a context fail, which pulls shattering down to 0.73 - Dimensions used (PR) 1.01: almost all of the variance lies along one direction, context (99.3%); the small differences within each context add the 0.01. One direction supports one split, which is why only the context split works - How the toy numbers are computed. Each of the four conditions is a mean point in 3 units plus 15 simulated trials: the mean plus Gaussian noise (s.d. 0.10 here). Every readout is a logistic regression - Shattering: for each of the three balanced splits, fit on half of each condition's trials and test on the other half, then swap; repeat over 15 random halves and average. CCGP: fit on all trials of one condition per context, test on all trials of the other two; average over the four choices
- Four conditions at random positions in 3 neurons, with noisy trials. Random placement puts them in general position with probability 1 - Decoding and transfer are different questions. Context is decodable: trained on trials of all four conditions, a plane separates red from blue perfectly, as for every other split (shattering 1.0). CCGP asks something else: does a plane fitted on only two conditions work on the other two? - Here it does not. The drawn plane is fitted on A+ and C+, and C− and A− happen to lie on it, so about half of each one's trials fall on each side: 53% and 47% correct, 50% in all. The readout separates A+ from C+ but does not transfer to C− and A− - The four choices (fitted on → tested on: accuracy) - A+, C+ → C−, A−: 0.50 (the plane drawn) - A+, A− → C−, C+: 0.50 - C−, A− → A+, C+: 0.50 - C−, C+ → A+, A−: 0.50 - Mean: CCGP 0.50, chance - Dimensions used (PR) 2.77, close to the most four points can use (3). With every condition in its own direction, every split is separable, and no direction is shared between them - Same simulation as the clustered toy (four conditions, 15 trials each, logistic-regression readouts), with noise s.d. 0.04
- The gold plane is the decision boundary of the context readout fitted on A+ and C+ only. Any plane between the trials of those two conditions would do, so the fit lands slightly tilted from vertical (about 14°). It still classifies C− and A− correctly; the 3-D view exaggerates the tilt - Shattering here is the mean of three balanced splits: context 1.0, reward 1.0, diagonal 0.46. For images A and C, action is tied to context (release in context 1, hold in context 2), so it is not a separate split here - Between the two extremes: give each variable its own direction. With two variables, the four conditions sit at the corners of a square. This is the XOR layout of the last lecture. Context and reward are two perpendicular directions, so a readout for either transfers. Of the three balanced splits, the diagonal one (A+ and A− against C+ and C−) cannot be separated by a plane - Dimensions used (PR) 2.00, exactly: the four corners lie in a plane. Two directions give the two variables, and nothing is left for the diagonal split - Same simulation as the clustered toy, with noise s.d. 0.02
- Bending adds one dimension to the population, as the lifting unit $x_1x_2$ did for XOR: the embedding dimension goes from 2 to 3. The PR rises only to 2.45, not 3, because the bend is smaller than the sides of the square: the third direction carries less variance than the other two. In neurons, that extra direction comes from nonlinear mixed selectivity: units that respond to a combination of context and reward, here image A vs image C - PR helps shattering but says little about generalization. Generalization needs a coding direction shared across conditions, whatever the number of dimensions: the clustered toy has PR 1.01 and CCGP 1.0, the square 2.00 and 0.78, the bent square 2.45 and 0.77, general position 2.77 and 0.50. The demo shows the same: the cube has PR 3.00 and CCGP 0.95, the compromise 4.14 and 0.80 - PR and shattering go together across the four toys: PR 1.01 → shattering 0.73 (clustered), 2.00 → 0.82 (square), 2.45 → 1.00 (bent), 2.77 → 1.00 (general position). More dimensions, more splits a plane can separate. They are not the same measure: PR weights directions by their variance and ignores the noise, while shattering also depends on how small the trial noise is along the directions that separate each split. Bernardi et al. (2020) note that shattering is correlated with, but not identical to, dimensionality counted from principal components - The context and reward directions stay nearly parallel across conditions, so the readouts still transfer. This is the geometry Bernardi et al. (2020, Fig. 2d) propose for the recordings. Next: add the third variable, action, and the square becomes a cube - Same simulation as the clustered toy, with noise s.d. 0.022
- From the square to the cube: add the third variable, action. The four conditions of the toys become eight, at the corners of a cube, and the mix slider sets how much the cube is bent - Toy model, not fitted to data: eight conditions, each the mean response of 30 model units, plus trial-to-trial noise. Each condition has 20 training and 20 test trials, with Gaussian noise σ = 0.75 by default (the noise slider). Readouts are ridge regressions onto ±1 labels. Shattering averages all 35 balanced splits of the eight conditions; CCGP averages the 16 ways to hold out one condition per side. They are named as in the task, image and outcome; red is context 1, blue context 2 - Why a cube (Bernardi et al. 2020, Fig. 2c, a factorized geometry): three two-valued variables (context, action, reward) make 2 × 2 × 2 = 8 conditions. If each variable has its own direction in the population, the eight conditions sit at the corners of a cube. This is the abstract geometry; the square of the last lecture is the same idea with two variables - Mix = 0, the cube. Take any condition and its partner that differs from it only in context: same action, same reward, other context - In a cube, the step from a condition to its partner is the same step along the same edge, whichever pair you pick. On the map, the gold segments join these pairs: in the cube setting they are horizontal, parallel and of equal length - The map: horizontally, a context readout fitted on six conditions (filled), with its boundary; the two hollow conditions were held out. Vertically, action and reward, plus the direction in which the gold segments differ once the cube bends. The inset (Dimensions used) shows the variance of the eight means along each direction: three directions for the cube, participation ratio 3.00. It is the same PR as under the toy cubes: there it went 1.01 (clustered), 2.00 (square), 2.45 (bent), 2.77 (general position); here 3.00 (cube), 4.14 (compromise) and 7.00 (general position, the most eight points can use) - So one plane, perpendicular to that step, puts every red corner on one side and every blue corner on the other. A context readout fitted on some conditions works on the ones it never saw (CCGP about 0.95) - The cube's limit: groupings that mix the variables, such as context XOR reward, cannot be split by one plane (shattering about 0.72) - Mix = 1, general position: each condition moves to its own random direction. The steps between partners now point in different directions, so the gold segments are no longer parallel, and a plane that separates some pairs need not separate the others. Almost any grouping can be split (shattering about 0.99), but nothing transfers (CCGP about 0.5) - In between, each condition sits part of the way: (1 − mix) times its cube corner plus mix times its random direction. The random directions are drawn twice as far from the centre as the corners, so a small step already counts: at mix = 0.25 both shattering and CCGP are about 0.8. On the map the gold segments tilt a little, each its own way, and the held-out pair still lands on the correct sides; the inset gains four small bars (participation ratio 4.14). Bending adds dimensions without destroying the shared context direction. Bernardi et al. (2020, Fig. 2d) show the same: distort the cube enough and shattering becomes maximal while CCGP stays high, the pattern they measured in hippocampus and prefrontal cortex - All readouts are linear, fitted on training trials and scored on test trials. Decoding uses all eight conditions in training, so it never has to transfer, and it stays above 0.93 in every setting. Only CCGP tells the settings apart - The demo is on the course site with the other demos
- Paper: Bernardi et al. (2020), The geometry of abstraction in the hippocampus and prefrontal cortex, Cell 183:954–967. Open copy: https://pmc.ncbi.nlm.nih.gov/articles/PMC8451959/ - The window is the 900 ms ending 100 ms after the next image appears, before the response: the current trial's action and reward are not known yet, so the variables are the previous trial's. Context carries over across trials - Panel: Bernardi et al. (2020), Fig. 3e - Only the short lines are data. Black lines are CCGP, grey lines are shattering accuracy (the axis label says shattering dimensionality; it is the same average accuracy as on the earlier slides). The coloured and white dots come from the authors' model, not shown here - CCGP for context: the readout is trained on 3 conditions per context and tested on the held-out pair, one from each context. There are 4 × 4 = 16 ways to pick that pair, and the score is averaged over them - Values read from the figure: context CCGP is 0.96 in HPC, 0.74 in DLPFC and 0.81 in ACC. Reward is 0.79, 0.88 and 0.94. Previous action is 0.44, 0.75 and 0.90. Shattering is given in the paper's legend: 0.70, 0.75 and 0.74 - CCGP is compared with a geometric null model, which scores randomly arranged conditions. Its ±2 s.d. band is about 0.41–0.59. Context is above it in all three areas - Previous action in the hippocampus: CCGP 0.44, the orange cluster under the dashed line. It is below 0.5 but inside the null band, so it is not significantly different from chance: the four conditions on each side share no action direction. Action can still be decoded when all conditions are in the training set (Bernardi et al. 2020, Fig. 3a). The paper: action "was not in an abstract format in HPC despite being decodable using traditional methods" - A CCGP well below the null band would mean something else: the readout fitted on three conditions per side would classify the held-out pair the wrong way round, so the action direction would reverse across conditions
- Panels: Bernardi et al. (2020), Fig. 6a–b - Error trials are rare, so this analysis uses its own sample: 180 neurons per area, resampled. Its numbers are not directly comparable with the previous slide's - Average drop in context CCGP on error trials: 0.107 in HPC, 0.069 in DLPFC, 0.065 in ACC, all significant. Average drop in decoding accuracy: 0.069, 0.063 and 0.053, none significant - Decoding asks whether context can be read out. CCGP asks whether it is coded the same way across conditions, which is what transfer to new situations needs. Only the second one goes with the monkeys' errors - Why would the geometry break? The paper does not say, and it does not mention attention. These are not the errors right after a switch: errors within 5 trials of a switch were left out, so these are slips while the rules are stable. Possible reasons, none tested here: a lapse of attention or engagement on that trial, a change in arousal that changes how strongly neurons respond, or a moment of doubt about which context is in force - It is a correlation: the drop in CCGP could cause the error, or both could follow from a third factor such as a lapse of attention
- The four measures describe one set of responses at a time, such as monkey IT, the pixels, or one network layer - Participation ratio is compared with the number of units as PR/D. A variable is explicit when it is available to a downstream neuron, i.e. a linear readout can recover it. CCGP is compared with chance, 0.5, or with the majority-class baseline when the classes differ in size - A representation can score high on one measure and low on another. Flexibility and abstraction pull against each other; the recordings of Bernardi et al. (2020) score high on both - Last Thursday also covered two scores that compare two populations, RSA and CKA. The rest of today asks what any of these numbers has to beat
- The last lecture gave one number per question: readout accuracy, shattering, CCGP, RSA and CKA. Each is one number. Before reading it as "alike" or "decodable", we need to know what a meaningless pairing scores, whether a new sample of images would give the same answer, and how high the noise in the recordings lets a score go - The figure is the pair of RDMs from the last lecture: correlation distances (1 − r) between the 51 object means, in monkey IT (480 neurons) and in ResNet-18's late layer (4,096 units). Their RSA is 0.71 - The rest of today takes the three questions in order. Significance: shuffles give the null distribution. Reliability: cross-validation for a fitted model, the bootstrap for any score. Ceiling: the noise ceiling
- First question of three: what would the score be if there were nothing to find? Shuffles give that baseline for RSA and for readouts
- Simulation: 20,000 pairs of lists of independent random numbers, 10 or 100 items each. For 100 items the correlation has a standard deviation of about 0.1, so 95% of pairs fall within ±0.2; the 95th percentile is 0.16. For 10 items: ±0.6, 95th percentile 0.55 - Fewer items, more luck: a correlation of 0.4 means nothing with 10 movies and is far above chance with 100 - The 95th-percentile rule is a convention: we accept being fooled by luck 5% of the time - The next slide builds the null distribution from the data themselves, by shuffling
- The shuffle measures the luck for the actual data, with their quirks included. The 4,950 entries of an RDM come from only 100 images, so they are far from 4,950 independent pieces of evidence, and the luck is larger than that count suggests - Shuffling keeps everything about each representation (its units, its spread, its dimensionality) and breaks only the pairing of images. The shuffled scores show what those properties produce on their own, with no match between images - The 95th percentile is the value below which 95% of the shuffled scores fall. If the two systems did not match at all, a score above it would come up less than 5% of the time. This is the permutation test of Nili et al. (2014) - The last lecture's monkey IT vs ResNet-18 RSA, 0.712, against a 95th percentile of 0.055 over 1,000 shuffles of the object order - The figure: RSA between six networks and the IT of one person scanned with fMRI (Hebart et al. 2023), on 100 THINGS images averaged over 12 repeats. The trained networks score 0.16 to 0.19, the untrained one 0.06, and the shuffled scores stay near 0.02. The dashed and dotted lines return later in the lecture - You met the same idea with MDS. In the first assignment you refit MDS to a shuffled dissimilarity matrix, which has no structure to recover: 2-D stress 0.446, against 0.204 for the real one
- Training accuracy alone says nothing: with random labels the readout still reaches 1.00 on the training images. Only accuracy on separate test images, never used in the fit, shows what it learned. With true labels: 0.96 on test - As in the shattering reminder at the start of today: N points in general position in N − 1 dimensions can be split in every way. 102 training images in the space of 480 neurons are such points, so a readout can fit any labels, random ones included - This is one form of the curse of dimensionality: with more dimensions (neurons) than training examples, a readout has enough free weights to fit any labelling. The usual remedies are more training images, fewer dimensions (e.g. a few principal components) or regularization, and in every case a score on held-out images - **Generalization gap** = training accuracy − test accuracy. With true labels: 1.00 − 0.96 = 0.04, so what the readout learned holds on new images. With shuffled labels: 1.00 − 0.56 = 0.44: the readout memorized the 102 training images and learned nothing that carries over - Computed from the Bao et al. (2020) recordings: the geometry lecture's animate vs inanimate readout, now on all 480 IT neurons instead of two principal components, trained on 2 views of each of the 51 objects (102 images) and tested on the 306 test views. True labels: training 1.00, test 0.96. Over 1,000 shuffles of the training labels: training 1.00 every time, test 0.56 on average, 95th percentile 0.67 - Guessing at random scores 0.50 on average, and always answering "inanimate" scores 0.63, the majority-class baseline of the geometry lecture. On test images, a readout fit to shuffled labels does no better than guessing - Shuffling the labels keeps the neural responses and the number of animate and inanimate images, and changes only which label goes with which image. The test accuracy after a shuffle is what a readout scores when responses and labels are unrelated - The previous slide did the same for RSA: shuffling the image order in one RDM keeps the structure of each RDM but breaks the pairing between the two - Zhang et al. (2017) trained deep image networks on labels shuffled at random; they still reached 100% training accuracy
- Mean |cos φ| of two random Gaussian vectors: 0.637 in 2-D, 0.263 in 10-D, 0.081 in 100-D, 0.026 in 1,000-D, 0.013 in 4,096-D, close to √(2/(πD)), where D is the number of dimensions - The cosine of the angle φ between two vectors is their dot product divided by both lengths (the MDS lecture): 1 when they point the same way, 0 when they are perpendicular - Here the comparison is within one system: the cosine is between the response vectors of two stimuli, each with one number per unit. Comparisons between two systems come on the next slide - One form of the curse of dimensionality: with many dimensions and few samples, every sample looks about equally far from every other - The same fact makes high-dimensional spaces roomy: many nearly independent directions fit, so a population of many units can carry many variables at once
- Computed: two independent random matrices of N stimuli × 4,096 units: CKA 0.998 with N = 10, 0.976 with 100, 0.804 with 1,000, 0.506 with 4,000 - The previous slide was within one system; this one compares two. The Gram matrix of a system (last lecture) is the N × N table of dot products between its stimuli's response vectors. When the stimuli are nearly perpendicular and of similar length, the off-diagonal entries are near 0 and the diagonal holds each stimulus's squared length. After centering, any two such systems have nearly the same Gram matrix, so their CKA is close to 1 whatever the two systems are - The high score comes from having few stimuli. Each added stimulus adds a row of off-diagonal entries, which differ between unrelated systems and pull CKA down - The conditions: more units than stimuli, response vectors of similar length, and responses spread over many directions - RSA compares only the off-diagonal entries (the distances between different stimuli), so its shuffled baseline stays near 0 at every number of stimuli. With few stimuli the real RSA score is noisy instead - Huh et al. (2024) argued that large networks converge on one representation; Gröger et al. (2026) showed that much of that convergence disappears once every score is compared with its shuffled baseline
- A fitted model, such as a readout or a regression, must be scored on images it was not fit on. Cross-validation does that, and leakage is how it fails - Any score, fitted or not, also depends on which images the study happened to use. The bootstrap estimates how much it would change with another sample - The two answer different questions: cross-validation scores a fit on images it never saw; the bootstrap measures how much any score, fitted or not, varies with the sample of images
- The label-shuffle slide showed that a readout can fit even random labels, so only images held out of the fit tell us what it learned. Cross-validation holds every image out once - The numbers right of the rows: each fit's R² on its own test fold, what a single train/test split would report (0.53 to 0.64). Cross-validation pools all 1,224 held-out predictions: 0.60 - R² is the fraction of the site's variance across images that the predictions explain: 1 is perfect, 0 is no better than predicting the site's mean response. Here it equals the squared correlation r² between prediction and recorded response (r = 0.77). It can differ from r², and even go negative, when predictions are systematically off - Each view of an object is its own image, so test folds hold new views of known objects, not new objects
- When test information reaches the fit, it is called leakage. Z-scoring with the mean of all images, or fitting PCA on all images, lets the test images shape the features the model sees. The effect is often small. Choosing settings, or the most informative units, on the test images can be large: a readout can then look good on pure noise - With cross-validation, this holds inside each fold: the four training folds set the z-scoring, the PCA and the settings, and the fifth fold is only scored - The note at the bottom: the shipped IT file (responses_cleaned) was prepared once, on all 1,224 images: each unit's missing entries filled and the unit centred and scaled. Every split you make afterwards inherits that small leak. With 1,224 images, a mean and a spread estimated from all of them barely differ from those estimated from 80% of them, so the effect on your scores is tiny; the readouts in the assignment z-score again with the training rows only. In a study you publish, do every step inside the training images
- Computed on the human-IT data of the shuffle slide: ResNet-50 first in 55% of resamples, CLIP 18%, DINOv2 16%, ResNet-18 8%, ResNet-50 DINO 4%, the untrained network 0.3% (drawn as 0%) - The criterion is a paired comparison: in each resample, subtract one network's RSA from the other's. If the difference is above 0 in at least 95% of resamples (equivalently, its 95% interval stays above 0, here as a one-sided test), the first network clearly wins. ResNet-50 minus CLIP: higher in 77% of resamples, difference −0.036 to +0.083; minus DINOv2: 75%, −0.038 to +0.071; minus ResNet-18: 87%; minus ResNet-50 DINO: 93%; minus the untrained network: 99.4%, +0.043 to +0.215 - Why paired: both networks are scored on the same resampled images, so a hard image set lowers both. Comparing their two separate intervals would ignore that and be too strict - How to read it: if ResNet-50 were really the best match, nearly every resample (say 95% or more) would put it first. At 55% it is the most likely winner, but the gaps between the trained networks are within what a different sample of 100 images could change. That difference does hold: in 98–99% of resamples, every trained network scores above the untrained one - An image drawn twice appears twice in the RDM; the pairs of an image with its own copy are left out, since their distance is zero in every RDM - For a fitted model, the two can be combined: cross-validate within each resample - Resampling tests replication with new images of the same kind; it does not test other kinds of images - Schütt et al. (2023) resample images and participants together
- Third question: a brain recorded twice does not agree perfectly with itself, so no model can score 1. The noise ceiling says how high a score could go
- Split-half reliability is the split-half correlation of the first lecture, which made the point with one image: single presentations of the same image correlated at 0.17, averages of 25 at 0.69 - Computed from the Papale et al. (2025) THINGS recordings: 100 images × 30 repeats - The figure: the most reliable IT electrode of monkey N on each of the 30 repeats of three THINGS images, beside one unit of ResNet-18's last layer, identical on every repeat. Dashed lines are the means over the 30 repeats - The chipmunk is the electrode's strongest image. Stalagmite (mean 0.18) and pan (mean −0.09) differ on average, but single presentations swap their order 9 times in 30. One presentation cannot rank two images; averaging over repeats can - Even this electrode, the most reliable one, is this noisy; a typical electrode is noisier - No model can be expected to predict a recording better than the recording predicts itself - Averaging more repeats lowers the noise in the mean, which is why the split-half agreement rises with the number of repeats (next slide)
- Spearman–Brown corrects a split-half score to the full number of repeats - One recording has a different ceiling for each measure: compare a CKA score with a CKA ceiling, an RSA score with an RSA ceiling - Computed: for each number of repeats R per half, two disjoint random sets of R repeats are averaged and compared with CKA, over 20 random splits. ResNet-18 vs IT is the CKA between the network's 512 late-layer features and the IT responses averaged over all 30 repeats, on the same 100 images - The ceiling is a property of the recording, not of any model: the model is not on this plot. It is scored against the recording, and that score is read against the ceiling - A caution for CKA: noise in a recording averaged over few repeats spreads the images' response vectors in random directions, which makes them more nearly perpendicular and can raise a model's CKA a little (the same effect as on the few-stimuli slide). ResNet-18's CKA with monkey N's IT is 0.24 with one repeat and 0.19 with all 30. Score models against well-averaged recordings - A score of 0.3 means something different when the ceiling is 0.35 than when it is 0.95: with a low ceiling, the recording is too noisy to tell models apart - In the human fMRI figure earlier, the dotted line was the noise ceiling (0.24), measured with RSA between two halves of the 12 repeats; every network sat between the shuffled scores and the ceiling
- No neuron is matched across animals; the score compares how each recording arranges the images. The grid shows the design only: which recordings are compared with which, and what each comparison is for. On the diagonal, a recording compared with itself scores 1; how far below 1 a comparison between two halves of the same recording falls is its split-half reliability - Safaie et al. (2023): the structure of motor-cortex population activity is preserved across animals doing the same task; Perich (2025) reviews such cross-animal comparisons - A network scored against one monkey's IT should be read next to the other monkey's IT scored against it. Another animal is a reference, not a ceiling: noise lowers both recordings, so correct both scores for noise before comparing - In the human fMRI figure earlier, the dashed line was this reference: another person's IT
- Computed: monkey IT (Bao et al. 2020) against seven networks on the 51 object means, with RSA and linear CKA. RSA moves CLIP from third to first and ResNet-50 from second to fourth. The untrained network is last under every measure - Report more than one measure, say which changes your measure ignores, and treat a conclusion that holds under only one measure as provisional. Soni et al. (2024) find the same disagreement across measures, datasets and model families
- Scores: a floor (shuffled images or labels), held-out images and cross-validation, a ceiling (split-half noise ceiling, another brain), and robustness (resampling, several measures)
- Bernardi, Benna, Rigotti, Munuera, Fusi & Salzman (2020), The geometry of abstraction in the hippocampus and prefrontal cortex - Schütt, Kipnis, Diedrichsen & Kriegeskorte (2023), Statistical inference on representational geometries - Safaie et al. (2023), Preserved neural dynamics across animals performing similar behaviour - Papale, Wang, Self & Roelfsema (2025), An extensive dataset of spiking activity to reveal the syntax of the ventral stream. The monkey recordings behind the noise-ceiling slides - Zhang, Bengio, Hardt, Recht & Vinyals (2017), Understanding deep learning requires rethinking generalization - Before posting, skim the existing entries on your paper; your post must add something not already said
- Submitted on Canvas (Minute papers → Minute Paper 8); credit for a thoughtful attempt