- The last lecture ended with a linear readout fitted to monkey IT responses, and with XOR, a split of four toy stimuli that no linear decision boundary separates - Today's question: what does a population of neurons encode, and in what format? A readout is the tool we use to measure it: which splits of the stimuli a readout can learn from the responses tells us what the population makes available (shattering) - Then CCGP asks whether a readout trained on some objects works on objects it never saw - Then we compare two representations of the same images, such as a network layer and monkey IT
- The figure is from the last lecture: two units, two classes, and the line a fitted readout draws between them - The readout answers "yes" (class 1, the positive side) when w₁x₁ + w₂x₂ + b > 0, and "no" (class 2, the negative side) otherwise. Swapping the names of the two classes flips the signs of w and b, so which side is class 1 is a convention. The bias b is the threshold with its sign flipped. The weight vector w is perpendicular to the boundary and points into the "yes" side - When a probe learns a split from the responses, any neuron downstream could learn it too. How well the probe does on new stimuli, and where its boundary lies, tell us how the category is laid out in the responses
- These are the monkey IT recordings from the last lecture. Inanimate images are blue, as they were on that slide - A split puts every image in one of two classes: animate or not, face or not, and a random half of the 51 objects. All 24 views of an object get the object's color - The points stay where they are from panel to panel. Only the colors change
- Each boundary is a readout fitted on 18 views of every object and scored on the 6 views left out. Chance is the share of the larger class - Two classes are linearly separable when one linear decision boundary (a line in 2-D, a hyperplane with more neurons) puts them on opposite sides - The random labeling splits the objects into two halves at random, so it means nothing to the monkey. Yet a readout of all 480 neurons classifies 78% of new views correctly, far more than a readout of two components does - The readout is linear in every panel. Only the split changes, and with it how well the IT responses separate the two classes
- Flexibility is what the first half of the lecture measured. A population is flexible when a downstream neuron could learn almost any grouping of the stimuli, which is high shattering - Abstraction is the new property. A population codes a variable abstractly when the variable keeps the same format across conditions, so a readout trained on some conditions transfers to others - Spreading points out in many dimensions makes every grouping separable, but then a boundary learned on some conditions says nothing about the others. Giving each variable its own direction makes boundaries transfer, but some groupings, such as XOR, become impossible - Bernardi et al. (2020) asked where hippocampus and prefrontal cortex sit between these two extremes
- Each group of points is one stimulus. The spread within a group is noise - XOR (exclusive or) says "yes" when exactly one input is on. The colors here mark its complement, which has the same geometry
- We found 75% by trying every linear boundary: every orientation, and every position for each orientation - The boundary drawn is one of the best. Every linear boundary leaves at least one group on the wrong side
Each group of points is one stimulus shown many times; the spread within a group is trial-to-trial noise. From here on we work with one point per stimulus, its mean response, so the question becomes purely geometric: which splits of these points can a line or a plane separate
- The four points sit at the corners of a square, at plus or minus 1 on each unit - Before we answer for four stimuli, we count for three - A split puts each stimulus in one of two classes. Three stimuli give 2 × 2 × 2 = 8 splits, including the two where all three share a class. For those two, the boundary lies outside the points - In 2-D, points are in general position when no three lie on one line. In D dimensions, no D + 1 points lie on one hyperplane - In representation B, the two outlined splits put the middle stimulus against both ends. No linear boundary separates them - The readout is linear in both representations, and the stimuli are the same. Only the responses to them change
- Dimension here is the number of independent directions the three points span. Points in general position in a plane span 2; points on a line span 1 - With D dimensions, N points in general position can be split every way as long as N ≤ D + 1. Three points need 2 dimensions; the points on a line have only 1, so the two splits that separate the middle point from both ends fail
- The four points are the clean XOR points from a few slides back, in the plane of two units (x₁, x₂) - Two units never support all 16 splits of four stimuli, however the four points are placed - To separate the last two splits, the representation itself has to change: Tuesday's lifting unit did that. We return to it with real neurons later today
- Each pixel is black or white by a coin flip. Each image is a point in a space with one axis per pixel, 100 axes - The label is a second coin flip. The pixels carry no rule for the readout to find, so a readout that labels the images correctly has memorized them - The readout has 100 weights, one per pixel, and a threshold - A random labeling is the hardest kind of split, and the one where the count can be done exactly. It tells us how many arbitrary splits a representation of a given size supports - Zhang et al. (2017) did this with deep networks and ImageNet photos: with the labels shuffled, the networks still fit every training photo, and their accuracy on new photos fell to chance. We return to this in the lecture on generalization
- So far the units were built by hand, and the stimuli were clean points. Real neurons were not built for one split, and their responses vary from one showing to the next - A population of neurons has to make available whatever split a new task asks for. Shattering, measured with a readout on held-out trials, tells us how many splits it does
- Figure: Photographs reproduced from Retzius (1906); areas marked here - Both studies recorded outside visual cortex, in areas that hold rules, memories and outcomes - Left, the outer (lateral) face of a rhesus monkey brain; right, the inner (medial) face of a hemisphere cut down the middle. The front of the brain is to the right in both - Lateral prefrontal cortex lies around the principal sulcus (area 46). Its neurons hold task rules and items in working memory, and it is needed for flexible, rule-guided behavior. Review: Miller & Cohen (2001) - The hippocampus supports memory and builds maps, of space and of abstract relations between things. Review: Behrens et al. (2018). It sits inside the temporal lobe, so its dot marks a site under the surface - Anterior cingulate cortex (area 24), above the corpus callosum on the inner face, tracks outcomes and errors and adjusts behavior. Review: Heilbronner & Hayden (2016). That matters in a task where the monkey must notice from surprising outcomes that a hidden rule changed - Rigotti et al. (2013) recorded 237 lateral prefrontal neurons in two monkeys. Bernardi et al. (2020) recorded 1,378 neurons: 629 in the hippocampus, 414 in dorsolateral prefrontal cortex (areas 8, 9 and 46), 335 in anterior cingulate cortex - Both studies pool neurons recorded in different sessions as if they had been recorded together. This is called a pseudo-population - The photographs are by Gustaf Retzius (1906), of a real rhesus brain. The dots and labels are ours - Shattering accuracy turns the count into a measure on data - Pick a random split of the stimuli into two groups of equal size. Fit a readout on training trials and score it on held-out trials. Repeat for many splits and average - 0.5 is guessing, because the splits are balanced. 1.0 means every random split is learned - In the assignment, each stimulus is an object and each trial is one view of it - One name, two quantities. Bernardi et al. (2020) call this average accuracy "shattering dimensionality", because it rises with the dimension of the responses. Rigotti et al. (2013), later in this lecture, use the same name for a count of dimensions. In this course: **shattering accuracy** for the average accuracy, **shattering dimensionality** for the count
- Figure: Adapted from Schiller et al. (1976): the original plot, with our axis labels - The plot is the original from Schiller, Finlay and Volman (1976): one monkey V1 neuron, recorded at 36 orientations, 10 repeats each. It fires most near one orientation and little elsewhere - In the first lecture you saw the Hubel and Wiesel film: a cat V1 neuron firing for a bar of light at one orientation - One variable, orientation, describes this neuron well
- Figure: V1: adapted from Schiller et al. (1976). IT: computed from the Bao et al. (2020) recordings - The bars are one of the 480 IT neurons of the assignment (Bao et al., 2020), averaged over the 24 views of each object. We picked the most face-selective neuron: its mean over the 9 face objects is 16.0 spikes/s against 2.4 for the 42 other objects - Its five strongest objects are all human heads - Like the V1 neuron, it is described by one variable: face or not - The next slides ask what a neuron like this would do when the same objects are used in two tasks
- Figure: V1: adapted from Schiller et al. (1976) and Albrecht & Hamilton (1982), original plots with our axis labels. IT: computed from the Bao et al. (2020) recordings - Right: Albrecht and Hamilton (1982), one V1 neuron. The response grows with contrast and saturates, a monotonic tuning curve instead of a peaked one - All three neurons are described by one stimulus variable: orientation, object category, or contrast
- The task is from Warden & Miller (2010). The monkey holds a bar and fixates for 1 s - Object 1 appears for 500 ms, then a 1 s delay, then object 2 for 500 ms and another 1 s delay. The two objects always differ and come from four objects used that day - In recognition blocks, a test sequence of two objects follows. If it shows the same two objects in the same order, releasing the bar during the second test object earns juice. Otherwise the monkey keeps holding - In recall blocks, an array of three objects appears, the two samples and a distractor. The monkey looks at the two sample objects in the order they were shown - No cue said which task was on. Recognition and recall trials were interleaved in blocks of 100–150 trials (Rigotti et al. 2013) - Rigotti et al. (2013) analyzed 237 lateral prefrontal neurons from two monkeys - The conditions are counted during the second delay, after both objects. Shattering asks how many of the ways to split them into two groups a readout can learn
- The plot shows the firing rate (spikes per second) of one imagined neuron for each of the four objects, A to D, shown first. Red is the recognition task, blue the recall task, the colors of the paper - A purely selective neuron ignores the task, so its tuning curves in the two tasks coincide. They are drawn slightly apart so both are visible - These two neurons are hypothetical, drawn to show what each kind of tuning predicts
- In an additive neuron, the task adds the same amount to the response to every object. The recall line is the recognition line moved up - Such a neuron's response is a weighted sum of the two variables, like the x₁ + x₂ unit. Adding more additive neurons does not raise the rank of the population's responses
- Figure: Reproduced from Rigotti et al. (2013), Fig. 3a–d - One recorded neuron, 0.2 s after the first object, before the second: for object C it fires about 6 spikes per second in task 1 and about half that in task 2, while object A drives it more in task 2 - Rigotti et al. (2013, Fig. 3b) show this neuron 0.2 s after the first object appeared, before the second one. Bars are means with their standard errors. The asterisk marks the difference between the two tasks for object C - Compare with the two expectations. Pure: the red and blue bars would match. Additive: every blue bar would be the red bar plus the same amount. Here C drops from task 1 to task 2 while A rises - Such a neuron responds to a combination of object and task, as the x₁x₂ unit responds to a combination of x₁ and x₂
- Figure: Reproduced from Rigotti et al. (2013), Fig. 3a–d - Rigotti et al. (2013, Fig. 3d) show one lateral prefrontal neuron 1.8 s after the first object appeared, while the second object is on the screen - Each bar is one condition, in spikes per second, with its standard error. Red bars are task 1 (recognition), blue bars task 2 (recall). Under the bars, the first object (Cue 1) and the second (Cue 2) - In recall, the neuron fires most when A or D comes second, and most of all after C (the two labelled bars). In recognition, the same pairs drive it less - A V1 or IT neuron of the previous slides is described by one preferred value. This neuron is not - The next slides go back to the XOR stimuli to see what such mixing does for a population
- These are the idealized neurons of Rigotti et al. (2013, Fig. 1), drawn for the XOR stimuli of the last lecture. Earlier today x₁ and x₂ were two units' responses. Here they are the two variables of the task, and each panel is a neuron that responds to them: a neuron called x₁ (pure) responds to the first variable only - A pure unit follows one input. A linear mixed unit follows a weighted sum of both - The score under each panel is the best a single threshold on that unit can do on the four conditions (the corners). XOR asks whether the two inputs have the same sign - A pure unit gets 2 of 4 right, the linear unit 3 of 4
- x₁x₂ is the lifting unit of the last lecture. |x₁ − x₂| is zero when the inputs agree and large when they differ - The AND unit takes a weighted sum of the inputs, subtracts a threshold, and passes the result through a rectifier, max(0, ·), which sets negative values to zero. That is the model neuron of the next lecture - On its own the AND unit gets 3 of 4 right. Its value lies in what it adds to a population, next - Rigotti et al. (2013, Fig. 1a) draw the same kinds of idealized neurons over two stimulus features. Their nonlinear mixed neuron has circular contours, like the radial unit that lifts the two rings (appendix)
- Rank, from the last lecture, is the number of directions the points span after centering each column - The third column is the sum of the first two, so the four points stay on a plane in the space of the three units. The two XOR splits stay out of reach - Any weighted sum of x₁ and x₂ does the same, whatever the weights
- The x₁x₂ column is +1, −1, −1, +1. No weighted sum of the x₁ and x₂ columns produces that pattern, so the rank goes up to 3 - In three dimensions, four points in general position can be split every way, so a readout can learn all 16 splits
- The AND unit is 0, 0, 0, 1 at the four corners. That is not a weighted sum of the x₁ and x₂ columns either, so it raises the rank to 3 - The separating plane is tilted. A readout now weighs x₁, x₂ and the AND unit together to learn the XOR split - A caveat on "any". A nonlinear unit adds a dimension only if its responses to these conditions are not a weighted sum of the others. x₁³ is nonlinear, but at ±1 it equals x₁ and adds nothing. Units whose responses mix the variables, like x₁x₂, the AND unit or |x₁ − x₂|, do add one
- Bottom row: eight stimuli, the four positions of the square in two versions, for example two colors. The units x₁ and x₂ respond to position only, so each red stimulus lands on a blue one, and the best line classifies 50% - x₃ responds to the new property and is linear in it. With x₃ a plane separates the classes: 100%, and the rank of the eight responses is 3 - What helps is not nonlinearity as such. It is a unit whose responses are not a weighted sum of the others': a nonlinear function of the same inputs, like x₁x₂, or a response to something new, like x₃ - The prefrontal neurons later in this lecture make the same point with real recordings
- Figure: Reproduced from Rigotti et al. (2013), Fig. 4b - Each condition becomes one point, the mean response of the neurons in that condition. Shattering asks how many ways one readout can split those points - The log₂ turns the count into a number of dimensions. N spread-out points with N − 1 units can be split all 2^N ways, and log₂ 2^N = N. Rigotti et al. (2013) call this number the shattering dimensionality - The left axis is the number of splits a readout solves (N_c, log scale). The right axis is log₂ of that number - The same name is used for three related numbers: the count of splits a readout solves, its log₂ (Rigotti et al. 2013), and the mean held-out accuracy over a random sample of splits (Bernardi et al. 2020, and the assignment). All three rise together; check which one a paper reports - The window is the second delay, after both objects. There are 24 conditions (4 first objects × 3 second objects × 2 tasks), so 2²⁴, about 17 million splits, and at most 24 dimensions - Populations larger than the recorded one are built by resampling the recorded neurons - The grey line is a simulated population of pure-selectivity neurons, each coding one task variable, with noise matched to the data. It stays below 8 dimensions however many neurons are added - Mixed selectivity is a property of single neurons. Shattering is a property of the population - Criterion (Rigotti et al. 2013, Supplementary Methods M.7): a readout is trained for each split on the condition means of training trials, and tested on held-out trials. A split counts as implementable if the held-out error is below θ = 0.2–0.25, i.e. at least 75–80% correct. The authors report that the exact θ only rescales the noise and does not change the asymptotic count. With 2^24 possible splits they sample 100,000 at random
- Figure: Reproduced from Rigotti et al. (2013), Fig. 5a - Here the analysis uses the recall task only, which has 12 conditions (4 first objects × 3 second objects). The maximum is 12 - 121 recorded neurons had as many correct as error trials in the recall task and enter this analysis. Larger populations, up to 4,000 neurons, are resampled from them - Values read from the figure: on correct trials the curve ends near 11, close to the maximum of 12. On error trials it levels off near 7 - On error trials, a readout can still tell which objects were shown (Rigotti et al. 2013, Fig. 5b). The objects can be decoded, but the responses of the population are lower-dimensional - Rigotti et al. (2013, Fig. 5c–d) removed parts of each neuron's response. Without the linear part of the mixing, the drop remains. Without the nonlinear part, it disappears - What might be going on: Rigotti et al. (2013) do not identify a cause. What they show is that the objects are still encoded on error trials, and that what is missing is the nonlinear mixing, the conjunctions of first object × second object that a downstream circuit needs to plan the right pair of eye movements. Their hypothesis (Supplementary S.1): the ability to read out many combinations "occasionally goes awry, giving rise to error trials" - Candidate explanations, none tested in this paper: a lapse of engagement or attention on that trial, so the circuits that build the conjunctions are weakly driven; or noise in those circuits on that trial. The result is a correlation: the drop could cause the error, or both could reflect a third factor such as a lapse in attention
- Shattering measures flexibility: a downstream neuron can learn almost any grouping of the stimuli it has seen. It says nothing about stimuli it has not seen - An animal meets new objects and new situations all the time. A readout that had to be retrained for every new object would be of little use - So we ask a second question of the same population: if a readout learns "animate or not" from some objects, does it work on objects it never saw? That is CCGP. The two properties can pull against each other, as the next slides show
- Shattering counts how many groupings of the same stimuli can be solved. CCGP tests whether a readout trained on some objects works on new objects - On the left of the figure, the same points sit in the plane of two units, with three boundaries, one per split. Shattering fits a new readout for every split - The figure's score is the shattering accuracy of the earlier slide: the readout accuracy for each split, on held-out trials, averaged over splits - On the right, one readout is fitted on the filled points, then frozen and scored on the hollow points, conditions never used to fit it. That score is CCGP, cross-condition generalization performance - Train an animacy readout on dogs and chairs. Then test it on cats and tables without changing its weights - A high score means animacy is coded along the same direction for all these objects - In the assignment, the split holds out half of the animate and half of the inanimate objects, and the score is averaged over random halvings - The assignment compares CCGP with the majority-class baseline instead of 0.5, because its two classes have different numbers of objects
- The figure shows two toy populations of two units, 25 noisy trials per object. Filled red dots are dogs and filled blue dots are chairs. The gold boundary is a logistic-regression readout fitted on them. Hollow dots are cats and tables, never used to fit - On the left, the dog-to-cat difference (and chair-to-table) is parallel to the boundary: the readout weights give it zero weight. Animate against inanimate is the same step for both pairs, so the frozen boundary classifies 98% of cats and tables correctly - On the right, the same clouds are slid across the boundary. The four groups are just as separable, so shattering can be just as high, but the frozen boundary gets most new objects wrong (22%) - Responses scattered at random in many dimensions solve almost every split. But a boundary fitted on some objects says nothing about new ones, so CCGP is near baseline - One consistent direction gives high CCGP. The conditions then lie in fewer dimensions, so shattering is lower - A population is flexible when many splits can be solved (high shattering). It is abstract when one readout transfers to new conditions (high CCGP) - The readout in the figure is fitted by logistic regression, a readout whose output is a probability
- Bernardi et al. (2020) trained two monkeys on this task - On each trial the monkey holds a button and fixates. A fractal image appears for 500 ms. The monkey releases the button within 900 ms (R) or keeps holding (H), then gets juice (+) or nothing (−) - Call the images A to D from top to bottom. In context 1, A means release for juice, B hold for juice, C release for nothing, D hold for nothing - In context 2 every image flips its action, and two images also flip their reward. So action and reward are not tied together - A correct response to an unrewarded image avoids a timeout and a repeat of the trial, so the monkey still has a reason to get it right - Context is hidden. The rules flip every 50–70 trials without warning, and nothing on the screen says which context is active - There are eight conditions, 2 contexts × 4 images. Each image fixes the action and the reward, so the conditions are context × action × reward - The neurons are analysed from 800 ms before the next image to 100 ms after it appears. They still hold the previous trial's context, action and reward - From the task slide: in context 1, image A means release for juice and image C release for nothing. In context 2, image A means hold for nothing and image C hold for juice - So A+ and C− happen only in context 1, and A− and C+ only in context 2 - The paper's figures use "−" for an unrewarded condition; the dots are coloured by context
- Figure: Reproduced from Bernardi et al. (2020), Fig. 1c - The bars count only the first time each image appears after a switch. The values are read from the figure - The first image after a switch gets the old response, so it is almost always wrong. An unexpected outcome is the only sign that the context changed - Images 2 to 4 have not been seen in the new context. A monkey that relearned each image by trial and error would be at chance on them. These monkeys are above chance, so one surprise changes their response to the other images too - Bernardi et al. (2020) call this inference. It needs a variable for context that is separate from any one image
- This repeats the slide on shattering three stimuli. The rule generalizes: N stimuli in general position in D = N − 1 units can be split every way by a hyperplane - Bernardi et al. (2020) count only splits into two equal halves, so that guessing is 0.5 for every split
- Four conditions at random positions in 3 neurons, with noisy trials. Random placement puts them in general position with probability 1 - The gold plane is our analysis readout, not the monkey's decision. It decodes the context, 1 or 2, from the neurons: a variable the monkey is never shown - CCGP: fit on one condition per context (here A+ and C+), score with weights frozen on the other two (C− and A−). A+ and C+ is one of four such choices - The four choices (fitted on → tested on: accuracy) - A+, C+ → C−, A−: 0.50 (the plane drawn) - A+, A− → C−, C+: 0.50 - C−, A− → A+, C+: 0.50 - C−, C+ → A+, A−: 0.50 - Mean: CCGP 0.50, chance - Shattering, the held-out accuracy averaged over the three balanced splits, is 1.0
- The plane fitted on A+ and C+ classifies C− and A− perfectly, because each lies next to its context partner. Splits that separate conditions within a context fail, which pulls shattering down to 0.73
- This is the XOR layout from the start of the lecture. Context and value are two perpendicular directions, so a readout for either transfers. Of the three balanced splits, the diagonal one (A+ and A− against C+ and C−) cannot be separated by a plane
- Pushing the corners out of the plane adds a dimension, as the nonlinear units did for XOR. The context and value directions stay nearly parallel across conditions, so the readouts still transfer. This is the geometry Bernardi et al. (2020) propose for the recordings
- Figure: Reproduced from Bernardi et al. (2020), Fig. 3e - The figure is Bernardi et al. (2020, Fig. 3e), for the window from 800 ms before the next image to 100 ms after it appears - Only the short lines are data. Black lines are CCGP, grey lines are shattering accuracy (the axis label says shattering dimensionality; it is the same average accuracy as on the earlier slides). The coloured and white dots come from a model, explained on the next slide - CCGP for context: the readout is trained on 3 conditions per context and tested on the held-out pair, one from each context. There are 4 × 4 = 16 ways to pick that pair, and the score is averaged over them - Values read from the figure: context CCGP is 0.96 in HPC, 0.74 in DLPFC and 0.81 in ACC. Value is 0.79, 0.88 and 0.94. Previous action is 0.44, 0.75 and 0.90. Shattering is given in the paper's legend: 0.70, 0.75 and 0.74 - CCGP is compared with a geometric null model, which scores randomly arranged conditions. Its ±2 s.d. band is about 0.41–0.59. Context is above it in all three areas - Previous action in the hippocampus is below chance for CCGP. It can still be decoded when all conditions are in the training set (Bernardi et al. 2020, Fig. 3a)
- Each dot is the mean response of the population across that condition's trials, in the window before the next image - The dots live in the space of hundreds of neurons. The figure shows them in three dimensions to make the geometry visible. It is a schematic, not the data - A perfect cube makes every variable transfer (CCGP high). But XOR-like groupings stay unsolvable (shattering low) - Random points make every grouping solvable (shattering high). But nothing transfers (CCGP near baseline) - Bernardi et al. (2020, Fig. 3e) tested the perfect cube. They put the eight conditions at the corners of a box whose sides give the measured CCGPs, added trial noise, and computed shattering 100 times - The box gives shattering of about 0.60 in HPC, 0.62 in DLPFC and 0.65 in ACC (the white dots, read from the figure). The recordings give 0.70, 0.75 and 0.74, higher than every run of the box - Bernardi et al. (2020) propose that the recordings sit near the cube, with each corner pushed off a little in extra directions by nonlinear mixed selectivity. The context edges stay nearly parallel, so context CCGP stays high. The pushed corners make the other groupings separable too, so shattering is high - On error trials, context CCGP drops in all three areas. Decoding of context, with all conditions in training, does not drop - Two scores do not pin down a geometry: many arrangements give the same shattering and CCGP. The scores rule geometries out (random, clustered), and the cartoons are landmarks, not reconstructions. Bernardi et al. (2020) add three further lines of evidence: the parallelism score (PS), which directly measures how parallel the coding directions of a variable are across conditions; explicit models, such as the perfect box matched to the measured CCGPs (its shattering is too low) and random geometries as null models; and 3-D MDS plots of the eight condition means (their Fig. 3C–D), which show the arrangement changing over the trial as context becomes more abstract
- Toy model, not fitted to data. Eight conditions (context × action × reward) are the mean responses of 30 model units, plus trial-to-trial noise - The readouts are linear, fitted on training trials and scored on test trials - The random part is set twice as long as the cube, so that the mixed setting scores high on both - In the cube setting each variable has its own direction. One boundary per variable transfers to unseen conditions (CCGP about 0.95). But odd groupings cannot be split (shattering about 0.72) - In the random setting every condition sits in its own direction. Almost any grouping can be split (shattering about 0.99), but nothing transfers (CCGP about 0.5) - At the compromise (mixing 0.25) both are about 0.8. This is a pattern like the one Bernardi et al. (2020) measured in hippocampus and prefrontal cortex - Decoding here is the readout accuracy with all conditions in the training set. It is high in all three cases (above 0.93). Only CCGP tells them apart - The demo is on the course site with the other demos
- Figure: Reproduced from Courellis et al. (2024), Fig. 2e–f - Courellis et al. (2024) recorded single neurons in 17 adult patients with drug-resistant epilepsy. They had depth electrodes implanted for seizure monitoring, and microwires in the electrodes record single neurons - 2,694 neurons in 36 sessions. 494 were in the hippocampus, the others in the amygdala, frontal cortex and ventral temporal cortex - The design is the one of the monkey task. The two contexts invert every image–response pairing, and the context switches every 15–32 trials - The contexts also change which images give the large reward, so context, response and reward are not tied together - After one error, a patient who knows the structure can switch all four responses at once. As with the monkeys, the first trial after a switch is almost always wrong - Sessions are split by the first image not yet seen in the new context. In inference-present sessions (22), accuracy on it was about 0.89. In inference-absent sessions (14), it was about 0.47, near chance (values read from the figure) - The figure is Courellis et al. (2024, Fig. 2e–f), hippocampus, 0.2 to 1.2 s after the image appears. The number of neurons is matched between the two kinds of session - Each split of the 8 conditions into two groups of four is one variable, 35 in all. Panel e is decoding accuracy with all conditions in training, panel f is CCGP - Stim pair groups images A and C against B and D. In each context, those images share a response - Values read from the figure: context CCGP rose from about 0.50 to 0.63, stim pair from about 0.56 to 0.70. Shattering, the mean decoding accuracy over the 35 splits, rose from 0.57 to 0.62 (given in the text) - Parity is the XOR-like split of the cube. Its decoding rose too, a sign of nonlinear mixing - No other recorded area showed this change. On error trials in inference-present sessions, the hippocampus looked like it did in inference-absent sessions - The same geometry appeared in patients who learned the structure from verbal instructions instead of by trial and error - In the authors' words, "only the neural representations formed in the hippocampus simultaneously encode several task variables in an abstract, or disentangled, format" - Bernardi et al. (2020) found abstract context in the hippocampus of trained monkeys. Here it appears in people, and only in sessions where they inferred the context
- The four measures describe one set of responses at a time, such as monkey IT, the pixels, or one network layer - A representation can score high on one measure and low on another. Bernardi et al.'s recordings score high on shattering and CCGP together - The rest of the lecture puts two representations side by side - If time runs short, the comparison section below opens the next lecture.
- So far each score read one population. Now two systems see the same images, such as a network layer's units and IT's neurons, or a network and people judging the objects - We cannot match unit 17 to neuron 17, so every score in this section compares how each system arranges the images - A comparison score takes two systems that saw the same images - IT has 480 neurons and a late network layer has thousands of units. No neuron corresponds to a particular unit, so a comparison score must work without such a match - Each score ignores some differences. For example, CKA ignores a rotation of the units. Ask whether your question should ignore it too - The same four questions apply to any comparison score you meet in papers
- The RDM is the table from the MDS lecture - The same 51 Bao objects are seen by two systems, 480 neurons in a monkey's IT and the 4,096 units of ResNet-18's late layer - Neuron 17 does not correspond to unit 17, so we compare what each system says about every pair of objects - The correlation distance 1 − r is 0 when two objects evoke the same pattern across the neurons, and larger as the patterns differ - Each object's response is averaged over its 24 views before the RDM is built. The assignment builds RDMs on individual images. Both are valid. Say which you used - Any change that keeps the order of the dissimilarities leaves RSA unchanged - You computed RSA in the first assignment against human judgments. Here the second system is a monkey's IT - Spearman, as in the MDS lecture: a bent but rising relation between the two RDMs keeps every rank (Spearman 1). Only a swap of order lowers it - We use Spearman because the two systems share no scale - We use the upper triangle of each RDM (1,275 pairs), because an RDM is symmetric with a zero diagonal - Computed on the 51 object means with every unit z-scored, RSA is 0.712 (0.595 without z-scoring) - Shuffling the object order of one RDM gives values near 0. This shuffle check returns in the next lecture - RSA is not a percentage of shared information
- RSA starts from distances between images. The MDS lecture had a second way to compare two response vectors, the dot product. CKA is the comparison score built on it, an alternative to RSA rather than a new idea - The correlation distance 1 − r of the RDM divides out each vector's length and mean - The dot product of two response vectors is large when both are long and point the same way. The cosine similarity divides the lengths out - Keeping the lengths means that an image that drives the neurons strongly counts for more. Ranks would throw that away - RSA compares the order of the distances. CKA compares the dot products - The assignment computes RSA on correlation distance - The assignment first passes raw responses to CKA. Later it also z-scores each unit, which makes CKA ignore each unit's scale - Davari et al. (2023) showed that CKA is sensitive to a few outlying points and high-variance directions. It can change a lot without a change in what a network computes - With few images and many units, CKA between unrelated representations can come out high. The next lecture shows why
- Layer counts: ResNet-18 = 1 stem convolution + 8 blocks × 2 convolutions + 1 final fully connected layer = 18. ResNet-50 = 1 stem convolution + 16 blocks × 3 convolutions (3, 4, 6 and 3 blocks in its four stages) + 1 fully connected layer = 50 - Center each unit first, by subtracting its mean response over the images. Uncentered CKA is dominated by the mean response and reports about 1 for almost anything - The Gram matrix has one row and one column per image, whatever the number of units. Two systems with different numbers of units give tables of the same size - To compare the two tables, CKA treats each one as a long list of numbers and measures their similarity, the cosine of the angle between the two lists - The heat map shows the average-pooled output of every block of an ImageNet-trained ResNet-18 (stem + 8 blocks) and ResNet-50 (stem + 16 blocks), on 1,854 THINGS images - Each ResNet-18 block matches best the ResNet-50 blocks at roughly the same relative depth (gold rings) - Most off-band cells are still above 0.5. Compare cells with each other instead of with zero - The assignment provides `linear_cka` and uses it on IT, on network layers and across networks
- RSA and CKA compare two response matrices. People's choices are behavior. They do not form a response matrix, so the comparison needs a score that works on choices - Cosine similarity (MDS lecture) is the dot product with both lengths divided out. It is 1 when two vectors point the same way - In the figure, ResNet-18's late layer puts cow and horse closest (0.77), so it picks the axe - People agree with each other on about two thirds of triplets. A network's agreement should be compared with that level instead of with 1 - The assignment computes this agreement for six networks on 88,448 test judgments
- Bernardi et al. (2020): read the Introduction and Figures 1–3. Skip the Methods - Rigotti, Barak, Warden, Wang, Daw, Miller & Fusi (2013), The importance of mixed selectivity in complex cognitive tasks - Courellis, Minxha, Cardenas, Kimmel, Reed, Valiante, Salzman, Mamelak, Fusi & Rutishauser (2024), Abstract representations emerge in human hippocampal neurons during inference - Johnston & Fusi (2023), Abstract representations emerge naturally in neural networks trained to perform multiple tasks - Kriegeskorte, Mur & Bandettini (2008), Representational similarity analysis: connecting the branches of systems neuroscience. This paper introduced RSA - Kornblith, Norouzi, Lee & Hinton (2019), Similarity of neural network representations revisited - Hebart, Zheng, Pereira & Baker (2020), Revealing the multidimensional mental representations of natural objects underlying human similarity judgements - Muttenthaler, Dippel, Linhardt, Vandermeulen & Kornblith (2023), Human alignment of neural network representations - Before posting, skim the existing entries on your paper. Your post must add something not already said
- Submitted on Canvas (Minute papers → Minute Paper 7). Credit for a thoughtful attempt
- Each score is the best over every linear boundary: a line in 2-D, a plane for the 3-D roll - The best linear boundary classifies 75% of the XOR points, 68% of the ring points, 90% of the moon points and 75% of the Swiss-roll points - The Swiss roll is the one Isomap unrolled: 800 points on a rolled-up sheet. Here the points on the first half of the sheet are blue and those on the second half red
- XOR: the product x₁x₂ is positive where the inputs share a sign and negative where they differ - Rings: the radial unit grows with the distance from the centre, so the outer ring rises above the inner one - Moons: the new unit has Gaussian tuning, like the tuning curves of the first lecture: exp(−‖x − c‖² / 2σ²), centred on c = (0, 0.25), the LEFT tip of the red moon, where it pokes into the hollow of the blue moon, with width σ = 0.5. We chose c and σ by hand. Red points at that tip respond about 0.8; the nearest blue points, about 0.6 away, respond at most 0.5 (median blue 0.23). The rest of the red moon, far to the right, responds near 0, but there x₁ and x₂ already separate it. The plane uses all three units together, not the new unit alone - Swiss roll: a unit that reports position along the sheet does what Isomap did in the t-SNE and UMAP lecture. If you know the manifold, or have learned it, a coordinate along it can make a class linearly separable that is not separable in the original coordinates. One threshold on that unit splits the roll - We built each unit by hand for its problem. In a later lecture a network's hidden layer learns such units from examples
- Muttenthaler et al. (2023) tested networks trained with labels, with self-supervision and with image–text pairs - Matched pairs of networks that differed only in the objective differed in agreement - A training objective is what a network is trained to do, such as predicting labels, matching two views of an image, or matching images to captions. How training works is the subject of a later lecture - The odd-one-out choice depends on which objects sit closest, inside a category as well as across categories
- The design is the task of Rigotti et al. (2013), with 2 tasks × 4 first objects × 3 second objects. The second object is one of the three objects not shown first, so its identity takes four values - Each population has 200 simulated neurons and no trial noise. The three populations have the same total variance - With noiseless responses, the rank is the number of components with nonzero variance. With recorded neurons every component has some variance, so in practice we count the components needed for, say, 90% of the variance: the embedding dimension of last lecture - A pure neuron responds with one value per level of its variable, for example one value per first object. A linear mixed neuron adds one value for the task, one for the first object and one for the second object - Why 7. The task has 2 values, which after centering give 1 direction. The first object has 4 values, 3 directions. The second object, 3 more. 1 + 3 + 3 = 7 - The participation ratio (last lecture) is 6.2 and 5.9. The seven directions do not share the variance equally
- How the toy data were made. We simulate three separate populations, each of 200 neurons, all responding to the same 24 conditions (2 tasks × 4 first objects × 3 second objects). Each is a response matrix of 24 conditions × 200 neurons, with no trial noise - Pure population: each neuron is assigned at random to one variable (task, first object or second object). It gets one random value (drawn from a normal distribution) for each level of that variable, and responds with that value whatever the other variables are - Linear mixed population: each neuron gets one random value per task, one per first object and one per second object, and responds with their sum - Nonlinear mixed population: each neuron gets an independent random value for each of the 24 conditions, the extreme case. Real neurons sit between this and the linear case - The three populations are scaled to have the same total variance, so their scree plots can be compared - 24 points span at most 23 directions after centering, so 23 is the maximum. The participation ratio is 20.7 - The measurement is the slide "Mixed neurons make a high-dimensional population". Rigotti et al. (2013) simulated pure-selectivity neurons too. Their count levels off far below the maximum, and the recorded neurons do not
- The model is the one in the demo you just tried, at three settings of its slider λ. Each condition's mean is a point in the space of 30 units, and each trial adds noise - The cube puts context, action and reward along three perpendicular directions - The RDM holds the Euclidean distance between every pair of condition means. The MDS map places the eight means in two dimensions to match those distances. The scree plot is PCA on the eight means - The gold segments join pairs that differ only in context. In the 30 units they are exactly parallel on the cube (0 degrees apart). The 2-D map cannot show that. It squeezes a 3-D cube into a plane, with stress 0.24
- Each condition's mean is 0.75 times its cube corner plus 0.25 times a random vector of its own, twice as long as the corner. This is the demo's slider at 0.25 - The four extra components are small (4% of the variance each), so the participation ratio rises only from 3.0 to 4.1 - The gold context segments now differ by 53 degrees on average in the full 30-unit space
- Eight points all at the same distance from each other form a regular simplex. It needs 7 dimensions, so any 2-D map distorts it. The MDS stress, 0.31, is the highest of the three - The participation ratio climbs 3.0, 4.1, 7.0 while context CCGP falls 0.95, 0.80, 0.51. PR counts directions. It does not say whether a readout for context transfers - Shattering and CCGP come from the demo's own computation, with 35 balanced splits and 16 held-out pairs per variable
- Minute paper 5 asked how to check whether t-SNE islands are really in the network's responses. Here the islands are real in all three populations, yet the map cannot tell which population codes context the same way for every condition - t-SNE keeps each trial's nearest neighbours and discards the distances between islands (the t-SNE lecture). The cube's conditions are closer to each other than the random population's, so its islands touch - Neighbours agree with the condition labels for 90% of cube trials in the map (87% in the original 30 units), 87% for the compromise (79%), and 100% for random (100%). Each count uses the 5 nearest neighbours