Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 6 · Theme 1: Representational spaces

Reading and comparing representations

Thursday, October 1 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Picking up: what does a population of neurons encode?

Reminder: a linear readout

Two clouds of points in the space of three units x1, x2, x3, blue upper right and red lower left, separated by a gold plane labelled w dot x plus b equals 0; an arrow labelled w points from the plane toward the blue points

w⋅x+b  {>0:class 1 (positive side)<0:class 2 (negative side)\mathbf{w}\cdot\mathbf{x} + b \;\begin{cases} > 0: & \text{class 1 (positive side)}\\ < 0: & \text{class 2 (negative side)}\end{cases}

  • A linear readout weights the DD responses, sums them (w⋅x\mathbf{w}\cdot\mathbf{x}), and compares the sum with a threshold, −b-b
  • Its decision boundary, w⋅x+b=0\mathbf{w}\cdot\mathbf{x} + b = 0, is a hyperplane of dimension D−1D-1 (a line in 2-D, a plane in 3-D). Which side is class 1 is arbitrary
  • Fitting places that boundary: w\mathbf{w} and bb come from labeled responses (next week)
  • It stands in for a downstream neuron: what it decodes, that neuron could learn. As a test, it is a probe

Where we left off

Three panels side by side, each showing the same monkey IT responses to the 1,224 Bao object images on their first two principal components, one point per image, colored by a different split. Animate or not: animate images red, inanimate blue; the inanimate images fill the upper left, the animate ones form two groups, one at the right and one at the bottom. Face or not: only the bottom group is red. A random half of the objects: red and blue points mixed everywhere
  • Each panel shows the same 1,224 images (51 objects × 24 views), each image a point in the first two principal components of 480 IT neurons' responses
  • Each panel splits the 51 objects into two classes in a different way, red and blue. Which splits can a linear readout learn?

Where we left off: the readouts

The same three panels with a gold readout line in each and its accuracy on new views under the panel. Animate or not: the line separates most red from blue; 93%, chance 63%. Face or not: the line cuts off the bottom group; 100%, chance 82%. A random half of the objects: red and blue stay mixed on both sides of the line; 57%, chance 51%
  • Animate or not is nearly linearly separable: the readout classifies 93% of new views. Face or not is linearly separable on new views: 100%
  • A random labeling (classes by coin flip) is not linearly separable: 57% (chance 51%). All 480 neurons: 99%, 100%, 78%
  • How the neurons respond decides which splits a readout can learn

Two properties of a representation: flexible and abstract

Flexible Abstract
What it means almost any grouping of the stimuli can be read out a readout learned on some stimuli still works on stimuli it never saw
What it supports a fixed association from each input to an output an explicit variable, such as context, that switches the associations
Measured by shattering: how many splits a readout can learn (first half) CCGP: does one readout carry over to new stimuli (second half)

First flexibility, then abstraction. The two pull against each other

XOR: not linearly separable

XOR: four noisy groups of points near the corners of a square in the plane of x1 and x2, blue where the two inputs share a sign and red where they differ
  • XOR, from the last lecture: two toy units respond to four stimuli, each shown many times with noise
  • A point is blue when the two units' responses have the same sign, and red when the signs differ
  • Can a linear decision boundary put all blue points on one side and all red points on the other?

XOR: not linearly separable

XOR: four noisy groups of points near the corners of a square in the plane of x1 and x2, blue where the two inputs share a sign and red where they differ; the best straight boundary, in gold, classifies 75%
  • XOR, from the last lecture: two toy units respond to four stimuli, each shown many times with noise
  • A point is blue when the two units' responses have the same sign, and red when the signs differ
  • No linear boundary can. The best one classifies 75% of the points: three of the four groups
  • XOR is not linearly separable

Can separability measure flexibility?

The XOR groups again, four noisy groups at the corners of a square
  • A representation is flexible if a downstream neuron could learn many different tasks from it. Here, a task = a split of the stimuli into two groups
  • To count them, simplify: each noisy group is one stimulus, shown many times

Can separability measure flexibility?

The same panel with each group replaced by one point at its mean: four points at the corners of a square
  • A representation is flexible if a downstream neuron could learn many different tasks from it. Here, a task = a split of the stimuli into two groups
  • To count them, simplify: each noisy group is one stimulus, shown many times
  • Remove the noise: keep one point per stimulus, its mean
  • Then count the splits a linear readout can separate

Flexibility: three stimuli, two representations

Two blocks of eight small panels, one panel per way to color three stimuli blue or red. Top block, Representation A, general position, 8 of 8: the three stimuli at the corners of a triangle, and every panel has a straight line with the blue stimuli on one side and the red on the other. Bottom block, Representation B, three points on one line, 6 of 8: the same three stimuli on a diagonal line; the two panels where the middle stimulus has the other color from both ends have no line and a thick dashed outline
  • 3 stimuli give 2³ = 8 splits, or labelings (a random labeling picks one at random)
  • Two hypothetical representations of the same 3 stimuli in 2 units, one point per stimulus (its mean response; noise not shown)
  • A, points in general position (no three on one line): all 8 splits are linearly separable, so a readout shatters them. B, points on one line: 6 of 8 (dashed: no line separates the colors)

Higher-dimensional, more flexible

Two blocks of eight small panels, one panel per way to color three stimuli blue or red. Top block, Representation A, general position, 8 of 8: the three stimuli at the corners of a triangle, and every panel has a straight line with the blue stimuli on one side and the red on the other. Bottom block, Representation B, three points on one line, 6 of 8: the same three stimuli on a diagonal line; the two panels where the middle stimulus has the other color from both ends have no line and a thick dashed outline
  • A: three points in general position span 2 dimensions. B: three points on a line span 1
  • Shattering characterizes the representation: A's extra dimension makes every split learnable (8 of 8 against 6 of 8), so A is more flexible

Shattering four stimuli: 14 of 16 splits

Sixteen small panels in a four by four grid, one for each way to color four stimuli at the corners of a square blue or red, under the header Four stimuli, 14 of 16. Fourteen panels have a straight line separating blue from red. The two panels in which opposite corners share a color have no line and a thick dashed outline
  • Four stimuli in 2 units, at the corners of a square: 2⁴ = 16 splits
  • A linear boundary separates 14 of them. The two it cannot separate are the XOR splits, which put opposite corners together
  • Even in general position (no three on one line), four points in a plane give 14 of 16, never more
  • Two units cap flexibility: shattering four stimuli needs a third dimension, a unit that is not a weighted sum of the first two

Random labels: splits that mean nothing

Eight random 10 by 10 pixel images in a row, each pixel black or white. Each image has a colored frame and a label under it: class 2 (red) or class 1 (blue); five are class 1 and three class 2
  • Each stimulus: a random 10 × 10 pixel image (100 units), its class set by a coin flip. Getting the training images right means memorizing them; on new images a readout is at chance
  • Capacity: how many random labels a classifier can memorize. A linear readout: about two per dimension (Cover 1965)
  • Networks: more weights, more capacity. Zhang et al. (2017): deep networks memorized ImageNet training photos with shuffled labels, and were at chance on new photos
  • Shattering measures capacity. Generalizing needs more than capacity: the trade-off between the two is a future lecture

How flexible are neurons in the brain?

Recorded neurons are noisy and many. How many splits can a readout learn from their responses?

Shattering in real brains: two monkey studies

Two photographs of a rhesus monkey brain (Retzius 1906): the outer side and the inner side, front to the right. Three labelled dots mark approximate recording sites: lateral prefrontal cortex (area 46, Rigotti 2013 and Bernardi 2020), anterior cingulate (area 24, Bernardi 2020) and hippocampus (Bernardi 2020)
  • Prefrontal cortex: rules, working memory. Hippocampus: memory, maps. Anterior cingulate: outcomes, errors
  • Rigotti et al. (2013): lateral prefrontal cortex. Bernardi et al. (2020): all three

Simple tuning: one neuron, one variable

Left panel only, titled V1: one orientation. Original printout of one monkey V1 neuron: total spikes over 10 trials against stimulus orientation, 0 to 348 degrees, with one sharp peak near 80 degrees and a low floor elsewhere
  • V1 (first lecture): fires most for a bar at one orientation

Simple tuning: one neuron, one variable

Two panels. Left, V1: one orientation. Original printout of one monkey V1 neuron: total spikes over 10 trials against stimulus orientation, 0 to 348 degrees, with one sharp peak near 80 degrees and a low floor elsewhere. Right, IT: one kind of object. Bar chart of one IT neuron's mean firing rate to each of 51 objects, grouped by category; the nine face bars reach 14 to 18 spikes per second, and nearly every other bar is below 5
  • V1 (first lecture): fires most for a bar at one orientation
  • IT (last lecture): fires for faces, little for anything else

Simple tuning: one neuron, one variable

Three panels. Left, V1: one orientation. Original printout of one monkey V1 neuron: total spikes over 10 trials against stimulus orientation, 0 to 348 degrees, with one sharp peak near 80 degrees and a low floor elsewhere. Middle, IT: one kind of object. Bar chart of one IT neuron's mean firing rate to each of 51 objects, grouped by category; the nine face bars reach 14 to 18 spikes per second, and nearly every other bar is below 5. Right, V1: more contrast. Original plot of one V1 neuron's response, as percent of its maximum, against contrast from 0 to about 55 percent: it rises steeply up to about 15 percent contrast, then levels off near 100 percent
  • V1 (first lecture): fires most for a bar at one orientation
  • IT (last lecture): fires for faces, little for anything else
  • V1, contrast: the response rises, then levels off (monotonic tuning)

The task: remembering a sequence in prefrontal cortex

Timeline of the sequence-memory task in three left-aligned rows. Sample: remember which two objects, in which order: fixate 1000 ms, object A 500 ms, delay 1000 ms, object C 500 ms, delay 1000 ms. Recognition block, in red: a test sequence A, delay, C; same two objects in the same order, release the bar; different, keep holding. Recall block, in blue: three objects B, A and C around the fixation point; arrow 1 to A, then arrow 2 to C: look at the two objects in the order seen

  • Recognition: keep the two objects in memory and compare them with a test sequence. Recall: plan two eye movements, in order, to the objects in an array
  • One task needs a memory to compare, the other a plan of movements. You might expect two separate groups of neurons
  • A condition is one combination of the task variables: 4 first × 3 second objects × 2 tasks = 24
  • Yet the same prefrontal neurons take part in both. How can one population encode objects and task so that the monkey can switch?

What would one neuron do in this task?

Two panels. Left, Pure, the object only: firing rate of hypothetical neuron 1 for first objects A to D, one tuning curve per task, recognition red and recall blue; the two curves coincide, peaking near 6 spikes per second at C. Right, two pure neurons: the 8 conditions plotted as neuron 1 against neuron 2; each red point sits on its blue partner, so no decision boundary separates the two tasks

Hypothetical neurons, synthetic data

  • Simple tuning, as in V1 or IT, predicts one tuning curve over the first object
  • Pure selectivity: the neuron's response depends on the first object only, so its tuning curves in the two tasks coincide

What would one neuron do in this task?

The same two panels with neuron 1 now additive. Left: the blue recall curve is the red recognition curve moved up by 1.2 spikes per second at every object. Right: every blue point moves right by 1.2, and a gold decision boundary from upper left to lower right separates red from blue

Hypothetical neurons, synthetic data

  • Simple tuning, as in V1 or IT, predicts one tuning curve over the first object
  • Pure selectivity: the neuron's response depends on the first object only, so its tuning curves in the two tasks coincide
  • Additive mixed selectivity: the task adds the same amount to the response to every object, so the two tuning curves are parallel
  • With such a neuron, a linear readout can tell the two tasks apart

The recorded neurons: mixed

Rigotti et al. (2013) Fig. 3, four panels from two recorded prefrontal neurons. Top: neuron 1, firing rate over time and bar charts per first object in the recognition task (red) and the recall task (blue); it fires most to object C in recognition, about half as much in recall. Bottom: neuron 2 in all 24 conditions, preferring objects A and D as the second object in the recall task
  • Recorded prefrontal neurons are neither pure nor additive: how a neuron responds to an object depends on the task
  • Nonlinear mixed selectivity: the neuron responds to combinations of object and task, like the lifting unit x1x2x_1x_2 of the last lecture
  • Many prefrontal neurons do. One preferred stimulus cannot describe them

The recorded neurons: mixed

The same Rigotti et al. (2013) Fig. 3, with the lower-right panel outlined in gold: neuron 2 in all 24 conditions, bars per combination of first object, second object and task; in the recall task it fires most when A or D comes second, most of all after C
  • Recorded prefrontal neurons are neither pure nor additive: how a neuron responds to an object depends on the task
  • Nonlinear mixed selectivity: the neuron responds to combinations of object and task, like the lifting unit x1x2x_1x_2 of the last lecture
  • Many prefrontal neurons do. One preferred stimulus cannot describe them
  • Outlined: a second neuron, all 24 conditions. Its response to object 2 depends on object 1 and the task
  • Is this mixing noise, or part of the code?

One unit alone: pure and linear units

Three panels over the plane of x1 and x2, each coloring one unit's response from low (light) to high (dark), with the XOR stimuli as blue (same sign) and red (different signs) dots near the four corners. x1, pure: the response changes from left to right only, gold threshold vertical. x2, pure: from bottom to top only, gold threshold horizontal. x1 plus x2, linear: along the diagonal, gold threshold a diagonal line. Under the panels, how many of the four conditions each unit gets right alone: 50%, 50% and 75%. A colour bar at the right
  • Now x1x_1 and x2x_2 are two task variables (like object and task), not neurons. Each panel is one neuron's response to them; gold, its best threshold
  • No pure or linear unit solves XOR alone

One unit alone: nonlinear units

The same layout with three nonlinear units, left to right. x1 times x2: high in the two corners where the inputs share a sign, low in the other two; its gold threshold follows both axes and gets 100%. The absolute difference of x1 and x2: zero along the diagonal, high toward the other two corners; 100%. The AND unit, max of 0 and x1 plus x2 minus 1: zero everywhere except toward the upper right corner; 75%
  • Nonlinear mixed units: x1x2x_1x_2 (the lifting unit of the last lecture) and ∣x1−x2∣|x_1-x_2| split XOR alone; AND (a rectified sum) fires when both are high
  • AND alone: 3 of 4. Added to x1x_1 and x2x_2, all 16 splits become separable (next slides)

The response matrix: linear mixing

Left, a response matrix with four rows, the four XOR conditions (minus minus, minus plus, plus minus, plus plus), and three columns, the units x1, x2 and x1 plus x2, each cell colored and printed with its value. Under it: rank 2, 14 of 16 splits. Right, the four conditions as points in three dimensions with axes x1, x2 and x1 plus x2, labelled by their corner; all four lie on one tilted plane, drawn lightly
  • The x1+x2x_1+x_2 column is the sum of the other two columns, so the rank stays 2. A linear readout can separate 14 of the 16 splits
  • Rank, as last lecture: the number of principal components with nonzero variance. These responses are noiseless, so we can read it off exactly
  • A unit whose responses are a weighted sum of the other units' responses never raises the rank

The response matrix: add the lifting unit

The response matrix with a fourth column, x1 times x2: plus 1, minus 1, minus 1, plus 1. Under it: rank 3, 16 of 16 splits. Right, the points in three dimensions with axes x1, x2 and x1 times x2: the two same-sign corners are raised, the two others lowered, and a plane at x1 times x2 equal to zero separates them
  • Add x1x2x_1x_2, the lifting unit of the last lecture: no weighted sum of the other columns gives it. The rank rises to 3, and all 16 splits are separable
  • Its responses are not a weighted sum of the others', so the points leave the plane

The response matrix: add the AND unit

The response matrix with the AND unit as fourth column instead of x1 times x2: 0, 0, 0, 1. Under it: rank 3, 16 of 16 splits. Right, the points in three dimensions with axes x1, x2 and the AND unit: only the plus plus corner is raised, and a tilted plane separates the two same-sign corners from the two others
  • Put the AND unit in place of x1x2x_1x_2. Its responses are not a weighted sum of the others' either: rank 3 again, and all 16 splits are separable
  • AND alone cannot do XOR. A readout of all three can: x1+x2−4 ANDx_1 + x_2 - 4\,\mathrm{AND} is −2-2 for same signs and 00 for different signs, so one threshold splits them
  • Many nonlinear units raise the rank. The product unit from the last lecture is only one of them

Summary: new units must not be weighted sums

Three panels. The two from before, plus a third: eight stimuli whose red and blue versions share positions in the x1, x2 plane, separated by a new unit x3 that responds to a different property
  • Top: adding x1x2x_1x_2 raises the rank to 3 (16 of 16); adding x1+x2x_1+x_2, a weighted sum, leaves it at 2 (14 of 16)
  • A unit can also help by responding to something new. Bottom: a linear unit x3x_3, tuned to a new property, separates stimuli that share positions
  • A new unit helps only if its responses are not a weighted sum of the others': then the rank goes up

Mixed neurons make a high-dimensional population

Rigotti et al. (2013) Fig. 4b, second delay of the sequence task. Line plot against the number of neurons read out, 0 to 4,000. Left axis: number of two-class splits a readout solves, N sub c, on a log scale from 10 to 10 to the 7; right axis: number of dimensions, 6 to 24. Black line, recorded data: rises steadily to about 23 dimensions at 4,000 neurons, just under the dashed maximum of 24. Grey line, simulated pure-selectivity neurons: levels off near 7 dimensions from about 500 neurons on. Error bars on each point
  • Shattering dimensionality counts flexibility: log⁡2\log_2 of the number of splits a readout solves (solved: ≥ 75–80% on held-out trials)
  • 24 conditions, 2²⁴ splits: solving all of them gives log₂ 2²⁴ = 24, the maximum (dashed line)
  • Recorded prefrontal neurons (black): close to 24, almost any grouping is readable
  • Pure selectivity, simulated (grey): at most 8 (3 + 3 + 1 directions for the two objects and the task, plus 1)
  • Nonlinear mixed selectivity makes the responses of the population high-dimensional

High dimensionality goes with correct choices

Rigotti et al. (2013) Fig. 5a, recall task. Line plot against the number of neurons read out, 0 to 4,000. Left axis: number of two-class splits a readout solves, N sub c, on a log scale; right axis: number of dimensions, 2 to 12. Black line, correct trials: rises to about 11 dimensions, close to the dashed maximum of 12. Grey line, error trials: levels off near 7. Error bars on each point
  • Recall task only: 12 conditions, so at most 12 dimensions (dashed line)
  • Correct trials (black): close to 12
  • Error trials (grey): about 7, i.e. about 2⁷ of the 2¹² splits are readable (2¹¹ on correct trials). Which objects were shown can still be decoded (Fig. 5b)
  • The drop is in the nonlinear mixed part of the responses: remove it and correct and error trials look alike
  • Less flexible on error trials: the objects are encoded, but fewer groupings are readable

Second property: abstraction

Shattering measured capacity: how many groupings can be learned, even random ones. Abstraction asks about generalization: does one readout carry over to stimuli it never saw? The two trade off

One boundary for stimuli never seen: CCGP

A readout for animate against inanimate is fitted on dogs and chairs, then applied to cats and tables. Left, shattering refits a readout for every split of the same objects; right, CCGP fits once and keeps the weights fixed for objects it never saw

  • CCGP (cross-condition generalization performance): fit an "animate or not" readout on images of dogs and chairs, then test it, weights fixed, on images of cats and tables
  • A high CCGP means the code for animacy is abstract: not tied to the training examples (dogs, chairs), so it carries over to new ones (cats, tables) (minute paper)

CCGP: when does one boundary transfer?

Two toy panels in the plane of two units x1 and x2, with noisy clouds. Filled red dots are dogs and filled blue dots chairs; a gold boundary is fitted on them. Hollow red dots are cats and hollow blue dots tables, never used to fit. Left, cats overlap the dogs and tables overlap the chairs; the dog-to-cat shift, shown by a grey arrow, runs along the boundary; on cats and tables the boundary scores 98%. Right, the same dogs, chairs and boundary, but cats and tables are shifted across the boundary; 22%
  • Filled: dogs, chairs (fit on). Hollow: cats, tables (never seen). Left: cats differ from dogs along the boundary, which the readout ignores: 98%. Right: across it: 22%
  • Abstract: objects differ only in directions the readout ignores (minute paper)

The task: a hidden context

Schematic of one trial: inter-trial interval 1750 ms, hold the button and fixate 400 ms, one of four images A to D for 500 ms, hold or release within 900 ms, a 500 ms wait, then the outcome, juice or nothing. Below, two tables, each row tagged with its condition. Context 1 (red): A release, juice, A plus; B hold, juice, B plus; C release, none, C minus; D hold, none, D minus. Context 2 (blue): A hold, none, A minus; B release, juice, B plus; C hold, juice, C plus; D release, none, D minus. Between them: switch every 50 to 70 trials, no cue, inferred from outcomes

Schematic of the task of Bernardi et al. (2020)

  • An uncued context flips the rules. Each variable splits the 8 conditions into two classes
  • Memorized associations are not enough: the monkey must hold context as a variable

The monkeys infer the hidden context

Bernardi et al. (2020) Fig. 1c. Bar chart of average performance in percent, 0 to 100, against image number around a context switch, marked by a magenta line. Last image before the switch: about 95%. First image after it: about 9%. Second, third and fourth images: about 56, 74 and 87%. A dashed line marks 50%, chance for a choice between release and hold. Error bars are 95% confidence intervals
  • Accuracy on the last image before a switch, then on the first showing of each image after it
  • 1st image after the switch: about 9% correct. Nothing warned the monkey
  • 2nd to 4th images, not yet seen in the new context: above chance (dashed line, 50%)
  • Bernardi et al. call this inference: it suggests a variable for context, coded separately from any one image
  • After one surprise, the monkeys switch all the rules

Recall: points in general position split every way

The shattering figure from earlier in the lecture: two blocks of eight small panels, one per way to color three stimuli blue or red. Top, three stimuli in general position: every panel has a line separating blue from red, 8 of 8. Bottom, three stimuli on one line: 6 of 8
  • Earlier: 3 stimuli in 2 units, in general position: a line separates all 8 splits
  • In general: NN stimuli in general position in D=N−1D = N - 1 units, every split is separable
  • One more unit, one more stimulus: 4 conditions at random in 3 neurons can be split all 16 ways by a plane
  • Bernardi's shattering accuracy: a readout for every balanced split (3 with 4 conditions, 35 with 8), scored on held-out trials, averaged. 0.5 = guessing
  • Next, four toy geometries and the scores each would produce. Then the scores measured in monkeys, and which geometry they match

Four toy geometries, 1 of 4: points at random

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Flexible, not abstract
  • Red/blue: context 1/2. Context readout (gold plane): fitted on A+, C+ (filled), tested on C−, A− (hollow)
  • N=4N = 4 conditions placed at random in D=3D = 3 neurons, so in general position. D=N−1D = N - 1, as for the three stimuli in two units: every split is separable
  • Shattering 1.0. Generalization (CCGP) 0.50: the plane cuts right through C− and A−

Four toy geometries, 2 of 4: clustered by context

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Abstract, not flexible
  • Red/blue: context 1/2. Context readout (gold plane): fitted on A+, C+ (filled), tested on C−, A− (hollow)
  • Each context's two conditions sit close together
  • Generalization (CCGP) 1.0: context transfers. Shattering 0.73: conditions within a context cannot be told apart

Four toy geometries, 3 of 4: a square

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context, CCGP value

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Abstract, but not fully flexible
  • Red/blue: context 1/2. Context readout (gold plane): fitted on A+, C+ (filled), tested on C−, A− (hollow)
  • Corners of a square: one side is context ({A+, C−} vs {A−, C+}), the other value, i.e. reward ({A+, C+} vs {A−, C−}). A readout for one ignores the other
  • Generalization: CCGP 0.78 (context), 0.72 (value). Shattering 0.82: the diagonal XOR split fails

Four toy geometries, 4 of 4: a square, bent

Schematic in the space of three units x1, x2, x3: four conditions A+, C− (red, context 1) and A−, C+ (blue, context 2), each a mean with its noisy trials. A+ and C+ are filled (fitted on), C− and A− hollow (never seen). A gold plane, the context readout fitted on A+ and C+. Bars at right: shattering, CCGP context, CCGP value

Simulated: a hypothetical population of 3 neurons (the axes), one dot per trial, after Bernardi et al. (2020), Fig. 2

  • Flexible and abstract at once
  • Red/blue: context 1/2. Context readout (gold plane): fitted on A+, C+ (filled), tested on C−, A− (hollow)
  • The same square, corners pushed out of the plane
  • Shattering 1.0: the diagonal split separates too. Generalization still holds: CCGP 0.77, 0.73

Hippocampus: context transfers, action does not

Bernardi et al. (2020) Fig. 3e. For hippocampus (HPC), dorsolateral prefrontal cortex (DLPFC) and anterior cingulate cortex (ACC), a vertical axis from 0.4 to 1 labelled shattering dimensionality and CCGP, with a dashed line at 0.5. Short black lines mark the measured CCGP of three variables, each on a set of coloured dots from 100 runs of a perfect-cube model: context in red (HPC 0.96, DLPFC 0.74, ACC 0.81), value of the previous trial in purple (0.79, 0.88, 0.94) and action of the previous trial in orange (0.44, 0.75, 0.90). Short grey lines mark the measured shattering dimensionality (0.70, 0.75, 0.74). Clusters of white dots, the cube model's shattering dimensionality, sit lower, at about 0.60, 0.62 and 0.65
  • Black lines: CCGP. Red: context. Orange: previous action. Purple: value (reward or not)
  • Grey lines: shattering accuracy. The dots come from a model (next slide)
  • Shattering (grey): 0.70 HPC, 0.75 DLPFC, 0.74 ACC, far above 0.5. Context CCGP (red): 0.96, 0.74, 0.81, above the null band (0.41–0.59). Previous action in HPC: 0.44
  • Context transfers in all three areas, with shattering high too: like toy geometry 4 (bent square), not 1 (random) or 2 (clustered)

Why both can be high: the compromise geometry

Schematic of the compromise geometry: eight conditions as dots in the space of three units, colored by context, near the corners of a cube but each pushed off a little. Gold edges join conditions that differ only in context and stay nearly parallel; grey edges join the others. Text above: near the cube's corners, each pushed off a little, most two-class splits separable (shattering high) and the direction still nearly parallel (CCGP high)
  • Dots: conditions. Color: context (schematic)
  • Cube (Rigotti's linear mixing): each variable adds its own fixed step, whatever the others are. Gold edges: the context step
  • Same step everywhere (parallel edges): one boundary per variable works for unseen combinations, high CCGP
  • Corners pushed off the cube (nonlinear mixing) make odd groupings solvable (high shattering)
  • A perfect cube with the measured CCGPs has lower shattering (white dots, last slide). The recordings sit near the cube, not on it

Try it: flexible or abstract?

  • Toward random: shattering rises, CCGP falls. Decoding (all conditions in training) stays high

Human hippocampus: context becomes abstract

Courellis et al. (2024) Fig. 2e–f, hippocampus while the image is on, inference-absent against inference-present sessions. e, decoding accuracy from 0.4 to 0.8 for all 35 ways to split the 8 conditions in half, grey open circles, with named splits in colour: context (red-orange) rises from about 0.57 to 0.72, stim pair (purple) from 0.62 to 0.77, parity (yellow) from 0.54 to 0.64, outcome (blue) falls from 0.62 to 0.56, response (green) rises from 0.46 to 0.50. Black lines mark shattering dimensionality, 0.57 and 0.62 in the text. f, CCGP from 0.3 to 0.7: context rises from about 0.50 to 0.63, stim pair from 0.56 to 0.70, parity from 0.46 to 0.53, outcome falls from 0.56 to 0.41, response stays near 0.36. Grey bars show the 5th to 95th percentile of a null distribution, about 0.40 to 0.60. The legend is at right

  • Patients with implanted electrodes, the monkeys' task. Each dot: one of 35 half-splits of the 8 conditions. Grey: chance
  • When people infer the switch, context (red) and stim pair (purple) become abstract

Recap: measures of one population

Measure What it says about the representation Compared with
Participation ratio How many dimensions it uses the number of units DD, as PR/D\mathrm{PR}/D
Readout accuracy Whether a variable is explicit: available to a downstream neuron majority-class baseline
Shattering accuracy How flexible it is: how many different groupings it makes available guessing, 0.5
CCGP How abstract it is: whether a variable keeps one format across conditions chance, 0.5

Each score describes one population of neurons or units. How alike are two populations?

By the way: comparing two representations

A tool for Assignment 2, not part of today's story: besides RSA, a second way to compare two systems, CKA

Recall from the MDS lecture: RSA

Two 51 by 51 dissimilarity matrices for the same Bao objects in the same category order, with object thumbnails along the edges, from monkey IT (480 neurons, left) and ResNet-18's late layer (4,096 units, right). Both show a light block for faces and lighter blocks within animals and within vehicles

Computed: correlation-distance RDMs of the 51 object means, monkey IT (480 neurons) and ResNet-18 late layer (4,096 units)

  • One RDM per system: correlation distance 1−r1 - r (rr: Pearson) for each pair of objects. RSA: Spearman ρ\rho of the two RDMs, 0.71 here. Only the rank order of distances counts

Another popular score: CKA

  • Each builds one matrix per system, then compares the two matrices. They differ at both steps:
RSA CKA (centered kernel alignment)
1. Build the matrix RDM: a distance per pair, here 1−r1 - r Gram matrix: a dot product per pair (units centered)
2. Compare the two matrices rank order (Spearman ρ\rho) values (normalized dot product)
Unchanged by any change that keeps the order of the distances uniform rescaling, rotation of the units
  • Assignment 2 uses both. By convention, CKA is more common for comparing network layers (Kornblith et al. 2019), RSA for brains and judgments

CKA on two networks, block by block

Heatmap of linear CKA between every block of an ImageNet-trained ResNet-18 (x axis, stem and blocks 1 to 8) and ResNet-50 (y axis, stem and blocks 1 to 16), measured on 1,854 THINGS images. A dark band runs from bottom-left to top-right: each ResNet-18 block is most similar to the ResNet-50 block at the same relative depth, marked by gold rings, and early-versus-late pairs are least similar

Computed: ResNet-18 vs ResNet-50, 1,854 THINGS images

  • ResNet-18 against ResNet-50, every block against every block, 1,854 THINGS images. A block is a group of layers with a skip connection: 8 blocks of 2 layers in ResNet-18, 16 blocks of 3 in ResNet-50
  • The highest CKA is between blocks at the same relative depth

Do networks judge objects the way people do?

Three THINGS object photographs in a row, a triplet; under them the rule: a representation's odd one out is the object that is not in the most similar pair, by the cosine of the angle between response vectors; a mark shows which object people picked

Public-domain THINGS photographs (CC0). Odd-one-out rule applied to ResNet-18's late layer

  • Odd-one-out: shown three objects, a person picks the one that belongs least. The THINGS odd-one-out data hold 4.7 million such choices over 1,854 objects (Hebart et al. 2023)
  • A network picks the object left out of its most similar pair (cosine similarity). Agreement is the fraction of triplets where network and person pick the same. Chance is 1/3

Papers from today

Minute paper

Submit on Canvas → Minute papers → Minute Paper 7 (access code read out in class).

Write three brief points in your own words:

  1. A readout trained to tell dogs from chairs is tested, weights fixed, on cats and tables.
    What does a high score tell you about the population, and what is that measure called?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Appendix

Extra slides: not covered in class, useful for the assignment and the research trail

Not linearly separable: four examples

Four panels in a row. XOR: four noisy groups at the corners, a dashed best line, best line 75%. Two rings: a blue inner ring inside a red outer ring, best line 68%. Two moons: two interleaved crescents, blue above and red below, best line 90%. Swiss roll: the rolled sheet in 3-D, the inner half of the sheet blue and the outer half red, best plane 75%
  • Four toy data sets: XOR; two rings; two moons; and the Swiss roll from the lecture on maps, split into the inner and the outer half of the sheet
  • In each, no linear decision boundary puts the blue points and the red points on opposite sides

A third unit lifts each one

The same four examples, each with one new unit. XOR plus x1 times x2: in 3-D the blue groups rise above a gold plane and the red ones drop below it; one plane 100%. Two rings plus x1 squared plus x2 squared: the outer ring rises above the inner ring, and a flat gold plane separates them; 100%. Two moons plus a Gaussian-tuned unit: the points near the unit's preferred point rise, and a tilted gold plane separates the moons; 100%. Swiss roll plus position along the roll: the sheet unrolled flat, position along the roll left to right and height bottom to top, with a vertical gold line splitting it; one threshold 100%
  • Each panel adds one unit that is not a weighted sum of the inputs (named above each panel)
  • Position along the roll is the coordinate Isomap recovers when it unrolls the manifold (the t-SNE and UMAP lecture)
  • With the right extra unit, a plane separates the two classes in all four data sets

What sets a network's agreement with people?

  • Muttenthaler et al. (2023) scored many networks on the THINGS odd-one-out choices
  • The training objective and training data mattered more than architecture or size
  • Classification accuracy only says on which side of a category boundary each image falls. It does not say how objects are arranged within a category
  • Two networks with the same accuracy can differ in how they arrange a category's objects. Odd-one-out choices show the difference

Scaling up: pure and linear populations

Two columns of panels for 200 simulated neurons in the 24 conditions of the sequence task, a third column left empty. Top row, response matrices, 24 conditions by 200 neurons, recognition rows above recall rows, colored from low to high. Bottom row, scree plots of the variance fraction of components 1 to 23. Pure selectivity: seven bars, then nothing; rank 7, PR 6.2. Linear mixed: seven bars, then nothing; rank 7, PR 5.9

Simulated: 24 conditions, 200 neurons

  • Pure: one variable per neuron. Linear mixed: a weighted sum of all three
  • PCA of 24 conditions × 200 noiseless neurons: without nonlinear mixing, only 7 dimensions

Scaling up: a nonlinear mixed population

The same panels with the third column filled: nonlinear mixed selectivity. Its response matrix looks as busy as the other two. Its scree plot has 23 bars that fall slowly from 0.07 to 0.02; rank 23, PR 20.7

Simulated: 24 conditions, 200 neurons

  • Nonlinear mixed: its own response to each condition. The matrices look alike
  • Only nonlinear mixing uses all 23 dimensions

Can the Theme 1 tools see it? 1 of 3: the cube

One row of panels for the cube population of the demo model, eight conditions A to H. Distances (RDM): an 8 by 8 matrix of Euclidean distances between condition means, with four distinct distance levels. MDS map: red dots A to D (context 1) and blue dots E to H (context 2) around a ring, short gold segments joining conditions that differ only in context and pointing in different directions; stress 0.24. Scree: three equal bars for components 1 to 3, none after; rank 3, PR 3.0. At right: shattering 0.72, context CCGP 0.95

Simulated: the demo's toy population, 30 units, 8 conditions A–H (context × action × reward)

  • Cube (λ = 0): one direction per variable. Gold segments (same action and reward, other context) are parallel
  • On the 2-D map those parallel edges are lost (stress 0.24)

Can the Theme 1 tools see it? 2 of 3: the compromise

The same row for the compromise population, lambda 0.25. RDM: distances more even. MDS map: again a ring of eight dots with short gold segments; stress 0.25. Scree: three large bars and four small ones; rank 7, PR 4.1. At right: shattering 0.80, context CCGP 0.80

Simulated: the demo's toy population, 30 units, 8 conditions A–H (context × action × reward)

  • Compromise (λ = 0.25): each condition = 0.75 × its cube corner + 0.25 × a random vector
  • Shattering rises (0.72 → 0.80), context CCGP falls (0.95 → 0.80)
  • The RDM and the map barely change. The readout scores do

Can the Theme 1 tools see it? 3 of 3: random

The same row for the random population, lambda 1. RDM: every off-diagonal entry the same, every condition equally far from every other. MDS map: gold segments crossing at odd angles; stress 0.31. Scree: seven equal bars; rank 7, PR 7.0. At right: shattering 0.99, context CCGP 0.51

Simulated: the demo's toy population, 30 units, 8 conditions A–H (context × action × reward)

  • Random (λ = 1): random vectors only; all conditions equally far apart
  • Shattering 0.99. Context CCGP 0.51, so nothing transfers
  • RDM, map and PR describe the geometry. Only CCGP tests transfer

Single trials in t-SNE: islands mislead

Three t-SNE maps of single trials, 40 per condition, one dot per trial, blue for context minus and red for context plus, each condition's letter beside its group. Cube: the trials form a ring of touching groups, red and blue alternating. Compromise: groups that touch and partly mix. Random: eight tight, well-separated islands. Under the maps: cube shattering 0.72, context CCGP 0.95; compromise 0.80 and 0.80; random 0.99 and 0.51

Simulated toy population (the three geometries above), single trials; t-SNE perplexity 30

  • Random: cleanest islands, nothing transfers (0.51). Cube: islands touch, context transfers best
  • A t-SNE map suggests groups. A readout on the responses tests them

- The last lecture ended with a linear readout fitted to monkey IT responses, and with XOR, a split of four toy stimuli that no linear decision boundary separates - Today's question: what does a population of neurons encode, and in what format? A readout is the tool we use to measure it: which splits of the stimuli a readout can learn from the responses tells us what the population makes available (shattering) - Then CCGP asks whether a readout trained on some objects works on objects it never saw - Then we compare two representations of the same images, such as a network layer and monkey IT

- The figure is from the last lecture: two units, two classes, and the line a fitted readout draws between them - The readout answers "yes" (class 1, the positive side) when w₁x₁ + w₂x₂ + b > 0, and "no" (class 2, the negative side) otherwise. Swapping the names of the two classes flips the signs of w and b, so which side is class 1 is a convention. The bias b is the threshold with its sign flipped. The weight vector w is perpendicular to the boundary and points into the "yes" side - When a probe learns a split from the responses, any neuron downstream could learn it too. How well the probe does on new stimuli, and where its boundary lies, tell us how the category is laid out in the responses

- These are the monkey IT recordings from the last lecture. Inanimate images are blue, as they were on that slide - A split puts every image in one of two classes: animate or not, face or not, and a random half of the 51 objects. All 24 views of an object get the object's color - The points stay where they are from panel to panel. Only the colors change

- Each boundary is a readout fitted on 18 views of every object and scored on the 6 views left out. Chance is the share of the larger class - Two classes are linearly separable when one linear decision boundary (a line in 2-D, a hyperplane with more neurons) puts them on opposite sides - The random labeling splits the objects into two halves at random, so it means nothing to the monkey. Yet a readout of all 480 neurons classifies 78% of new views correctly, far more than a readout of two components does - The readout is linear in every panel. Only the split changes, and with it how well the IT responses separate the two classes

- Flexibility is what the first half of the lecture measured. A population is flexible when a downstream neuron could learn almost any grouping of the stimuli, which is high shattering - Abstraction is the new property. A population codes a variable abstractly when the variable keeps the same format across conditions, so a readout trained on some conditions transfers to others - Spreading points out in many dimensions makes every grouping separable, but then a boundary learned on some conditions says nothing about the others. Giving each variable its own direction makes boundaries transfer, but some groupings, such as XOR, become impossible - Bernardi et al. (2020) asked where hippocampus and prefrontal cortex sit between these two extremes

- Each group of points is one stimulus. The spread within a group is noise - XOR (exclusive or) says "yes" when exactly one input is on. The colors here mark its complement, which has the same geometry

- We found 75% by trying every linear boundary: every orientation, and every position for each orientation - The boundary drawn is one of the best. Every linear boundary leaves at least one group on the wrong side

Each group of points is one stimulus shown many times; the spread within a group is trial-to-trial noise. From here on we work with one point per stimulus, its mean response, so the question becomes purely geometric: which splits of these points can a line or a plane separate

- The four points sit at the corners of a square, at plus or minus 1 on each unit - Before we answer for four stimuli, we count for three - A split puts each stimulus in one of two classes. Three stimuli give 2 × 2 × 2 = 8 splits, including the two where all three share a class. For those two, the boundary lies outside the points - In 2-D, points are in general position when no three lie on one line. In D dimensions, no D + 1 points lie on one hyperplane - In representation B, the two outlined splits put the middle stimulus against both ends. No linear boundary separates them - The readout is linear in both representations, and the stimuli are the same. Only the responses to them change

- Dimension here is the number of independent directions the three points span. Points in general position in a plane span 2; points on a line span 1 - With D dimensions, N points in general position can be split every way as long as N ≤ D + 1. Three points need 2 dimensions; the points on a line have only 1, so the two splits that separate the middle point from both ends fail

- The four points are the clean XOR points from a few slides back, in the plane of two units (x₁, x₂) - Two units never support all 16 splits of four stimuli, however the four points are placed - To separate the last two splits, the representation itself has to change: Tuesday's lifting unit did that. We return to it with real neurons later today

- Each pixel is black or white by a coin flip. Each image is a point in a space with one axis per pixel, 100 axes - The label is a second coin flip. The pixels carry no rule for the readout to find, so a readout that labels the images correctly has memorized them - The readout has 100 weights, one per pixel, and a threshold - A random labeling is the hardest kind of split, and the one where the count can be done exactly. It tells us how many arbitrary splits a representation of a given size supports - Zhang et al. (2017) did this with deep networks and ImageNet photos: with the labels shuffled, the networks still fit every training photo, and their accuracy on new photos fell to chance. We return to this in the lecture on generalization

- So far the units were built by hand, and the stimuli were clean points. Real neurons were not built for one split, and their responses vary from one showing to the next - A population of neurons has to make available whatever split a new task asks for. Shattering, measured with a readout on held-out trials, tells us how many splits it does

- Figure: Photographs reproduced from Retzius (1906); areas marked here - Both studies recorded outside visual cortex, in areas that hold rules, memories and outcomes - Left, the outer (lateral) face of a rhesus monkey brain; right, the inner (medial) face of a hemisphere cut down the middle. The front of the brain is to the right in both - Lateral prefrontal cortex lies around the principal sulcus (area 46). Its neurons hold task rules and items in working memory, and it is needed for flexible, rule-guided behavior. Review: Miller & Cohen (2001) - The hippocampus supports memory and builds maps, of space and of abstract relations between things. Review: Behrens et al. (2018). It sits inside the temporal lobe, so its dot marks a site under the surface - Anterior cingulate cortex (area 24), above the corpus callosum on the inner face, tracks outcomes and errors and adjusts behavior. Review: Heilbronner & Hayden (2016). That matters in a task where the monkey must notice from surprising outcomes that a hidden rule changed - Rigotti et al. (2013) recorded 237 lateral prefrontal neurons in two monkeys. Bernardi et al. (2020) recorded 1,378 neurons: 629 in the hippocampus, 414 in dorsolateral prefrontal cortex (areas 8, 9 and 46), 335 in anterior cingulate cortex - Both studies pool neurons recorded in different sessions as if they had been recorded together. This is called a pseudo-population - The photographs are by Gustaf Retzius (1906), of a real rhesus brain. The dots and labels are ours - Shattering accuracy turns the count into a measure on data - Pick a random split of the stimuli into two groups of equal size. Fit a readout on training trials and score it on held-out trials. Repeat for many splits and average - 0.5 is guessing, because the splits are balanced. 1.0 means every random split is learned - In the assignment, each stimulus is an object and each trial is one view of it - One name, two quantities. Bernardi et al. (2020) call this average accuracy "shattering dimensionality", because it rises with the dimension of the responses. Rigotti et al. (2013), later in this lecture, use the same name for a count of dimensions. In this course: **shattering accuracy** for the average accuracy, **shattering dimensionality** for the count

- Figure: Adapted from Schiller et al. (1976): the original plot, with our axis labels - The plot is the original from Schiller, Finlay and Volman (1976): one monkey V1 neuron, recorded at 36 orientations, 10 repeats each. It fires most near one orientation and little elsewhere - In the first lecture you saw the Hubel and Wiesel film: a cat V1 neuron firing for a bar of light at one orientation - One variable, orientation, describes this neuron well

- Figure: V1: adapted from Schiller et al. (1976). IT: computed from the Bao et al. (2020) recordings - The bars are one of the 480 IT neurons of the assignment (Bao et al., 2020), averaged over the 24 views of each object. We picked the most face-selective neuron: its mean over the 9 face objects is 16.0 spikes/s against 2.4 for the 42 other objects - Its five strongest objects are all human heads - Like the V1 neuron, it is described by one variable: face or not - The next slides ask what a neuron like this would do when the same objects are used in two tasks

- Figure: V1: adapted from Schiller et al. (1976) and Albrecht &amp; Hamilton (1982), original plots with our axis labels. IT: computed from the Bao et al. (2020) recordings - Right: Albrecht and Hamilton (1982), one V1 neuron. The response grows with contrast and saturates, a monotonic tuning curve instead of a peaked one - All three neurons are described by one stimulus variable: orientation, object category, or contrast

- The task is from Warden & Miller (2010). The monkey holds a bar and fixates for 1 s - Object 1 appears for 500 ms, then a 1 s delay, then object 2 for 500 ms and another 1 s delay. The two objects always differ and come from four objects used that day - In recognition blocks, a test sequence of two objects follows. If it shows the same two objects in the same order, releasing the bar during the second test object earns juice. Otherwise the monkey keeps holding - In recall blocks, an array of three objects appears, the two samples and a distractor. The monkey looks at the two sample objects in the order they were shown - No cue said which task was on. Recognition and recall trials were interleaved in blocks of 100–150 trials (Rigotti et al. 2013) - Rigotti et al. (2013) analyzed 237 lateral prefrontal neurons from two monkeys - The conditions are counted during the second delay, after both objects. Shattering asks how many of the ways to split them into two groups a readout can learn

- The plot shows the firing rate (spikes per second) of one imagined neuron for each of the four objects, A to D, shown first. Red is the recognition task, blue the recall task, the colors of the paper - A purely selective neuron ignores the task, so its tuning curves in the two tasks coincide. They are drawn slightly apart so both are visible - These two neurons are hypothetical, drawn to show what each kind of tuning predicts

- In an additive neuron, the task adds the same amount to the response to every object. The recall line is the recognition line moved up - Such a neuron's response is a weighted sum of the two variables, like the x₁ + x₂ unit. Adding more additive neurons does not raise the rank of the population's responses

- Figure: Reproduced from Rigotti et al. (2013), Fig. 3a–d - One recorded neuron, 0.2 s after the first object, before the second: for object C it fires about 6 spikes per second in task 1 and about half that in task 2, while object A drives it more in task 2 - Rigotti et al. (2013, Fig. 3b) show this neuron 0.2 s after the first object appeared, before the second one. Bars are means with their standard errors. The asterisk marks the difference between the two tasks for object C - Compare with the two expectations. Pure: the red and blue bars would match. Additive: every blue bar would be the red bar plus the same amount. Here C drops from task 1 to task 2 while A rises - Such a neuron responds to a combination of object and task, as the x₁x₂ unit responds to a combination of x₁ and x₂

- Figure: Reproduced from Rigotti et al. (2013), Fig. 3a–d - Rigotti et al. (2013, Fig. 3d) show one lateral prefrontal neuron 1.8 s after the first object appeared, while the second object is on the screen - Each bar is one condition, in spikes per second, with its standard error. Red bars are task 1 (recognition), blue bars task 2 (recall). Under the bars, the first object (Cue 1) and the second (Cue 2) - In recall, the neuron fires most when A or D comes second, and most of all after C (the two labelled bars). In recognition, the same pairs drive it less - A V1 or IT neuron of the previous slides is described by one preferred value. This neuron is not - The next slides go back to the XOR stimuli to see what such mixing does for a population

- These are the idealized neurons of Rigotti et al. (2013, Fig. 1), drawn for the XOR stimuli of the last lecture. Earlier today x₁ and x₂ were two units' responses. Here they are the two variables of the task, and each panel is a neuron that responds to them: a neuron called x₁ (pure) responds to the first variable only - A pure unit follows one input. A linear mixed unit follows a weighted sum of both - The score under each panel is the best a single threshold on that unit can do on the four conditions (the corners). XOR asks whether the two inputs have the same sign - A pure unit gets 2 of 4 right, the linear unit 3 of 4

- x₁x₂ is the lifting unit of the last lecture. |x₁ − x₂| is zero when the inputs agree and large when they differ - The AND unit takes a weighted sum of the inputs, subtracts a threshold, and passes the result through a rectifier, max(0, ·), which sets negative values to zero. That is the model neuron of the next lecture - On its own the AND unit gets 3 of 4 right. Its value lies in what it adds to a population, next - Rigotti et al. (2013, Fig. 1a) draw the same kinds of idealized neurons over two stimulus features. Their nonlinear mixed neuron has circular contours, like the radial unit that lifts the two rings (appendix)

- Rank, from the last lecture, is the number of directions the points span after centering each column - The third column is the sum of the first two, so the four points stay on a plane in the space of the three units. The two XOR splits stay out of reach - Any weighted sum of x₁ and x₂ does the same, whatever the weights

- The x₁x₂ column is +1, −1, −1, +1. No weighted sum of the x₁ and x₂ columns produces that pattern, so the rank goes up to 3 - In three dimensions, four points in general position can be split every way, so a readout can learn all 16 splits

- The AND unit is 0, 0, 0, 1 at the four corners. That is not a weighted sum of the x₁ and x₂ columns either, so it raises the rank to 3 - The separating plane is tilted. A readout now weighs x₁, x₂ and the AND unit together to learn the XOR split - A caveat on "any". A nonlinear unit adds a dimension only if its responses to these conditions are not a weighted sum of the others. x₁³ is nonlinear, but at ±1 it equals x₁ and adds nothing. Units whose responses mix the variables, like x₁x₂, the AND unit or |x₁ − x₂|, do add one

- Bottom row: eight stimuli, the four positions of the square in two versions, for example two colors. The units x₁ and x₂ respond to position only, so each red stimulus lands on a blue one, and the best line classifies 50% - x₃ responds to the new property and is linear in it. With x₃ a plane separates the classes: 100%, and the rank of the eight responses is 3 - What helps is not nonlinearity as such. It is a unit whose responses are not a weighted sum of the others': a nonlinear function of the same inputs, like x₁x₂, or a response to something new, like x₃ - The prefrontal neurons later in this lecture make the same point with real recordings

- Figure: Reproduced from Rigotti et al. (2013), Fig. 4b - Each condition becomes one point, the mean response of the neurons in that condition. Shattering asks how many ways one readout can split those points - The log₂ turns the count into a number of dimensions. N spread-out points with N − 1 units can be split all 2^N ways, and log₂ 2^N = N. Rigotti et al. (2013) call this number the shattering dimensionality - The left axis is the number of splits a readout solves (N_c, log scale). The right axis is log₂ of that number - The same name is used for three related numbers: the count of splits a readout solves, its log₂ (Rigotti et al. 2013), and the mean held-out accuracy over a random sample of splits (Bernardi et al. 2020, and the assignment). All three rise together; check which one a paper reports - The window is the second delay, after both objects. There are 24 conditions (4 first objects × 3 second objects × 2 tasks), so 2²⁴, about 17 million splits, and at most 24 dimensions - Populations larger than the recorded one are built by resampling the recorded neurons - The grey line is a simulated population of pure-selectivity neurons, each coding one task variable, with noise matched to the data. It stays below 8 dimensions however many neurons are added - Mixed selectivity is a property of single neurons. Shattering is a property of the population - Criterion (Rigotti et al. 2013, Supplementary Methods M.7): a readout is trained for each split on the condition means of training trials, and tested on held-out trials. A split counts as implementable if the held-out error is below θ = 0.2–0.25, i.e. at least 75–80% correct. The authors report that the exact θ only rescales the noise and does not change the asymptotic count. With 2^24 possible splits they sample 100,000 at random

- Figure: Reproduced from Rigotti et al. (2013), Fig. 5a - Here the analysis uses the recall task only, which has 12 conditions (4 first objects × 3 second objects). The maximum is 12 - 121 recorded neurons had as many correct as error trials in the recall task and enter this analysis. Larger populations, up to 4,000 neurons, are resampled from them - Values read from the figure: on correct trials the curve ends near 11, close to the maximum of 12. On error trials it levels off near 7 - On error trials, a readout can still tell which objects were shown (Rigotti et al. 2013, Fig. 5b). The objects can be decoded, but the responses of the population are lower-dimensional - Rigotti et al. (2013, Fig. 5c–d) removed parts of each neuron's response. Without the linear part of the mixing, the drop remains. Without the nonlinear part, it disappears - What might be going on: Rigotti et al. (2013) do not identify a cause. What they show is that the objects are still encoded on error trials, and that what is missing is the nonlinear mixing, the conjunctions of first object × second object that a downstream circuit needs to plan the right pair of eye movements. Their hypothesis (Supplementary S.1): the ability to read out many combinations "occasionally goes awry, giving rise to error trials" - Candidate explanations, none tested in this paper: a lapse of engagement or attention on that trial, so the circuits that build the conjunctions are weakly driven; or noise in those circuits on that trial. The result is a correlation: the drop could cause the error, or both could reflect a third factor such as a lapse in attention

- Shattering measures flexibility: a downstream neuron can learn almost any grouping of the stimuli it has seen. It says nothing about stimuli it has not seen - An animal meets new objects and new situations all the time. A readout that had to be retrained for every new object would be of little use - So we ask a second question of the same population: if a readout learns "animate or not" from some objects, does it work on objects it never saw? That is CCGP. The two properties can pull against each other, as the next slides show

- Shattering counts how many groupings of the same stimuli can be solved. CCGP tests whether a readout trained on some objects works on new objects - On the left of the figure, the same points sit in the plane of two units, with three boundaries, one per split. Shattering fits a new readout for every split - The figure's score is the shattering accuracy of the earlier slide: the readout accuracy for each split, on held-out trials, averaged over splits - On the right, one readout is fitted on the filled points, then frozen and scored on the hollow points, conditions never used to fit it. That score is CCGP, cross-condition generalization performance - Train an animacy readout on dogs and chairs. Then test it on cats and tables without changing its weights - A high score means animacy is coded along the same direction for all these objects - In the assignment, the split holds out half of the animate and half of the inanimate objects, and the score is averaged over random halvings - The assignment compares CCGP with the majority-class baseline instead of 0.5, because its two classes have different numbers of objects

- The figure shows two toy populations of two units, 25 noisy trials per object. Filled red dots are dogs and filled blue dots are chairs. The gold boundary is a logistic-regression readout fitted on them. Hollow dots are cats and tables, never used to fit - On the left, the dog-to-cat difference (and chair-to-table) is parallel to the boundary: the readout weights give it zero weight. Animate against inanimate is the same step for both pairs, so the frozen boundary classifies 98% of cats and tables correctly - On the right, the same clouds are slid across the boundary. The four groups are just as separable, so shattering can be just as high, but the frozen boundary gets most new objects wrong (22%) - Responses scattered at random in many dimensions solve almost every split. But a boundary fitted on some objects says nothing about new ones, so CCGP is near baseline - One consistent direction gives high CCGP. The conditions then lie in fewer dimensions, so shattering is lower - A population is flexible when many splits can be solved (high shattering). It is abstract when one readout transfers to new conditions (high CCGP) - The readout in the figure is fitted by logistic regression, a readout whose output is a probability

- Bernardi et al. (2020) trained two monkeys on this task - On each trial the monkey holds a button and fixates. A fractal image appears for 500 ms. The monkey releases the button within 900 ms (R) or keeps holding (H), then gets juice (+) or nothing (−) - Call the images A to D from top to bottom. In context 1, A means release for juice, B hold for juice, C release for nothing, D hold for nothing - In context 2 every image flips its action, and two images also flip their reward. So action and reward are not tied together - A correct response to an unrewarded image avoids a timeout and a repeat of the trial, so the monkey still has a reason to get it right - Context is hidden. The rules flip every 50–70 trials without warning, and nothing on the screen says which context is active - There are eight conditions, 2 contexts × 4 images. Each image fixes the action and the reward, so the conditions are context × action × reward - The neurons are analysed from 800 ms before the next image to 100 ms after it appears. They still hold the previous trial's context, action and reward - From the task slide: in context 1, image A means release for juice and image C release for nothing. In context 2, image A means hold for nothing and image C hold for juice - So A+ and C− happen only in context 1, and A− and C+ only in context 2 - The paper's figures use "−" for an unrewarded condition; the dots are coloured by context

- Figure: Reproduced from Bernardi et al. (2020), Fig. 1c - The bars count only the first time each image appears after a switch. The values are read from the figure - The first image after a switch gets the old response, so it is almost always wrong. An unexpected outcome is the only sign that the context changed - Images 2 to 4 have not been seen in the new context. A monkey that relearned each image by trial and error would be at chance on them. These monkeys are above chance, so one surprise changes their response to the other images too - Bernardi et al. (2020) call this inference. It needs a variable for context that is separate from any one image

- This repeats the slide on shattering three stimuli. The rule generalizes: N stimuli in general position in D = N − 1 units can be split every way by a hyperplane - Bernardi et al. (2020) count only splits into two equal halves, so that guessing is 0.5 for every split

- Four conditions at random positions in 3 neurons, with noisy trials. Random placement puts them in general position with probability 1 - The gold plane is our analysis readout, not the monkey's decision. It decodes the context, 1 or 2, from the neurons: a variable the monkey is never shown - CCGP: fit on one condition per context (here A+ and C+), score with weights frozen on the other two (C− and A−). A+ and C+ is one of four such choices - The four choices (fitted on → tested on: accuracy) - A+, C+ → C−, A−: 0.50 (the plane drawn) - A+, A− → C−, C+: 0.50 - C−, A− → A+, C+: 0.50 - C−, C+ → A+, A−: 0.50 - Mean: CCGP 0.50, chance - Shattering, the held-out accuracy averaged over the three balanced splits, is 1.0

- The plane fitted on A+ and C+ classifies C− and A− perfectly, because each lies next to its context partner. Splits that separate conditions within a context fail, which pulls shattering down to 0.73

- This is the XOR layout from the start of the lecture. Context and value are two perpendicular directions, so a readout for either transfers. Of the three balanced splits, the diagonal one (A+ and A− against C+ and C−) cannot be separated by a plane

- Pushing the corners out of the plane adds a dimension, as the nonlinear units did for XOR. The context and value directions stay nearly parallel across conditions, so the readouts still transfer. This is the geometry Bernardi et al. (2020) propose for the recordings

- Figure: Reproduced from Bernardi et al. (2020), Fig. 3e - The figure is Bernardi et al. (2020, Fig. 3e), for the window from 800 ms before the next image to 100 ms after it appears - Only the short lines are data. Black lines are CCGP, grey lines are shattering accuracy (the axis label says shattering dimensionality; it is the same average accuracy as on the earlier slides). The coloured and white dots come from a model, explained on the next slide - CCGP for context: the readout is trained on 3 conditions per context and tested on the held-out pair, one from each context. There are 4 × 4 = 16 ways to pick that pair, and the score is averaged over them - Values read from the figure: context CCGP is 0.96 in HPC, 0.74 in DLPFC and 0.81 in ACC. Value is 0.79, 0.88 and 0.94. Previous action is 0.44, 0.75 and 0.90. Shattering is given in the paper's legend: 0.70, 0.75 and 0.74 - CCGP is compared with a geometric null model, which scores randomly arranged conditions. Its ±2 s.d. band is about 0.41–0.59. Context is above it in all three areas - Previous action in the hippocampus is below chance for CCGP. It can still be decoded when all conditions are in the training set (Bernardi et al. 2020, Fig. 3a)

- Each dot is the mean response of the population across that condition's trials, in the window before the next image - The dots live in the space of hundreds of neurons. The figure shows them in three dimensions to make the geometry visible. It is a schematic, not the data - A perfect cube makes every variable transfer (CCGP high). But XOR-like groupings stay unsolvable (shattering low) - Random points make every grouping solvable (shattering high). But nothing transfers (CCGP near baseline) - Bernardi et al. (2020, Fig. 3e) tested the perfect cube. They put the eight conditions at the corners of a box whose sides give the measured CCGPs, added trial noise, and computed shattering 100 times - The box gives shattering of about 0.60 in HPC, 0.62 in DLPFC and 0.65 in ACC (the white dots, read from the figure). The recordings give 0.70, 0.75 and 0.74, higher than every run of the box - Bernardi et al. (2020) propose that the recordings sit near the cube, with each corner pushed off a little in extra directions by nonlinear mixed selectivity. The context edges stay nearly parallel, so context CCGP stays high. The pushed corners make the other groupings separable too, so shattering is high - On error trials, context CCGP drops in all three areas. Decoding of context, with all conditions in training, does not drop - Two scores do not pin down a geometry: many arrangements give the same shattering and CCGP. The scores rule geometries out (random, clustered), and the cartoons are landmarks, not reconstructions. Bernardi et al. (2020) add three further lines of evidence: the parallelism score (PS), which directly measures how parallel the coding directions of a variable are across conditions; explicit models, such as the perfect box matched to the measured CCGPs (its shattering is too low) and random geometries as null models; and 3-D MDS plots of the eight condition means (their Fig. 3C–D), which show the arrangement changing over the trial as context becomes more abstract

- Toy model, not fitted to data. Eight conditions (context × action × reward) are the mean responses of 30 model units, plus trial-to-trial noise - The readouts are linear, fitted on training trials and scored on test trials - The random part is set twice as long as the cube, so that the mixed setting scores high on both - In the cube setting each variable has its own direction. One boundary per variable transfers to unseen conditions (CCGP about 0.95). But odd groupings cannot be split (shattering about 0.72) - In the random setting every condition sits in its own direction. Almost any grouping can be split (shattering about 0.99), but nothing transfers (CCGP about 0.5) - At the compromise (mixing 0.25) both are about 0.8. This is a pattern like the one Bernardi et al. (2020) measured in hippocampus and prefrontal cortex - Decoding here is the readout accuracy with all conditions in the training set. It is high in all three cases (above 0.93). Only CCGP tells them apart - The demo is on the course site with the other demos

- Figure: Reproduced from Courellis et al. (2024), Fig. 2e–f - Courellis et al. (2024) recorded single neurons in 17 adult patients with drug-resistant epilepsy. They had depth electrodes implanted for seizure monitoring, and microwires in the electrodes record single neurons - 2,694 neurons in 36 sessions. 494 were in the hippocampus, the others in the amygdala, frontal cortex and ventral temporal cortex - The design is the one of the monkey task. The two contexts invert every image–response pairing, and the context switches every 15–32 trials - The contexts also change which images give the large reward, so context, response and reward are not tied together - After one error, a patient who knows the structure can switch all four responses at once. As with the monkeys, the first trial after a switch is almost always wrong - Sessions are split by the first image not yet seen in the new context. In inference-present sessions (22), accuracy on it was about 0.89. In inference-absent sessions (14), it was about 0.47, near chance (values read from the figure) - The figure is Courellis et al. (2024, Fig. 2e–f), hippocampus, 0.2 to 1.2 s after the image appears. The number of neurons is matched between the two kinds of session - Each split of the 8 conditions into two groups of four is one variable, 35 in all. Panel e is decoding accuracy with all conditions in training, panel f is CCGP - Stim pair groups images A and C against B and D. In each context, those images share a response - Values read from the figure: context CCGP rose from about 0.50 to 0.63, stim pair from about 0.56 to 0.70. Shattering, the mean decoding accuracy over the 35 splits, rose from 0.57 to 0.62 (given in the text) - Parity is the XOR-like split of the cube. Its decoding rose too, a sign of nonlinear mixing - No other recorded area showed this change. On error trials in inference-present sessions, the hippocampus looked like it did in inference-absent sessions - The same geometry appeared in patients who learned the structure from verbal instructions instead of by trial and error - In the authors' words, "only the neural representations formed in the hippocampus simultaneously encode several task variables in an abstract, or disentangled, format" - Bernardi et al. (2020) found abstract context in the hippocampus of trained monkeys. Here it appears in people, and only in sessions where they inferred the context

- The four measures describe one set of responses at a time, such as monkey IT, the pixels, or one network layer - A representation can score high on one measure and low on another. Bernardi et al.'s recordings score high on shattering and CCGP together - The rest of the lecture puts two representations side by side - If time runs short, the comparison section below opens the next lecture.

- So far each score read one population. Now two systems see the same images, such as a network layer's units and IT's neurons, or a network and people judging the objects - We cannot match unit 17 to neuron 17, so every score in this section compares how each system arranges the images - A comparison score takes two systems that saw the same images - IT has 480 neurons and a late network layer has thousands of units. No neuron corresponds to a particular unit, so a comparison score must work without such a match - Each score ignores some differences. For example, CKA ignores a rotation of the units. Ask whether your question should ignore it too - The same four questions apply to any comparison score you meet in papers

- The RDM is the table from the MDS lecture - The same 51 Bao objects are seen by two systems, 480 neurons in a monkey's IT and the 4,096 units of ResNet-18's late layer - Neuron 17 does not correspond to unit 17, so we compare what each system says about every pair of objects - The correlation distance 1 − r is 0 when two objects evoke the same pattern across the neurons, and larger as the patterns differ - Each object's response is averaged over its 24 views before the RDM is built. The assignment builds RDMs on individual images. Both are valid. Say which you used - Any change that keeps the order of the dissimilarities leaves RSA unchanged - You computed RSA in the first assignment against human judgments. Here the second system is a monkey's IT - Spearman, as in the MDS lecture: a bent but rising relation between the two RDMs keeps every rank (Spearman 1). Only a swap of order lowers it - We use Spearman because the two systems share no scale - We use the upper triangle of each RDM (1,275 pairs), because an RDM is symmetric with a zero diagonal - Computed on the 51 object means with every unit z-scored, RSA is 0.712 (0.595 without z-scoring) - Shuffling the object order of one RDM gives values near 0. This shuffle check returns in the next lecture - RSA is not a percentage of shared information

- RSA starts from distances between images. The MDS lecture had a second way to compare two response vectors, the dot product. CKA is the comparison score built on it, an alternative to RSA rather than a new idea - The correlation distance 1 − r of the RDM divides out each vector's length and mean - The dot product of two response vectors is large when both are long and point the same way. The cosine similarity divides the lengths out - Keeping the lengths means that an image that drives the neurons strongly counts for more. Ranks would throw that away - RSA compares the order of the distances. CKA compares the dot products - The assignment computes RSA on correlation distance - The assignment first passes raw responses to CKA. Later it also z-scores each unit, which makes CKA ignore each unit's scale - Davari et al. (2023) showed that CKA is sensitive to a few outlying points and high-variance directions. It can change a lot without a change in what a network computes - With few images and many units, CKA between unrelated representations can come out high. The next lecture shows why

- Layer counts: ResNet-18 = 1 stem convolution + 8 blocks × 2 convolutions + 1 final fully connected layer = 18. ResNet-50 = 1 stem convolution + 16 blocks × 3 convolutions (3, 4, 6 and 3 blocks in its four stages) + 1 fully connected layer = 50 - Center each unit first, by subtracting its mean response over the images. Uncentered CKA is dominated by the mean response and reports about 1 for almost anything - The Gram matrix has one row and one column per image, whatever the number of units. Two systems with different numbers of units give tables of the same size - To compare the two tables, CKA treats each one as a long list of numbers and measures their similarity, the cosine of the angle between the two lists - The heat map shows the average-pooled output of every block of an ImageNet-trained ResNet-18 (stem + 8 blocks) and ResNet-50 (stem + 16 blocks), on 1,854 THINGS images - Each ResNet-18 block matches best the ResNet-50 blocks at roughly the same relative depth (gold rings) - Most off-band cells are still above 0.5. Compare cells with each other instead of with zero - The assignment provides `linear_cka` and uses it on IT, on network layers and across networks

- RSA and CKA compare two response matrices. People's choices are behavior. They do not form a response matrix, so the comparison needs a score that works on choices - Cosine similarity (MDS lecture) is the dot product with both lengths divided out. It is 1 when two vectors point the same way - In the figure, ResNet-18's late layer puts cow and horse closest (0.77), so it picks the axe - People agree with each other on about two thirds of triplets. A network's agreement should be compared with that level instead of with 1 - The assignment computes this agreement for six networks on 88,448 test judgments

- Bernardi et al. (2020): read the Introduction and Figures 1–3. Skip the Methods - Rigotti, Barak, Warden, Wang, Daw, Miller & Fusi (2013), The importance of mixed selectivity in complex cognitive tasks - Courellis, Minxha, Cardenas, Kimmel, Reed, Valiante, Salzman, Mamelak, Fusi & Rutishauser (2024), Abstract representations emerge in human hippocampal neurons during inference - Johnston & Fusi (2023), Abstract representations emerge naturally in neural networks trained to perform multiple tasks - Kriegeskorte, Mur & Bandettini (2008), Representational similarity analysis: connecting the branches of systems neuroscience. This paper introduced RSA - Kornblith, Norouzi, Lee & Hinton (2019), Similarity of neural network representations revisited - Hebart, Zheng, Pereira & Baker (2020), Revealing the multidimensional mental representations of natural objects underlying human similarity judgements - Muttenthaler, Dippel, Linhardt, Vandermeulen & Kornblith (2023), Human alignment of neural network representations - Before posting, skim the existing entries on your paper. Your post must add something not already said

- Submitted on Canvas (Minute papers → Minute Paper 7). Credit for a thoughtful attempt

- Each score is the best over every linear boundary: a line in 2-D, a plane for the 3-D roll - The best linear boundary classifies 75% of the XOR points, 68% of the ring points, 90% of the moon points and 75% of the Swiss-roll points - The Swiss roll is the one Isomap unrolled: 800 points on a rolled-up sheet. Here the points on the first half of the sheet are blue and those on the second half red

- XOR: the product x₁x₂ is positive where the inputs share a sign and negative where they differ - Rings: the radial unit grows with the distance from the centre, so the outer ring rises above the inner one - Moons: the new unit has Gaussian tuning, like the tuning curves of the first lecture: exp(−‖x − c‖² / 2σ²), centred on c = (0, 0.25), the LEFT tip of the red moon, where it pokes into the hollow of the blue moon, with width σ = 0.5. We chose c and σ by hand. Red points at that tip respond about 0.8; the nearest blue points, about 0.6 away, respond at most 0.5 (median blue 0.23). The rest of the red moon, far to the right, responds near 0, but there x₁ and x₂ already separate it. The plane uses all three units together, not the new unit alone - Swiss roll: a unit that reports position along the sheet does what Isomap did in the t-SNE and UMAP lecture. If you know the manifold, or have learned it, a coordinate along it can make a class linearly separable that is not separable in the original coordinates. One threshold on that unit splits the roll - We built each unit by hand for its problem. In a later lecture a network's hidden layer learns such units from examples

- Muttenthaler et al. (2023) tested networks trained with labels, with self-supervision and with image–text pairs - Matched pairs of networks that differed only in the objective differed in agreement - A training objective is what a network is trained to do, such as predicting labels, matching two views of an image, or matching images to captions. How training works is the subject of a later lecture - The odd-one-out choice depends on which objects sit closest, inside a category as well as across categories

- The design is the task of Rigotti et al. (2013), with 2 tasks × 4 first objects × 3 second objects. The second object is one of the three objects not shown first, so its identity takes four values - Each population has 200 simulated neurons and no trial noise. The three populations have the same total variance - With noiseless responses, the rank is the number of components with nonzero variance. With recorded neurons every component has some variance, so in practice we count the components needed for, say, 90% of the variance: the embedding dimension of last lecture - A pure neuron responds with one value per level of its variable, for example one value per first object. A linear mixed neuron adds one value for the task, one for the first object and one for the second object - Why 7. The task has 2 values, which after centering give 1 direction. The first object has 4 values, 3 directions. The second object, 3 more. 1 + 3 + 3 = 7 - The participation ratio (last lecture) is 6.2 and 5.9. The seven directions do not share the variance equally

- How the toy data were made. We simulate three separate populations, each of 200 neurons, all responding to the same 24 conditions (2 tasks × 4 first objects × 3 second objects). Each is a response matrix of 24 conditions × 200 neurons, with no trial noise - Pure population: each neuron is assigned at random to one variable (task, first object or second object). It gets one random value (drawn from a normal distribution) for each level of that variable, and responds with that value whatever the other variables are - Linear mixed population: each neuron gets one random value per task, one per first object and one per second object, and responds with their sum - Nonlinear mixed population: each neuron gets an independent random value for each of the 24 conditions, the extreme case. Real neurons sit between this and the linear case - The three populations are scaled to have the same total variance, so their scree plots can be compared - 24 points span at most 23 directions after centering, so 23 is the maximum. The participation ratio is 20.7 - The measurement is the slide "Mixed neurons make a high-dimensional population". Rigotti et al. (2013) simulated pure-selectivity neurons too. Their count levels off far below the maximum, and the recorded neurons do not

- The model is the one in the demo you just tried, at three settings of its slider λ. Each condition's mean is a point in the space of 30 units, and each trial adds noise - The cube puts context, action and reward along three perpendicular directions - The RDM holds the Euclidean distance between every pair of condition means. The MDS map places the eight means in two dimensions to match those distances. The scree plot is PCA on the eight means - The gold segments join pairs that differ only in context. In the 30 units they are exactly parallel on the cube (0 degrees apart). The 2-D map cannot show that. It squeezes a 3-D cube into a plane, with stress 0.24

- Each condition's mean is 0.75 times its cube corner plus 0.25 times a random vector of its own, twice as long as the corner. This is the demo's slider at 0.25 - The four extra components are small (4% of the variance each), so the participation ratio rises only from 3.0 to 4.1 - The gold context segments now differ by 53 degrees on average in the full 30-unit space

- Eight points all at the same distance from each other form a regular simplex. It needs 7 dimensions, so any 2-D map distorts it. The MDS stress, 0.31, is the highest of the three - The participation ratio climbs 3.0, 4.1, 7.0 while context CCGP falls 0.95, 0.80, 0.51. PR counts directions. It does not say whether a readout for context transfers - Shattering and CCGP come from the demo's own computation, with 35 balanced splits and 16 held-out pairs per variable

- Minute paper 5 asked how to check whether t-SNE islands are really in the network's responses. Here the islands are real in all three populations, yet the map cannot tell which population codes context the same way for every condition - t-SNE keeps each trial's nearest neighbours and discards the distances between islands (the t-SNE lecture). The cube's conditions are closer to each other than the random population's, so its islands touch - Neighbours agree with the condition labels for 90% of cube trials in the map (87% in the original 30 units), 87% for the compromise (79%), and 100% for random (100%). Each count uses the 5 nearest neighbours