For two samples, (tiger) and (elephant):
Two features (plot): eagle or penguin?
| Pair | Euclidean | Cosine* | Correlation* |
|---|---|---|---|
| tiger–eagle | 23.8 | 0.021 | 0† |
| tiger–penguin | 54.1 | 0.009 | 0† |
4,096 network responses: penguin or gorilla?
| Pair | Euclidean | Cosine* | Correlation* |
|---|---|---|---|
| tiger–penguin | 122.8 | 0.649 | 0.911 |
| tiger–gorilla | 125.7 | 0.635 | 0.897 |
* dissimilarities, and · † with two components, always: correlation needs more than two
A long tradition in psychology: what we know of objects is arranged as points in a space, near when alike. That psychological space cannot be measured directly. What can be measured is a judgment. Peterson and colleagues asked people: how similar are these two images, from 0 to 10? Ten raters for every pair of 120 animal photographs.
Mean rating A–B: 10.0 · A–C: 2.9 · B–C: 2.2 · one number per pair: a dissimilarity matrix, from which the space is to be recovered
Will it work? Yes, if the matrix is, or is well approximated by, a table of distances between points in that space. What a table of distances must satisfy, and which of our measures satisfy it:
| Requirement | Mathematical statement | Euclidean | Cosine | Correlation |
|---|---|---|---|---|
| Nonnegativity · no negative distances | ✓ Yes | ✓ Yes | ✓ Yes | |
| Identity · zero only for an image and itself | ✓ Yes | ✗ No | ✗ No | |
| Symmetry · same both ways | ✓ Yes | ✓ Yes | ✓ Yes | |
| Triangle · no shortcut via a third image | ✓ Yes | ✗ No | ✗ No | |
| A distance in the full sense (a metric)? | All four | ✓ Yes | ✗ No | ✗ No |
✓ holds throughout the domain; ✗ can fail even when the measure is defined. On our images: 2.3 % of pixel triples break the triangle inequality under cosine, none in the trained late layer, and a brighter copy of an image sits at cosine distance 0 from it. You test this yourself: Assignment 1 bonus B1 · worked examples in the appendix.
Given a dissimilarity matrix over stimuli, with entries .
Find coordinates, one point per stimulus,
such that the map distances reproduce the measured dissimilarities :
Real dissimilarities are noisy; often no exact Euclidean configuration exists. Minimize the mismatch instead:
sklearn doesn_init)sklearn can print it. Rule of thumb (Kruskal 1964): 0.05 excellent · 0.10 fair · 0.20 poor. Assignment 1 §2e adds a real check: the stress of the same fit on a shuffled matrix, which has no structure to recoverIf a story about your map depends on which way is "up", it is a story about the plot.


Assignment 1 §2g fits three curves to the same (distance, similarity) pairs:
Two systems, the same stimuli, two matrices: from the judgments, from a layer. Flatten each upper triangle into a list of numbers and correlate the two lists:
In practice: iu = np.triu_indices(N, k=1), then spearmanr(D1[iu], D2[iu]).
The operation has a name: representational similarity analysis, RSA

Seven seeds, none required. Branch from any of them; pick a paper your neighbour has not picked.
Canvas submission: Canvas → Minute papers → Minute Paper 3 (access code read out in class)
Write three brief points in your own words:
Credit for a thoughtful attempt, not for being correct.
Subtraction: the displacement from image 1 (tiger) to image 2 (elephant):
is itself a vector: from the tiger's point to the elephant's.
ReLU (rectified linear unit): input , output .
Similarity ratings come from people, not from a formula. A network's RDM inherits its properties from the measure ; a rater is free to answer anything. Before we recover a space from such numbers, what must they satisfy? For tiger, penguin, gorilla ():
The first three make a dissimilarity matrix. The fourth is what points in a space obey: without it, no map reproduces the entries exactly. Which measures satisfy which: the distance table in the lecture. Whether people do: the assignment bonus, and the Tversky slides later in this appendix
np.polyfit for the line, curve_fit for the two curvesSplit-half reliability: correlate two independent sets of ratings for the same pairs. That agreement is the noise ceiling: how well the measurement predicts itself, the bar no model should beat.
| Comparison | Statistic |
|---|---|
| Rating half 1 versus rating half 2 | Reliability of the human measurement |
| Network distances versus human dissimilarities | Representational alignment |
In practice: scipy.cluster.hierarchy.linkage(squareform(D), method="average").
Step 3 is the whole design decision:

the dashed line is the cut; four colors are the four proposed groups
| MDS | Hierarchical clustering | |
|---|---|---|
| Output | continuous coordinates | a nested tree |
| Structure assumed | a smooth space — items vary along dimensions | discrete categories — items belong to groups |
| Fits well | perceptual continua (hue, pitch, size) | taxonomies, category structure |
| Free parameter | number of dimensions | linkage rule + where to cut |
| Fails on | superordinate terms, deep hierarchies | continuous variation |
Run both. Where they disagree is the interesting question.
Every method so far assumes dissimilarity is a distance. Tversky, with data: human judgments violate all three axioms.
Minimality,
Symmetry,

lamp ~ moon (both luminous) · moon ~ ball (both round) · lamp ≁ ball
One correlation is a point estimate. Papers add three things; each has a catch.

each letter is one video, colored by its dominant category
The 2-D layout here uses a nonlinear cousin of MDS, taught in a later lecture.
- Tuesday ended at the representational dissimilarity matrix: every pair of images compared, without saying what the comparison is - First half: the comparison, three measures and their properties - Second half: the reverse construction; given an RDM, recover the space the images live in
- Question 1 of Tuesday's minute paper referred to slides not reached in class and is credited to everyone; the answer is the first half of today - The questions are quoted from the minute papers, lightly shortened. The first is the gap Tuesday left: how a pair of rows becomes one number of the RDM - Channels, in a convolutional network: one channel is a sheet of units sharing the same weights, the same feature detector applied at every position; the sheet is a map of where that feature occurs. Each unit is one entry of x, one input dimension in the machine-learning sense. The feature carries no label; the lecture on convolutional networks examines what it detects. A recording channel is one electrode, an unrelated use of the word - Neural vectors: one entry per neuron, each entry a firing rate averaged over that image's presentations; the Assignment 1 monkey data are those averages - Real values: dividing 0–255 by 255 changes units, not information; every image is rescaled the same way, so which images resemble which does not change
- Minute papers asked for the linear algebra to slow down. The pace stays: the course has a lot to cover, and the material is front-loaded; the objects of Theme 1 are reused all semester - Vector operations and matrix products now, the eigendecomposition next week (PCA), gradients in Theme 2 (learning). Nothing beyond that is required for the assignments - Office hours Wednesday 1 PM in Carney 419 (email first so I expect you). TA office hours: vote on Ed; they will be scheduled once there is demand - Two handouts, on the course site under Recitations and on Canvas: linear algebra for the course (vectors, dot products, norms, matrices, the three measures, worked numbers) and programming for Assignment 1 (NumPy, the library calls, the checks). The notation reference defines every symbol in the order the course uses it - Assignment 1 Part 1 is the steepest part of the whole assignment; Parts 2 and 3 reuse what Part 1 builds
- The step Tuesday skipped: two rows of the response matrix, compared, fill one cell of the RDM - The response matrix is the Assignment 1 activation file: 120 images × 4,096 stored late-layer activations of ResNet-18 - The same picture holds for neurons (120 × 482), voxels and human ratings; only the measure d and the meaning of a row change
- Zero diagonal: an image compared with itself shows no difference, d_ii = 0. Symmetry: a comparison does not depend on which image is named first, d_ik = d_ki. Nonnegativity: nothing is less different than identical, d_ik ≥ 0. Any measure that computes a difference between two vectors gives all three; the three of today do - Ratings need not obey them: a rater can say tiger-to-penguin 3 and penguin-to-tiger 5, or rate a photograph against itself below 10. The usual repair averages the two directions and sets the diagonal to zero; the appendix shows where people depart from the constraints - D is images by images; entry (i, k) compares the response vectors of images i and k. Lower-case d is the dissimilarity measure; different measures keep different properties of the responses, and choosing d is a scientific decision - The matrix is the one you build in Assignment 1: 120 images ordered by eight animal categories, correlation distance on the trained network's 4,096 late-layer activations. Within-category mean 0.82, between 0.92; primates and carnivores form the clearest blocks. Colour scale clipped to the 2nd–98th percentile - Pairwise human ratings fit the same format with no response vector measured, so systems can be compared without equating their units or components
- Euclidean distance is the length of the difference vector, the straight-line separation between two points; the double bars denote vector length, also called the Euclidean norm - In two dimensions Pythagoras gives sqrt(55.6² + 21.4²) = 59.5 for tiger and elephant; in D dimensions the sum of squared component differences runs over all D components - Reversing the subtraction changes the direction of the difference but not its length - Euclidean distance counts a change of scale as a difference: doubling every entry of the tiger's vector (a brighter copy of the same image) moves its point to 2x, at distance ‖x‖ = 83.3 from the original, even though the pattern across components is unchanged; cosine similarity ignores this change - Component scales matter: changing a feature's units changes its contribution; multiplying every vector by the same positive constant multiplies all distances by that constant, and adding the same vector to every image leaves all distances unchanged
- Multiplying every pixel by a constant c multiplies both features by c (the mean and the standard deviation of the pixels both scale), so the point moves along the ray through the origin. The points are measured on the scaled images themselves: at 1.5x a few of the 784 resampled pixels saturate at white, so that point sits a hair below the ray (108.1, 58.2 instead of 109.0, 60.9) - Euclidean distance reports a change of brightness as a difference between images; the same happens for a population that fires harder or a layer with larger activations. The next measures compare directions instead
- The dot product of two response vectors is the sum of the entry-wise products; the identity with the cosine of the angle between them is what makes it a measure of alignment. Numbers: tiger (72.7, 40.6), elephant (128.3, 19.2), dot product 10,104, lengths 83.3 and 129.7, cos φ = 0.936, φ = 20.7° - Dividing the lengths out gives cosine similarity, the next slide; the projection picture returns when a neuron's weight vector plays the role of x^(1)
- The vector here is the image itself, 784 pixel values, not the two summary features of the previous slides. Correlation subtracts the mean over a vector's own entries; on the (mean intensity, contrast) plane those entries are the mean and the contrast, so centring there does not remove a shift of the mean. On the pixels the invariance is exact: x + c·1 minus its mean equals x minus its mean - Cosine similarity of 0.991 and 0.978 is close to 1 but not 1: cosine sees the offset, correlation does not. Euclidean distance grows linearly with the offset, 28·c for a 784-pixel image - The pixel values are not clipped: the tiger's brightest pixel is 217, so +60 takes it to 277, above the 255 a camera would record. With unclipped values the three centred images are identical and r is exactly 1; the displayed +60 image shows its brightest pixels saturated at white - The same pair of nuisances exists for neurons: a change of gain (the previous slide) and a change of baseline firing (this one); which measure you choose decides which of the two counts as a difference
- Cosine similarity is the dot product divided by both vector lengths, cos(phi): 1 when two vectors point the same way, 0 when they are perpendicular. Doubling every response leaves it unchanged (next slide, middle panel: Euclidean distance 13.8, cosine similarity 1). Two vectors on the same line through the origin have cosine similarity 1 and cosine dissimilarity 0; their Euclidean distance equals their difference in length - Pearson correlation is cosine similarity after each pattern has had its own mean subtracted, across its units. Adding a constant to every unit leaves it unchanged (next slide, right panel: cosine similarity drops to 0.978, correlation stays 1). Centering is done within each image across its units, not within each unit across images - These are similarities, large when patterns are alike. Distances are dissimilarities, large when patterns differ. To put every measure in one RDM we use dissimilarities, so cosine and correlation enter as 1 minus the similarity - Practical difference: cosine ignores overall gain (a brighter image, a more excitable neuron); correlation also ignores a shared offset (a baseline firing rate, a mean activation). Euclidean distance keeps both - A version of correlation applied to ranks is used later today, when two RDMs are compared. Cosine is undefined for the zero vector and correlation for a constant vector
- In the two-feature plot the eagle's point is nearer the tiger's (Euclidean 23.8 against 54.1), but the penguin's arrow is more nearly parallel to the tiger's (7.5 degrees against 11.9), so cosine dissimilarity ranks the penguin closer (0.009 against 0.021). Correlation cannot rank them: with only two components, centering leaves (a, −a), so every pair has r = ±1; correlation is a measure for long vectors - The same happens with the network's late-layer responses: Euclidean distance ranks the penguin closer to the tiger than the gorilla (122.8 against 125.7); cosine (0.649 against 0.635) and correlation (0.911 against 0.897) rank the gorilla closer - The differences are small, and with real responses the measure often decides close calls. Euclidean distance asks about position; cosine and correlation ask about direction alone - Assignment 1 builds all three measures and applies them to pixels and to early, middle and late activations; expect the rankings to disagree
- The count is automatic: open-access papers whose full text mentions representational dissimilarity matrices or representational similarity analysis, the sentences around the matrix classified by keyword; a paper counts once per measure, so the shares exceed 100 percent - The estimate is rough in two directions: a measure named elsewhere in the paper is missed, and "Euclidean" near an RDM can refer to something other than the RDM's own measure; the ranking is robust, the percentages are not
- Top row, the 120 animals of Assignment 1. Pixels: correlation distance on the 28 × 28 grey version of each image; within-category mean 0.985, between 0.995, no category structure. People: 10 minus the mean rating over ten raters; within 3.7, between 8.6, the strongest block structure on the slide, with birds, primates and reptiles the tightest. ResNet-18 late layer, 4,096 stored units, correlation distance: within 0.82, between 0.92 - Bottom row, the 1,224 object images of Bao and colleagues (Nature 2020), six categories. Pixels: within 0.445, between 0.466. Monkey IT, 482 neurons, correlation distance on the trimmed firing rates: within 0.91, between 1.02, faces and animals the clearest blocks. The same ResNet-18 layer on these images: within 0.60, between 0.75, with faces the tightest block - Each row is one stimulus set seen by three systems; the two rows are different image sets, so their matrices are not comparable with each other, only within a row - Every panel obeys the same three constraints, zero diagonal, symmetry, nonnegativity; that is all the second half of today needs from a matrix
- The second half inverts the first: start from the matrix and find points whose distances reproduce it. Only the table is used, never the responses - The matrix is the human dissimilarity D = 10 − mean rating for six of the 120 images; the map is the two-dimensional configuration from metric MDS (scikit-learn), with three of the fifteen map distances printed on their segments - Six points on a plane have twelve free coordinates (fewer after removing rotation and translation) against fifteen distances to match, and ratings are not exact distances, so every pair is off by a little; stress measures that mismatch
- The idea of a psychological space is old: objects, colours, sounds and words as points, with nearby points alike, so that generalization, confusion and categorization are facts about distance. Shepard built a career on it. No behavioural method reads the coordinates of a point in that space - What can be measured is what a distance would produce: a judgment of how alike two things are. Psychophysics has many designs for this (direct ratings, confusions, odd-one-out choices, sorting); all end as one number per pair, a dissimilarity matrix. MDS goes from that table back to points - Peterson, Abbott & Griffiths (2018) collected ratings on a 0 to 10 scale for every pair of 120 animal photographs, ten raters per pair, 7,140 pairs in all. The numbers on the slide are their mean ratings: the two tigers A and B were judged maximally similar (10.0), and each tiger about equally dissimilar from the monkey (2.9 and 2.2). Tuesday's opening slide used the same two tigers with a third tiger in place of the monkey - Behaviour is a representation too: the pattern of ratings describes how the images are organised in people's heads without saying anything about neurons - Assignment 1 uses these ratings and asks which measurement, pixels, IT neurons or a network layer, comes closest to reproducing them; other formats such as odd-one-out triplets give similar structure
- MDS recovers a space in which the images are points whose separations reproduce the RDM entries. Points on a page satisfy all four requirements, so an RDM that violates one cannot be reproduced exactly by any arrangement of points; the recovered map distorts it. For a network's RDM this depends on the measure d; for a human matrix nothing guarantees it (appendix) - The rows: nonnegativity, no pair of images is less than zero apart; identity, an image is at distance zero only from itself, which cosine breaks because a brighter copy of the tiger is at cosine dissimilarity zero from the original; symmetry, tiger to elephant equals elephant to tiger; triangle, no route through a third image is shorter than the direct comparison, which cosine violates on 2.3 % of pixel triples in our images - Only Euclidean distance is a metric: it is the length of a difference vector. Cosine and correlation discard gain (correlation also offset), so identity fails, since x and 2x count as the same pattern, and so does the triangle inequality - They remain useful. Cosine dissimilarity is a monotone function of the angle between the vectors, and the angle is a proper distance (the arc between two points on a sphere), so cosine and correlation order pairs of images as a distance would; comparing RDMs uses that order. √(2(1 − cos φ)) is the Euclidean distance between the two length-normalized vectors, and it is a metric. Expect small distortions in a map recovered from cosine or correlation dissimilarities - A cross means the requirement can fail, not that every pair violates it; one counterexample disproves a metric, and no finite check proves one. Assignment 1 builds these checks, with a numerical guard for zero or constant vectors
- Stimulus i is the point y^(i), a vector with m entries (here m = 2); the map distance between two points carries a hat to keep it apart from the measured entry d_ik - Each arrow is a correspondence the map has to honour for all fifteen pairs: an entry in the table is a distance on the map
- MDS takes only the dissimilarity matrix D and returns one point per stimulus in m dimensions, placed so that map distances reproduce the measured dissimilarities; the raw responses, the channels and the stimuli are never used - Analogy: given a table of driving distances between cities, draw the map
- Ekman (1954) had observers judge the similarity of 14 monochromatic lights from 434 to 674 nm; Shepard's (1962) MDS of those judgments returns a circle ordered by wavelength - The recovered structure is a closed curve: one dimension (position around the ring) drawn in two, which is why the fit needs two dimensions - The closure of the circle, violet next to red, is a psychological fact with no physical counterpart: the stimulus dimension is a line, the percept is a ring
- Physical analogy: each pair (i, k) is a spring with rest length d_ik, stress is the total potential energy, and the layout relaxes into the least-strained configuration - Lowering a number by moving the points a little at a time is the same downhill idea that trains the networks of Theme 2 - A local minimum is a layout that no small move improves, although a better layout exists elsewhere
- scikit-learn's MDS solver, run from one fixed random start and stopped after 0, 10 and 300 rounds; each frame is the solver's actual state - Moving each point by an amount computed from the mismatch of its distances is the optimisation idea of Theme 2
- After ten rounds the stress has fallen from 217 to 64: the coarse layout, which animal is near which, is already in place, and the remaining rounds shrink the small mismatches
- Convergence means no small move of any point reduces the stress; 22 is what remains because the fifteen ratings are not exactly the distances of any six points on a plane - Different random starts can converge to different local minima; scikit-learn's n_init runs several and keeps the best, which is what Assignment 1 does
- Raw stress has squared-distance units; dividing by the summed squared map distances removes the overall scale, so Stress-1 can be compared across analyses. In nonmetric MDS the fitted targets are disparities, a monotone transformation of the dissimilarities - Kruskal's (1964) rule of thumb (0.05 excellent, 0.10 fair, 0.20 poor) is not a significance test: item count, target dimension, ties and the optimiser all move the value, which is why Assignment 1 also refits on shuffled dissimilarities - Stress values in this lecture come from scikit-learn 1.8
- The common error is reading the horizontal axis of an MDS plot as if it were a principal component; PCA gives ordered, meaningful directions and MDS does not, which is the cleanest way to tell the two methods apart before the PCA lecture - On dimensionality, the answer is the elbow in the stress-against-m curve together with a stated choice
- Shepard's 1962 result: ordinal information alone, given enough stimuli, pins down a metric configuration almost uniquely; that is why non-metric MDS became the standard tool in psychology - Which falling curve is recovered is itself a finding; Shepard's law, two slides on, is the most famous answer. Shepard 1962; Kruskal 1964
- Fifty-one published similarity matrices, collected by Michael Lee (UC Irvine); every fit was computed in advance with the same scikit-learn call as Assignment 1 - Ekman's colours: stress 0.29 in one dimension, 0.03 in two, and the two-dimensional solution is the colour circle, 434 nm next to 674 nm - Rothkopf's Morse code: one dimension already orders the signals by length (E and T at one end, the digits at the other); the second separates dots from dashes - The stress curve is the tool for choosing dimensions; look for the elbow, not for a threshold; the kinship terms and Kruschke's rectangles show clear elbows, the Romney word lists do not - The dendrogram is the second hypothesis about the same table; the clusters it returns need not match the neighbourhoods of the map
- MDS is fit to the ordinal structure only, so the exponential shape of generalization against distance is an empirical outcome - The constant c is a sensitivity parameter that varies across tasks; the exponential form does not - The law was only ever tested on simple, low-dimensional stimuli; whether it survives real-world stimuli, and why the exponential arises at all, are open questions; Tenenbaum and Griffiths (*Behavioral and Brain Sciences* 2001) give the Bayesian account of the exponential. Shepard, *Science* 1987
- Assignment 1 does this: three fits to the same pairs of MDS distance and rated similarity, each reported with R², the share of the spread in similarity the curve accounts for (1 perfect, 0 no better than the mean) - The line is the null: constant-rate decay, no law worth stating. Any bending curve beats it, so the real comparison is a second bending curve with the same number of parameters - The Gaussian a·exp(−b·d²) decays and flattens like the exponential but is flat at the origin; Shepard's claim is that generalisation is steepest at zero distance - The curves on the slide are analytic with a = b = 1, not fits to assignment data. Shepard 1987 - Preview of Assignment 1: in the psychological space recovered from the ratings, the exponential is the best of the three curves. Against distances in a network layer the same ratings give a different answer, because the x-axis changed: distance in a 4,096-unit layer is bent relative to psychological distance. Shepard's law is a claim about the psychological space
- Two systems, the same stimuli in the same order, two matrices of the same size. Both show blocks in the same places, primates and carnivores clearest, and the network's are weaker. This section puts a number on that resemblance - The units differ: a rating scale on the left, a correlation distance on the right. Whatever number is chosen has to survive that
- The change of level: in the first half a measure compared two response vectors and produced one entry of an RDM; now a measure compares two RDMs and produces one number for a pair of systems - Nothing new is needed: the upper triangle of an RDM is a vector of N(N−1)/2 entries, and Euclidean, cosine and correlation apply to any two vectors of the same length - Rank correlation is the default because the two RDMs have unrelated units. Pearson asks a stricter question, whether the dissimilarities are linearly related, and gives a different number on the same pair of matrices (0.57 against 0.42) - Assignment 1 computes this for the human RDM against a trained and an untrained network
- The same idea underlies non-metric MDS (only the order of the distances matters) and the comparison of two RDMs on the next slides (only the order of the pairs matters); the middle panel is what Spearman actually computes - Warping: two systems whose dissimilarities are related by any rising curve, exp(0.35x) here, have Spearman 1 and Pearson below 1; the swap in the right panel is the kind of disagreement Spearman is sensitive to
- The operation has a name, representational similarity analysis (RSA); model comparison and ranking, Brain-Score, noise-ceiling inference, significance testing and the zoo of alternative measures come in a later lecture - This is the operation that produced the number relating the human RDM and the network RDM over the same 120 photographs - The diagonal is N pairs on which the two matrices agree perfectly by construction, so including it drags rho upward; the lower triangle doubles the apparent sample size, invisible until someone asks how likely the result is by chance - Spearman is used because the matrices share no common scale, and because Shepard's law is why rank information is the trustworthy part of a similarity measurement - Two 120 × 120 matrices in, one number out; Assignment 1 does this operation, the behavioural RDM against each layer of each network. Kriegeskorte, Mur & Bandettini 2008; Kriegeskorte & Kievit 2013
- Two systems can agree on every pair while measuring different things about the images. Two networks trained on different objectives can produce nearly the same RDM over a stimulus set that does not distinguish them; the RDM is a summary, and systems that differ underneath can match it - So "the model matches IT" is a weaker claim than it sounds - The split-half reliability of the human ratings, two slides on, gives Assignment 1 its comparison scale; it is a proxy, not a hard ceiling. Formal inference and model-ranking uncertainty come in the later model–brain lecture. Kriegeskorte & Kievit, *TICS* 2013
- Kriegeskorte, Mur, Ruff et al. (Neuron 2008): 92 colour object photographs, monkey IT single units (Kiani et al. 2007 data) and human IT fMRI voxels, both matrices computed with correlation distance and shown as percentiles - Both matrices show the animate/inanimate division and the face and body sub-blocks; the paper reports a linear correlation of 0.49 between the two matrices, the operation of the previous slides applied to two brains - Figure: Kriegeskorte et al., Neuron 2008, fig. 1, reproduced for teaching with the colour bar moved to the side
- The RDM correlation applied five times, each system against the same human RDM over the same 120 photographs; you can recompute every number from the Assignment 1 data - Raw pixels carry none of the human ordering, each stage of the trained network carries more, and the untrained copy of the best stage carries none, so the alignment is a property of the learned weights - The dashed line is the noise ceiling: split-half reliability, the two rater batches Peterson released correlated with each other, 0.82 on the Spearman scale. The averaged ratings are more reliable than either half, so "half of the ceiling" is an upper bound on what the model has reached - Assignment 1 repeats this with your own RDMs; it also asks a second, Shepard-like question: does the network's similarity fall off with distance in the human space the way human similarity does
- Same operation as the previous slide, five more models; features are the pooled penultimate layer (512 to 2,048 numbers per image). On the same pooled features ResNet-18 reaches 0.50, so the deeper ResNets add 0.02, not the 0.10 the bars suggest against the 4,096 sampled units of Assignment 1 - ResNet-50 and ResNet-152 (ImageNet V2 weights) both reach 0.52; the supervised ViT-B/16 (a vision transformer) and ConvNeXt-Base, more accurate on ImageNet than any ResNet, fall to 0.26 and 0.20 - DINOv2 (ViT-S/14, self-supervised on 142 million images, no labels) reaches 0.61, about three quarters of the 0.82 ceiling; it is the model whose features Assignment 1 also provides - This is the pattern of Linsley et al. 2026: past a point, ImageNet accuracy and human-likeness come apart; the training objective, not the size, is what moves alignment
- Parts 1 and 2 are fully covered as of today; Part 3 needs Tuesday's lecture (PCA) - Each notebook is self-contained, so you can work ahead; trying a part before its lecture is a good way to learn it
- The RDM is the common currency that makes brains, behaviour and models commensurable - MDS, Shepard's law, Tversky's violations, clustering and RSA are operations on, or checks of, the RDM
- None of these is required reading. Shepard 1987 and 1980 are the two classics behind today's lecture; the two Transmitter essays are short and set up two questions of later lectures (how to compare representations, what dimensionality means); the last three are recent studies that use MDS and RDM correlation - The research trail asks for a paper you chose and can explain; if several of you converge on the same one, branch from its references or from the papers that cite it
- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 3) - Question 1 is the one Tuesday's minute paper asked before the lecture reached it - Credit is for a thoughtful attempt, not for being correct; recurring questions are summarized anonymously and addressed on Ed or at the start of the next lecture
- Subtracting vectors gives the displacement from one image to the other: x^(2) − x^(1) = (128.3 − 72.7, 19.2 − 40.6) = (55.6, −21.4) for the tiger, image 1, and the elephant, image 2 - The difference is itself a vector, one entry per feature: drawn from the tiger's point, it ends at the elephant's, and its length is the distance between the two representations
- N enters the output quadratically, as N(N−1)/2 pairs, while the number of channels n does not enter at all: n numbers per image in, one number per pair out - Assignment 1 asks you to write these distances by hand, then checks them against this SciPy call
- All 4,096 stored late-layer (layer4) ResNet-18 activations for the tiger, without thresholding or normalization: 1,843 (45.0%) are exactly zero and the maximum is about 13.57; activations are model values, not firing rates or probabilities - Each ResNet-18 block ends with a rectified linear unit (ReLU) after the residual addition, and the layer4 activations are recorded after it; ReLU maps nonpositive inputs to zero and passes positive inputs unchanged (−2 → 0, 3 → 3), which is why so many entries are exactly zero, and a stored zero does not reveal its preactivation - ReLU is nonlinear because it is not additive: ReLU(−2 + 3) = 1 but ReLU(−2) + ReLU(3) = 3; it is piecewise linear, not one linear function - Nonlinearity is what gives depth its power: a stack of linear maps collapses to a single linear map, W2(W1 x) = (W2 W1) x, and with biases to a single affine map, whereas nonlinear units let successive layers build input-dependent representations. Max pooling is nonlinear too - A zero for one image is not an inactive feature: none of the 4,096 features is zero for all 120 assignment images
- Comparing the tiger's and gorilla's full late-layer vectors on the same 4,096 features, exactly 879 entries are zero in both images. ReLU makes such entries common; firing rates and BOLD values are almost never exactly equal - What goes wrong: correlation subtracts each vector's mean before comparing. Hundreds of shared zeros pull both means down, and every zero then counts as a below-average entry that the two images share, which manufactures agreement: r = 0.103 with all 4,096 entries, −0.011 on the 3,217 entries that are not zero in both. Correlation dissimilarity moves from 0.897 to 1.011. Euclidean distance (125.66) and cosine dissimilarity (0.635) do not move, because a shared zero adds nothing to a squared difference, a dot product or a norm - Correlation depends on which units are included. Dropping "units silent for this pair" gives every pair its own unit set, and the RDM no longer compares like with like - What to do: decide the unit set once, over the whole image set (for instance, drop units that are zero for every image; here there are none, so keep all 4,096), apply that same set to every pair, and state the choice with the result. With ReLU activations, cosine is the gain-free measure that is immune to this particular choice; correlation is fine provided the unit set is fixed. Assignment 1 examines retaining versus excluding inactive units
- The weights come from training: the network was shown labelled photographs and its weights adjusted, step by step, to reduce the classification loss. The RDM is a property of those weights, not of the architecture - The full 120 × 120 RDM (correlation distance on the 4,096 late-layer activations), computed twice: with random weights and with ImageNet-trained weights; images ordered by taxon, so a category is a block on the diagonal - Random weights: every off-diagonal entry is close to 0.15 (range 0.03 to 0.41), within and between categories alike; the untrained network gives almost the same pattern for every image - Trained: the mean distance rises to 0.90, and same-category pairs (0.82) are closer than different-category pairs (0.92), the block structure; amphibians (0.72) and rodents (0.76) are the tightest categories - Each panel has its own colour scale: on the trained panel's scale the random-weights matrix would be one pale square; on its own scale it is noise with no block structure.
- Every network RDM in this lecture was computed with weights that came from training; the RDM is not a property of the architecture alone. Supervised training pairs each image with a category label; the network outputs class probabilities, and a loss scores the prediction against the true label; cross-entropy, the usual classification loss, penalizes low probability on the correct class rather than counting misclassifications - A learning algorithm repeatedly updates the weights on batches of examples to lower the average loss; the outcome is a learned set of weights, not a guaranteed global optimum, and performance is measured on held-out images, images not used in training - Changing the weights changes the internal responses to the same image, and therefore the distances between image representations; the objective concerns labels, not human judgments or neural responses - Once trained, the weights are fixed while responses are extracted; presenting an image to a fixed network does not train it - The diagram is schematic: the four photographs are Assignment 1 images with their labels, the frog is the example currently passing through, and the probability bars are invented, not a measured prediction
- The change of source matters: a network's RDM inherits zero diagonal, symmetry and nonnegativity from the measure that made it (all three of ours satisfy them), and the triangle inequality when the measure is a distance. A human rating is a number a person chose; averaging over ten raters and symmetrising is what makes the Peterson matrix usable, and nothing in the procedure guarantees the triangle inequality - Tversky (1977) showed with data that human judgments violate symmetry, minimality and the triangle inequality; the appendix has the three demonstrations, and the Assignment 1 bonus tests the axioms on our own matrices - Why it matters: the second half of the lecture starts from an RDM and tries to recover the space behind it, one point per image. Points on a page obey all four properties automatically, so a matrix that violates one cannot be reproduced exactly; the mismatch is absorbed by stress
- Trap 1: forgetting `metric="precomputed"` gives a silently wrong answer rather than an error, because scikit-learn then computes Euclidean distances from D as if it were raw data; Assignment 1 hands you the MDS call ready-made for this reason - Trap 2: what `mds.stress_` reports depends on the call; with scikit-learn 1.8 and `normalized_stress="auto"`, the metric call prints raw stress and the nonmetric call prints Stress-1, so compare like with like
- Least squares is the whole method: the fitted curve is the member of the family (all a·exp(−b·d), say) with the smallest sum of squared vertical residuals; curve_fit searches a and b for that minimum, polyfit solves it directly for a line - R² compares that residual sum with the spread around the mean; it can be negative when a curve does worse than the flat line at the mean, and it is not a probability - Assignment 1 fits the raw pairs, all 7,140 of them. Published R² values are often computed on binned means, which flatters any curve; when you compare with a paper, check which was done
- Assignment 1 needs the ceiling operation: agreement with humans requires a reference for how consistently humans agree with one another, and the split-half correlation is that reference - The split-half value is a conservative proxy, not a hard maximum: under independent noise the half-sample denominator is too small, which inflates the reported fraction, hence "upper bound" - Do not turn a rank correlation into variance explained; Assignment 1 uses ResNet-18 and Spearman correlation on correlation distances, while Peterson's paper reports VGG features and a regression R², so the numbers are not comparable - The 0.42 and 0.82 are computed on the Assignment 1 data: the two released rater batches, and ResNet-18's late layer. Peterson et al. 2018
- In a dendrogram the horizontal axis is not meaningful: leaf order is arbitrary up to flipping any join, just as an MDS configuration is arbitrary up to rotation - Height is meaningful, since it is the distance at which two groups merged; the same warning as for MDS axes, in a different guise
- For the Assignment 1 bonus: SciPy's `linkage` expects the condensed form, so pass `squareform(D, checks=False)`, not the square matrix - Passing the square matrix is a classic twenty-minute debugging trap, because it is silently treated as N observations in N dimensions
- Every linkage rule has a knob, and the knob encodes an assumption about what a cluster is: single linkage chains, complete linkage forces compact groups - The same lesson as the choice of dissimilarity measure: the choice decides the answer
- The adjusted Rand index compares two partitions: 1 when they are identical, about 0 for chance agreement. 0.89 means the eight-cluster cut nearly recovers the eight taxa
- Drysdale et al. (*Nature Medicine* 2017) proposed neurophysiological subtypes of depression tied to treatment response, with a standard pipeline; the paper has thousands of citations - Hierarchical clustering has no null hypothesis: 1,000 points from a single Gaussian blob return a clean-looking dendrogram, because merging is all the algorithm can do - "We found four clusters" is not a finding; "the four clusters survive resampling, hold up in a held-out sample, or beat a null with no group structure" is - Dinga et al. (2019) reran the pipeline on the original data and new samples: the cluster structure was not statistically distinguishable from no clusters, and the treatment-prediction advantage did not hold up; later multi-site work was similarly unsupportive - The Assignment 1 bonus meets the same failure; like the number of dimensions m, the number of clusters is a modelling choice. Figure: Drysdale et al., *Nature Medicine* 2017, fig. 1e, Springer Nature, all rights reserved, used with critical commentary in a non-commercial lecture
- The choice of visualisation is itself a hypothesis about the domain, a smooth space for MDS and discrete categories for clustering - Shepard, "Multidimensional scaling, tree-fitting, and clustering", *Science* 1980, makes this argument; optional reading
- Tversky's argument sets the domain in which MDS is valid - Euclidean RDMs satisfy the metric axioms, but cosine and correlation dissimilarities need not, even when computed from response vectors; human judgments add further failures, including asymmetry - Both kinds of dissimilarity are computed this semester, so the distinction matters. Tversky, *Psychological Review* 1977
- The asymmetry is easy to reproduce: rate how similar North Korea is to China, then China to North Korea; the ratings differ - Tversky's explanation is featural, not spatial: similarity is a weighted contrast of common and distinctive features, and the features of the first-named item are weighted more heavily. Rothkopf 1957; Tversky 1977
- An MDS map of fruit names puts "fruit" in the centre, yet it can be the nearest neighbour of only two or three items, because a Euclidean space bounds how many points can share a nearest neighbour - A tree handles a superordinate term naturally and a flat map does not, which is why clustering follows MDS. Tversky 1977; Tversky & Hutchinson 1986
- Not part of Assignment 1; papers you may pick for the research trail use these terms - Bootstrapping over stimuli asks whether the result would survive a different sample of images from the same pool; bootstrapping over subjects asks the same about the people. Nili et al. 2014 (the RSA toolbox) and Schütt et al. 2023 (rsatoolbox, model-comparison inference with both bootstraps) are the standard references - Ceilings: a group RDM predicts each subject's RDM only so well; the upper bound uses the group including the subject, the lower bound excludes them; Nili et al. 2014 - Cross-validated (crossnobis) distances remove the positive bias that noise adds to every distance; Walther et al. 2016. Soni et al. 2024 show that the ranking of models against brain data changes with the similarity measure; the next lecture, on the geometry of neural representations, defines CKA; shape metrics come in the later model–brain lecture - Geometry is evidence about a representation, not a demonstration of shared mechanism; Kriegeskorte & Kievit 2013
- This is the first case with no physical stimulus dimension to fall back on: colour had wavelength, emotion has nothing - The 2-D layout in the figure is t-SNE, not MDS; what transfers is the recipe (judgments, pairwise structure, map), not the solver - The field's default account has two dimensions, valence and arousal; Cowen & Keltner need about 27 categories with continuous boundaries, more categories than the dimensional view wants and fuzzier edges than the basic-emotion view wants - The open weakness is whether the 27th dimension is real or an artefact of how many response options were offered. Cowen & Keltner, *PNAS* 2017, fig. 2A