Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 2 · Theme 1: Representational spaces

MDS: recovering psychological spaces

Thursday, September 17 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Last time, and today

Tuesday: a representation is a list of numbers, a point in a space; pixels, voxels, neurons and network units all give one; averaging over repetitions; the response matrix X\mathbf{X}; the RDM D\mathbf{D}, one comparison per pair of images.

Today, part 1: what the comparison is. Euclidean distance, cosine, correlation; which of them are distances; the RDMs of pixels, people, neurons and a network.

Today, part 2: from the RDM back to a space. The psychological space: each stimulus a point, the distance between two points how dissimilar they seem to a person. Recovered from judgments alone by multidimensional scaling.

Your questions from Tuesday

Four questions came up often enough to answer here:

  • "Where did the 4,096 numbers go when a row became one square of the RDM?": one number per pair, computed from both rows: the first slides today
  • "How different channels can provide different features": in a convolutional network, a channel is a whole sheet of units that share one set of weights, one feature detector applied at every position; the sheet shows where in the image that feature is present. Each unit is one entry of x\mathbf{x}. A recording channel is one electrode, an unrelated use of the word
  • "Are the neural vectors per neuron, averaged over trials?": yes: one mean rate per neuron per image; Assignment 1 gives you those means
  • "Why 0–1 rather than 0–255?": the same 256 levels divided by 255, nothing added or lost. Networks train and run in floating point: weights and activations are real numbers, far finer than 8-bit integers, so inputs are scaled to what the weights were trained on

"The linear algebra is going too fast"

  • It is, and the pace stays: there is a lot to cover. Hardest now, at the start
  • The list is short: vectors, dot products, lengths, matrices, the operations every network is built from. Next week adds the eigenvector; Theme 2 the gradient
  • Use the help. Office hours Wednesday 1 PM, Carney 419: nobody came to talk about linear algebra ☹. TA hours are not fixed yet: vote on an Ed thread and we will set them up; I will not ask the TAs to hold hours nobody attends. Notation reference on Canvas
  • Two handouts, to help: Linear algebra for this course (vectors, dot products, norms, matrices, the three measures, worked numbers) and Programming notes for Assignment 1 (NumPy, the calls, the checks). On the course site and on Canvas
  • Assignment 1 is hard in Part 1 and gets easier. Writing the code yourself is what makes the notation yours

Comparing two responses

From the response matrix to the RDM

Left: the response matrix, 120 rows, one image each, 4,096 columns of unit activations, drawn as a tall strip of values with two rows highlighted; an arrow labelled with the dissimilarity measure d leads to the right: a 120 by 120 matrix with the single entry for that pair of rows highlighted

  • Response matrix X\mathbf{X}: NN rows (images), DD columns (units); here N=120N = 120, D=4,096D = 4{,}096
  • Two rows, x(i)\mathbf{x}^{(i)} and x(k)\mathbf{x}^{(k)}, compared with the measure dd: one number, dikd_{ik}
  • Every pair, stored by row and column: the RDM D\mathbf{D}, N×NN \times N
  • DD numbers per image go in; one number per pair comes out

An RDM compares every pair of images

Representational dissimilarity matrix (RDM): every pairwise comparison.

D∈RN×N,dik=d(x(i),x(k))\mathbf{D}\in\mathbb{R}^{N\times N},\qquad d_{ik}=d\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)

  • Rows and columns: images, here ordered by category. Each entry: one comparison
  • dd: the dissimilarity measure. It decides which differences count
  • What makes it a dissimilarity matrix: zero diagonal, dii=0d_{ii}=0; symmetric, dik=dkid_{ik}=d_{ki}; nonnegative, dik≥0d_{ik}\ge0. A fourth property, the one points in a space obey, comes later today
A 120 by 120 representational dissimilarity matrix of the Assignment 1 images from the trained ResNet-18 late layer, correlation distance, images ordered by animal category with category boundaries drawn; light blocks along the diagonal show that primates, carnivores and hoofed mammals are more alike within their category than across

Euclidean distance

Length of the difference vector:

deuc(x(i),x(k))=∥x(k)−x(i)∥=∑j=1D(xj(k)−xj(i))2d_{\mathrm{euc}}\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\bigl\|\mathbf{x}^{(k)}-\mathbf{x}^{(i)}\bigr\|=\sqrt{\sum_{j=1}^{D}\bigl(x_j^{(k)}-x_j^{(i)}\bigr)^2}

Tiger and elephant, x(2)−x(1)=(55.6, −21.4)\mathbf{x}^{(2)}-\mathbf{x}^{(1)}=(55.6,\,-21.4):

deuc=55.62+21.42=59.5d_{\mathrm{euc}}=\sqrt{55.6^2+21.4^2}=59.5

  • Distance asks how far apart two points are; nothing else
The two-feature plot, mean intensity against contrast, with the tiger's and the elephant's points; the segment between them is labelled with its Euclidean length, 59.5

Same tiger, three brightnesses

  • Scale every pixel by 0.50.5, 11, 1.51.5: the same tiger, darker or brighter. The three points lie on one line through the origin
  • Euclidean distance between the darkest and the brightest: ∥1.5 x(1)−0.5 x(1)∥=∥x(1)∥=83.3\|1.5\,\mathbf{x}^{(1)}-0.5\,\mathbf{x}^{(1)}\|=\|\mathbf{x}^{(1)}\|=83.3, more than tiger to elephant (59.559.5)
  • What the three share is their direction; what differs is their length. Direction = the pattern. Length = overall strength, the gain (brightness here; firing rate or activation elsewhere)
  • A measure of the pattern should compare directions and ignore the gain
The feature-space plot with the tiger at half, normal and one-and-a-half brightness, three thumbnails whose points lie on one dotted line through the origin, the normal tiger marked by the blue vector x superscript 1

The dot product measures alignment

For two samples, x(1)\mathbf{x}^{(1)} (tiger) and x(2)\mathbf{x}^{(2)} (elephant):

x(1)⊤x(2)=∑j=1Dxj(1)xj(2)=∥x(1)∥ ∥x(2)∥cos⁡ϕ\mathbf{x}^{(1)\top}\mathbf{x}^{(2)}=\sum_{j=1}^{D}x^{(1)}_j x^{(2)}_j=\|\mathbf{x}^{(1)}\|\,\|\mathbf{x}^{(2)}\|\cos\phi

  • ϕ\phi is the angle between the two vectors: 0°0° when they point the same way, 90°90° when perpendicular
  • The dot product grows with both lengths and with the alignment cos⁡ϕ\cos\phi
  • Tiger and elephant: x(1)⊤x(2)=10,104=83.3×129.7×cos⁡ϕ\mathbf{x}^{(1)\top}\mathbf{x}^{(2)}=10{,}104 = 83.3\times129.7\times\cos\phi, so cos⁡ϕ=0.94\cos\phi=0.94, ϕ=21°\phi=21°
  • ∥x(2)∥cos⁡ϕ\|\mathbf{x}^{(2)}\|\cos\phi is the projection of x(2)\mathbf{x}^{(2)} onto the direction of x(1)\mathbf{x}^{(1)}, how far x(2)\mathbf{x}^{(2)} reaches along it: 121.4121.4
The feature-space plot with the tiger vector x superscript 1 and the elephant vector x superscript 2, the angle phi between them marked at the origin, and the projection of the elephant vector onto the tiger's direction shown as a magenta segment with a dashed perpendicular

Same tiger, plus a constant

  • Add 3030, then 6060 to every pixel (the 784 numbers of the 28 × 28 image): a washed-out tiger. Euclidean distance 840840 and 1,6801{,}680 (28×3028\times30, 28×6028\times60: this is the 784-pixel space); cos⁡ϕ=0.991\cos\phi = 0.991, 0.9780.978. Both report a change; the pattern did not change
  • Subtract each image's own mean from its pixels: the three become the same image, entry for entry
  • Cosine of the centred vectors is 11 for every pair: that is correlation, next slide
  • A measure of the pattern should ignore a gain (last slide) and an offset (this one): a brighter image, a shift in baseline firing rate
Top row: the tiger as a 28 by 28 grey image, then with 30 and with 60 added to every pixel, progressively washed out, means 73, 103 and 133. Bottom row: each image minus its own mean, three identical images with mean 0

Two similarity measures: cosine and correlation

scos⁡(x(i),x(k))=cos⁡ϕ=x(i)⊤x(k)∥x(i)∥ ∥x(k)∥r(x(i),x(k))=xc(i)⊤xc(k)∥xc(i)∥ ∥xc(k)∥,xc=x−xˉ 1s_{\cos}\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\cos\phi=\frac{\mathbf{x}^{(i)\top}\mathbf{x}^{(k)}}{\|\mathbf{x}^{(i)}\|\,\|\mathbf{x}^{(k)}\|}\qquad\qquad r\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\frac{\mathbf{x}_c^{(i)\top}\mathbf{x}_c^{(k)}}{\|\mathbf{x}_c^{(i)}\|\,\|\mathbf{x}_c^{(k)}\|},\quad \mathbf{x}_c=\mathbf{x}-\bar{x}\,\mathbf{1}

  • Cosine similarity: the dot product with both lengths divided out; only cos⁡ϕ\cos\phi remains. In [−1,1][-1,1]: 11 same direction, 00 perpendicular. The three tigers: 11
  • Correlation (Pearson rr): cosine similarity after subtracting each pattern's own mean xˉ\bar{x}
  • Both in [−1,1][-1,1], large when alike. Dissimilarity is the flip: dcos⁡=1−scos⁡d_{\cos}=1-s_{\cos}, dcorr=1−rd_{\mathrm{corr}}=1-r, in [0,2][0,2]

What each measure ignores

Three bar charts of a twelve-unit response pattern: the pattern itself, every unit doubled, and a constant added to every unit; under each, marks show which measures call it the same as the original: Euclidean distance none, cosine similarity only the doubled pattern, correlation both

  • Euclidean distance keeps gain and offset · cosine discards the gain · correlation discards both. Choose by what should not count
Assignment 1 §1b: the three functions, then all three matrices

Which is closer to the tiger? It depends on the measure

Two features (plot): eagle or penguin?

Pair Euclidean Cosine* Correlation*
tiger–eagle 23.8 0.021 0†
tiger–penguin 54.1 0.009 0†

4,096 network responses: penguin or gorilla?

Pair Euclidean Cosine* Correlation*
tiger–penguin 122.8 0.649 0.911
tiger–gorilla 125.7 0.635 0.897

* dissimilarities, 1−cos⁡ϕ1-\cos\phi and 1−r1-r · † with two components, r=±1r=\pm1 always: correlation needs more than two

The two-feature plot with three arrows from the origin: the tiger x superscript 1, the eagle x superscript 2, and the penguin x superscript 3. The eagle's point is nearer the tiger's, but the penguin's arrow is more nearly parallel to the tiger's

Which measure do papers actually use? All of them

Horizontal bar chart over 192 published RSA papers from 2021 to 2026: correlation distance 49 percent, Euclidean 38, Mahalanobis or crossnobis 8, cosine 12, measure not stated 20 percent; the title says 24 percent use two or more measures and that RDMs are compared with Pearson in 48 percent, Spearman in 40, Kendall in 7

  • Correlation distance first, Euclidean next, cosine and a few others behind. The field uses them interchangeably, and one paper in five never says which one it used
  • They are not interchangeable: the last slides showed each one ignoring something different. That is why you build all three in Assignment 1 and say which you used, and why
  • RDMs are compared with two different correlations about equally often; the assignment uses the rank-based one, the end of today
192 open-access papers that compute an RDM, 2021–2026, Europe PMC full text; the measure counted automatically · an estimate

Six RDMs, one format

Six dissimilarity matrices in two rows, images ordered by category with boundaries drawn, each on its own colour scale. Top row, the 120 animal photographs: pixels show no block structure, human judgments show strong light blocks on the diagonal, one per category, the ResNet-18 late layer shows partial blocks. Bottom row, the 1,224 object images of Bao and colleagues: pixels show faint structure, monkey IT neurons show clear blocks for faces and animals, the ResNet-18 late layer shows a strong face block
  • Two stimulus sets, five kinds of measurement, one format: an N×NN\times N matrix, the stimuli on both axes, one number per pair
  • Light blocks on the diagonal: the categories a system treats as alike. None in pixels; strong in people and in IT neurons; partial in the network
  • Units differ (a rating, a correlation distance); the format does not. The shared format is what lets two systems be compared
Assignment 1 data · people: Peterson, Abbott & Griffiths Cognitive Science 2018 · monkey IT: Bao et al. Nature 2020 · each panel on its own colour scale

From the matrix back to a space

From the matrix to a map

Left: the 6 by 6 human dissimilarity matrix of a tiger, a gorilla, a monkey, an eagle, a penguin and a frog, entries printed. Right: the two-dimensional map that MDS returns, a dot per animal with its thumbnail beside it, three map distances written on their segments: tiger to monkey 5.57, gorilla to monkey 1.15, eagle to penguin 3.72

  • So far: responses in, a matrix out. Now the reverse: a matrix in, points out. Six animals rated by people (dikd_{ik}: 0 identical, 10 nothing in common) to six points on a plane whose distances d^ik\hat d_{ik} come as close as possible
  • Gorilla–monkey rated 1.81.8, 1.151.15 on the map; eagle–penguin 5.25.2, 3.73.7. No plane reproduces all fifteen numbers at once; how close one can get is measured later today
Assignment 1 data, ratings from Peterson, Abbott & Griffiths Cognitive Science 2018 · metric MDS, m = 2

Human similarity judgments

A long tradition in psychology: what we know of objects is arranged as points in a space, near when alike. That psychological space cannot be measured directly. What can be measured is a judgment. Peterson and colleagues asked people: how similar are these two images, from 0 to 10? Ten raters for every pair of 120 animal photographs.

Two tiger photographs A and B and a monkey photograph C, from the Assignment 1 image set

Mean rating   A–B: 10.0  ·  A–C: 2.9  ·  B–C: 2.2  ·  one number per pair: a dissimilarity matrix, from which the space is to be recovered

Ratings: Peterson, Abbott & Griffiths Cognitive Science 2018, the data used in Assignment 1

Which measures are distances in the full sense?

Will it work? Yes, if the matrix is, or is well approximated by, a table of distances between points in that space. What a table of distances must satisfy, and which of our measures satisfy it:

Requirement Mathematical statement Euclidean deucd_{\mathrm{euc}} Cosine dcos⁡d_{\cos} Correlation dcorrd_{\mathrm{corr}}
Nonnegativity · no negative distances dik≥0d_{ik}\ge0 ✓ Yes ✓ Yes ✓ Yes
Identity · zero only for an image and itself dik=0  ⟺  x(i)=x(k)d_{ik}=0\iff\mathbf{x}^{(i)}=\mathbf{x}^{(k)} ✓ Yes ✗ No ✗ No
Symmetry · same both ways dik=dkid_{ik}=d_{ki} ✓ Yes ✓ Yes ✓ Yes
Triangle · no shortcut via a third image dil≤dik+dkld_{il}\le d_{ik}+d_{kl} ✓ Yes ✗ No ✗ No
A distance in the full sense (a metric)? All four ✓ Yes ✗ No ✗ No

✓ holds throughout the domain; ✗ can fail even when the measure is defined. On our images: 2.3 % of pixel triples break the triangle inequality under cosine, none in the trained late layer, and a brighter copy of an image sits at cosine distance 0 from it. You test this yourself: Assignment 1 bonus B1 · worked examples in the appendix.

Reading the map: the notation

The six-animal dissimilarity matrix on the left with three entries outlined, tiger–monkey, gorilla–monkey and eagle–penguin, each joined by an arrow to the segment between the corresponding two points on the map on the right, where every point is labelled y with its stimulus number and each segment is labelled d-hat with the pair's indices

  • Stimulus ii becomes a point y(i)\mathbf{y}^{(i)}; the entry dikd_{ik} becomes the distance d^ik=∥y(i)−y(k)∥\hat d_{ik}=\|\mathbf{y}^{(i)}-\mathbf{y}^{(k)}\|. The hat: a quantity of the map, not a measurement
  • Three of the fifteen entries and the three distances that stand for them. A good map makes every d^ik\hat d_{ik} close to its dikd_{ik}

Multidimensional scaling: the problem

Given a dissimilarity matrix D\textcolor{#3F7A6B}{\mathbf{D}} over NN stimuli, with entries dik\textcolor{#3F7A6B}{d_{ik}}.

Find coordinates, one point per stimulus,

Y∈RN×m,y(i)∈Rm,    m≪N\textcolor{#B5396B}{\mathbf{Y}} \in \mathbb{R}^{N\times m}, \qquad \textcolor{#B5396B}{\mathbf{y}^{(i)}} \in \mathbb{R}^{m}, \;\; m \ll N

such that the map distances d^ik\hat d_{ik} reproduce the measured dissimilarities dikd_{ik}:

d^ik  =  ∥y(i)−y(k)∥  ≈  dikfor all i<k\hat d_{ik} \;=\; \big\lVert \textcolor{#B5396B}{\mathbf{y}^{(i)}} - \textcolor{#B5396B}{\mathbf{y}^{(k)}} \big\rVert \;\approx\; \textcolor{#3F7A6B}{d_{ik}} \quad \text{for all } i<k

  • Usually m=2m=2 or 33, to look at it
  • Not required: X\mathbf{X}, the channels, a model of the stimuli. D\mathbf{D} is the only input.
  • When D\mathbf{D} comes from judgments, Y\mathbf{Y} is the psychological space: each stimulus a point, judged dissimilarity the distance between points
  • Solved by lowering a mismatch, one small step at a time: next slides

Worked example: the colour circle falls out

Left: Ekman's 14 by 14 matrix of judged dissimilarity between monochromatic lights from 434 to 674 nanometres, with a wavelength colour strip along each edge. Right: the two-dimensional non-metric MDS solution, each light drawn in its own colour: a circle, violet next to red

  • Input: people judging the similarity of 14 monochromatic lights, 434–674 nm. A table of numbers; the algorithm is never told the wavelengths
  • Output: a circle, ordered by wavelength, violet next to red: Newton's colour circle (1704), recovered from judgments alone
Data: Ekman J. Psychol. 1954, via Michael Lee's collection · MDS, m = 2 · analysis after Shepard Psychometrika 1962

Stress-based MDS — the optimization view

Real dissimilarities are noisy; often no exact Euclidean configuration exists. Minimize the mismatch instead:

Stress(Y)  =  ∑i<k(d^ik−dik)2,d^ik=∥y(i)−y(k)∥\mathrm{Stress}(\textcolor{#B5396B}{\mathbf{Y}}) \;=\; \sum_{i<k}\Big( \hat d_{ik} - \textcolor{#3F7A6B}{d_{ik}} \Big)^{2}, \qquad \hat d_{ik}=\big\lVert \textcolor{#B5396B}{\mathbf{y}^{(i)}} - \textcolor{#B5396B}{\mathbf{y}^{(k)}} \big\rVert

  • Stress: zero for a perfect map, growing with every mismatched pair
  • Stress depends on the whole layout Y\mathbf{Y}; the map is whatever makes it smallest
  • Start from any layout, move each point a little in the direction that lowers the stress, repeat until it stops falling: what sklearn does
  • How the direction is found: Theme 2. For now the solver is a black box
  • Local minima: different starts give different maps, so run several and keep the lowest stress (n_init)
Kruskal Psychometrika 1964 · usually reported in normalized form, as Kruskal stress-1

Watch it move

The six animals placed at random on the plane, every pair joined by a thin line; the title reads random start, stress 217
  • Start anywhere: six random points. The stress, the total squared mismatch over the fifteen pairs, is 217217
  • One round: every point moves a little toward where its distances to the other five would match the ratings
  • Repeat until the stress stops falling
Metric MDS (scikit-learn), six animals, human ratings · same axes on all three slides

Watch it move

The same six animals after ten rounds: the two primates have moved together, the two birds too, the frog has separated; stress 64
  • After ten rounds: stress 217→64217 \to 64. The primates have found each other, so have the birds
  • Most of the drop happens in the first few rounds; the rest is fine adjustment
Metric MDS (scikit-learn), six animals, human ratings · same axes on all three slides

Watch it move

The six animals at convergence, the map MDS returns for the six-animal ratings: gorilla and monkey close together, eagle and penguin close together; stress 22
  • Converged: stress 2222, and no move lowers it further. This is the six-animal map from the start of this section
  • A different random start ends in the same map, possibly rotated or mirrored, or in a slightly worse one: run several starts, keep the lowest stress
Metric MDS (scikit-learn), six animals, human ratings · same axes on all three slides

Normalized stress: one number for the fit

Stress ⁣ ⁣− ⁣ ⁣1=∑i<k(d^ik−dik)2∑i<kd^ik 2\mathrm{Stress\!\! -\!\!1}=\sqrt{\frac{\sum_{i<k}(\hat d_{ik}-d_{ik})^2}{\sum_{i<k}\hat d_{ik}^{\,2}}}

  • The mismatch, as a fraction of the map's own spread. 00: every distance matches. 0.30.3: mismatches about a third the size of the distances
  • Raw stress depends on the units of D\mathbf{D}; Stress-1 does not, so fits can be compared across data sets and dimensions
  • sklearn can print it. Rule of thumb (Kruskal 1964): 0.05 excellent · 0.10 fair · 0.20 poor. Assignment 1 §2e adds a real check: the stress of the same fit on a shuffled matrix, which has no structure to recover
  • The assignment gives you the call and takes you through it one step at a time

Reading an MDS solution — three warnings

  • The solution is not unique. Shift, rotate or reflect the configuration: pairwise distances unchanged, stress unchanged. The axes mean nothing. Only relative positions do
  • The dimension mm is your choice. Plot stress against mm: it always falls; look for the elbow. Two dimensions is a display convenience
  • Low stress ≠ true. Enough dimensions fit anything, noise included. Report normalized stress (Stress-1) and the number of dimensions

If a story about your map depends on which way is "up", it is a story about the plot.

Metric → non-metric: keep only the order

Two scatter plots of rated similarity, 0 to 10, against distance in the recovered map for the 120 animal ratings. Left, metric MDS: the fitted relation is a straight falling line, stress-1 0.31. Right, non-metric MDS: the fitted relation is a falling step curve, stress-1 0.23; the points bend because ratings pile up near 0

  • Metric MDS demands d^ik≈dik\hat d_{ik} \approx d_{ik}, with dik=10−d_{ik} = 10 - rating: a pair rated 2 must sit twice as far as a pair rated 6. People do not rate that way; most pairs pile up at 0–1
  • Non-metric MDS demands only the order: if dik<di′k′d_{ik} < d_{i'k'}, put i,ki,k closer than i′,k′i',k'. Any falling curve through the points is allowed, fitted along with the map
  • Same data: stress-1 0.31→0.230.31 \to 0.23. A 0–10 rating is an ordering, not a measurement

Explore: fifty classic similarity datasets

Data: Michael Lee's similarity-data collection (Ekman 1954, Rothkopf 1957, Rosenberg & Kim 1975, Romney et al. 1993, …) · open full screen

Shepard's universal law of generalization

Twelve small panels of generalization plotted against distance in psychological space — sizes, hues, phonemes, Morse code, human and pigeon data — each tracing the same exponential decay

  • Recover the psychological space by non-metric MDS (Shepard 1962, Kruskal 1964) from confusions or ratings; then plot measured generalization against distance in that space. Twelve datasets, one curve:

sik  =  e−c dik\textcolor{#3F7A6B}{s_{ik}} \;=\; e^{-\textcolor{#B5396B}{c}\, \textcolor{#3E6D8E}{d_{ik}}}

  • Sizes, lightnesses, spectral hues, vowels, consonants, Morse code, free-form shapes. Humans and pigeons.
  • Similarity is an exponential function of psychological distance, and the law is not circular: non-metric MDS fits rank order only, so the exponential shape was never built in.

Testing the law: line, exponential, Gaussian

Assignment 1 §2g fits three curves to the same (distance, similarity) pairs:

line: s=m d+cexponential: s=a e−bdGaussian: s=a e−bd2\text{line: } s = m\,d + c \qquad\quad \text{exponential: } s = a\,e^{-bd} \qquad\quad \text{Gaussian: } s = a\,e^{-bd^2}

  • The line is the baseline: constant-rate decay, no law to state
  • Any bend beats a line, so the fair rival is the Gaussian: same two parameters, but flat at d=0d=0 where the exponential is steepest
  • The test is at the origin: which shape do the pairs follow there?
  • Preview: in the space recovered from people the exponential wins; against a network's distances the picture changes, because a network's distance is a different ruler from psychological distance
Exponential and Gaussian similarity curves with the same value at zero but different initial slopes
Shepard Science 1987 · Assignment 1 §2g

Comparing two RDMs

Two matrices, one question

Two 120 by 120 dissimilarity matrices of the same animals, ordered by category with category boundaries drawn: human judgments on the left with strong light blocks along the diagonal, the ResNet-18 late layer on the right with weaker blocks in the same places

  • The same 120 animals, judged by people and by a network. How alike are these two matrices?
Assignment 1 data · ratings Peterson, Abbott & Griffiths Cognitive Science 2018 · ResNet-18 late layer, correlation distance · each on its own colour scale

An RDM is itself a vector

Two 120 by 120 dissimilarity matrices of the same animals side by side, human judgments and the ResNet-18 late layer, both with light blocks along the diagonal, weaker for the network
  • Each upper triangle laid out as a list: N(N−1)/2N(N-1)/2 numbers, one per pair. For N=120N = 120: 7,140 per system
  • Two lists of the same length: every measure from the first half of the lecture applies, one level up. Euclidean, cosine, correlation
  • No shared scale (ratings on 0–10; correlation distances on 0–2), so compare ranks: which pairs each system calls close
Numbers: Assignment 1 data · ResNet-18 late layer, correlation-distance RDM · human RDM D = 10 − S

Spearman's ρ: agree on the order, not the values

Three scatter plots of eight pairs' dissimilarities in two systems. Left: the second system's values are a bent but rising function of the first's, Pearson 0.95, Spearman 1. Middle: the same eight pairs plotted as ranks, a perfect straight line. Right: two neighbouring pairs swapped, Pearson 0.94, Spearman 0.98

  • Spearman's ρ\rho: replace every value by its rank (1st smallest, 2nd, …), then take Pearson's rr of the ranks
  • A system can warp the values, stretch the large ones, compress the small, and still order every pair the same way: ρ=1\rho = 1 while r<1r < 1
  • Only a change of order lowers ρ\rho. Non-metric MDS asks for the same thing of a map: the order of the distances
Toy values

Correlating two RDMs

Two systems, the same NN stimuli, two matrices: Dhuman\textcolor{#3F7A6B}{\mathbf{D}^{\text{human}}} from the judgments, Dmodel\textcolor{#3F7A6B}{\mathbf{D}^{\text{model}}} from a layer. Flatten each upper triangle into a list of N(N−1)/2N(N-1)/2 numbers and correlate the two lists:

ρ  =  corrrank(triu(Dhuman),  triu(Dmodel))\rho \;=\; \mathrm{corr}_{\text{rank}}\Big(\mathrm{triu}\big(\textcolor{#3F7A6B}{\mathbf{D}^{\text{human}}}\big),\; \mathrm{triu}\big(\textcolor{#3F7A6B}{\mathbf{D}^{\text{model}}}\big)\Big)

In practice: iu = np.triu_indices(N, k=1), then spearmanr(D1[iu], D2[iu]).

The operation has a name: representational similarity analysis, RSA

  • Upper triangle only. The diagonal is NN zeros in both matrices, agreement you did not measure; the lower half copies the upper. Either one inflates agreement or the sample size
  • Rank correlation (Spearman): the two systems share no scale, a distance between voxel patterns and one between unit activations being different kinds of number. The raw values are not comparable; their order is
  • The two matrices of a moment ago: ρ=0.42\rho = 0.42. Pearson rr on the same lists gives 0.570.57, a different question

What that number does and does not say

  • The two systems order the pairs alike: what one treats as close, the other tends to treat as close. A claim about geometry, and only that
  • ρ\rho alone means little. It needs how well people agree with each other (the noise ceiling, two slides on) and a second model: something to be large relative to
  • Assignment 1 reports Spearman ρ\rho between the human matrix and ResNet-18's correlation distances: expect about 0.4. Whether that is good: two slides on
  • Matching geometry is evidence about a representation, not a demonstration of shared mechanism.

Why anyone cares: two species, one structure

Two 92-by-92 dissimilarity matrices, monkey IT from single-unit recordings and human IT from fMRI, both showing the same blue animate block and warm inanimate block with face and body sub-blocks; a vertical colour bar on the right runs from blue, similar, to red, dissimilar

  • Same 92 images; a monkey (IT single units) and a human (IT fMRI voxels): two measurements, two 92×9292\times92 matrices, one block structure
  • r = 0.49 between the two matrices (Pearson, as the paper reports it), computed as on the previous slides: the upper triangles, correlated
Kriegeskorte et al. Neuron 2008, fig. 1 · colour: percentile of 1 − r

Which system is most human-like? One wins

Bar chart of Spearman correlation between the human RDM and five candidate RDMs on the 120 animal images: pixels 0.01, early layer 0.10, middle layer 0.17, late layer 0.42, late layer with random weights 0.03; a dashed line at 0.82 marks the agreement between two independent groups of raters

  • Pixels: 0.01. Nothing in the raw image orders the pairs the way people do
  • Depth helps: early 0.10 → middle 0.17 → late 0.42. The late layer wins
  • Same architecture, random weights: 0.03. The structure came from training, not from the wiring
  • Two independent rater groups agree at 0.82, the noise ceiling: the late layer reaches about half of it
Assignment 1 §2i numbers · pixels not in the assignment · ratings Peterson, Abbott & Griffiths Cognitive Science 2018

Do bigger models get closer? Some do

Bar chart of Spearman correlation with the human RDM for five models using pooled penultimate features: ResNet-50 0.52, ResNet-152 0.52, supervised ViT-B/16 0.26, ConvNeXt-B 0.20, self-supervised DINOv2 ViT-S 0.61 in green; dashed ceiling at 0.82

  • Deeper ResNets: 0.52. Depth alone buys little
  • Two newer, more accurate ImageNet models: 0.26 and 0.20. Bigger is not closer
  • A model trained without labels (DINOv2): 0.61, three quarters of the ceiling. How a model was trained matters more than how big it is
  • Transformers, self-supervised learning, why accuracy and human-likeness come apart: later in the course
Same 120 images and human RDM · pooled penultimate layer, correlation distance, Spearman · torchvision; DINOv2 ViT-S/14

Assignment 1: where you stand

  • It is long. Start now if you have not; it is due Tuesday, September 29
  • After today you have everything for Parts 1 and 2 (distances, RDMs, MDS, Shepard's law, the RDM comparison). Part 3 is Tuesday's lecture (PCA); the bonus has no lecture of its own
  • The assignment is written to be self-sufficient: every part explains what it needs. Starting a part before its lecture is allowed and often useful; the lecture then lands on something you have already tried
  • Bonus B2 (hierarchical clustering) is not taught in class: the clustering slides in the appendix are what you need
  • Research-trail entry 1: Friday, September 18

Recap

  • A dissimilarity matrix turns any responses (neurons, voxels, model units, judgments) into one N×NN\times N object, whatever the measurement system
  • MDS turns that matrix into a map by lowering the stress. Axes mean nothing; relative positions do
  • Shepard's law: similarity falls off exponentially with distance in psychological space
  • People's judgments need not obey the triangle inequality (appendix). Treat the map as a useful approximation
  • The same matrix can be read as a tree of nested categories instead of a flat space (clustering, in the appendix)
  • Representational similarity analysis: correlate the upper triangles of two such matrices; one number for how far two systems agree about what resembles what

Further reading for the research trail

Seven seeds, none required. Branch from any of them; pick a paper your neighbour has not picked.

Minute paper

Canvas submission: Canvas → Minute papers → Minute Paper 3 (access code read out in class)

Write three brief points in your own words:

  1. One sentence: when can two response vectors have cosine similarity 1 (cosine dissimilarity 0) and yet a large Euclidean distance?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Appendix · A difference defines a direction

Subtraction: the displacement from image 1 (tiger) to image 2 (elephant):

x(2)−x(1)=(128.3−72.719.2−40.6)=(55.6−21.4)\mathbf{x}^{(2)}-\mathbf{x}^{(1)}=\begin{pmatrix}128.3-72.7\\19.2-40.6\end{pmatrix}=\begin{pmatrix}55.6\\-21.4\end{pmatrix}

x(2)−x(1)\mathbf{x}^{(2)}-\mathbf{x}^{(1)} is itself a vector: from the tiger's point to the elephant's.

The two-feature plot, mean intensity against contrast, with three arrows: x superscript 1 to the tiger's point, x superscript 2 to the elephant's point, and their difference drawn from the tiger's point to the elephant's

Appendix · Computing an RDM from a response matrix

import numpy as np
from scipy.spatial.distance import pdist, squareform

R = np.load("responses.npy")     # (N stimuli, n channels)
print(R.shape)                   # (120, 4096) units  or  (120, 482) neurons

# pdist returns the upper triangle: N(N-1)/2 numbers, the pairs i < k
d_vec = pdist(R, metric="correlation")   # 1 - Pearson r, pattern to pattern
D     = squareform(d_vec)                # (N, N): symmetric, zero diagonal

assert np.allclose(D, D.T) and np.allclose(np.diag(D), 0)
  • pdist computes only the upper triangle, never the N×NN\times N matrix. squareform converts either way
  • Swap metric="euclidean" (or "cosine") to change the geometry; nothing downstream changes

Appendix · Why network responses contain zeros

ReLU (rectified linear unit): input zjz_j, output aj=max⁡(0,zj)a_j=\max(0,z_j).

The tiger image with all 4096 stored late ResNet-18 activations, 1843 exact zeros marked below the baseline, and the ReLU response curve
  • Units are selective: one tuned to headlights or mirrors is silent on 120 animals. Zeros are common: the code is sparse

Appendix · Shared zeros can change correlation

The tiger and gorilla photographs beside their full measured ResNet-18 activation vectors, with 879 shared-zero entries marked at matching positions
  • All 4,096 units: correlation dissimilarity 0.897. Without the 879 units silent for both: 1.011. Euclidean, cosine: unchanged
  • Issue: shared zeros pull both means down and manufacture agreement (rr = 0.10 with, −0.01 without)
  • Rule: choose the unit set once, over the whole image set, apply it to every pair, report it

Appendix · What training does to the RDM

Two 120-by-120 correlation-distance RDMs of the late ResNet-18 layer, images ordered by category with dark lines at the category boundaries, each on its own colour scale: with random weights the entries are noise with no block structure; with ImageNet-trained weights lighter blocks sit on the diagonal, one per category
  • Same 120 images, same architecture, two sets of weights
  • Random weights: every pair alike, mean 1−r=0.151-r = 0.15, the same within and between categories. No blocks, whatever the colour scale
  • Trained: pairs far apart, mean 0.900.90; within a category 0.820.82, between 0.920.92. The blocks on the diagonal are the categories
  • The RDM is a property of the weights, not of the architecture
Correlation distance on the 4,096 stored late-layer activations · random weights, seed 1291 · each panel on its own colour scale

Appendix · Training changes the representation

  • Every RDM so far came from a trained network
  • Weights: the adjustable numbers inside the network. Training: find weights that reduce the classification loss on labeled images
Supervised training: four labelled training photographs (frog, eagle, elephant, penguin) pass one at a time through a network to produce class predictions; comparing the prediction with the true label gives a loss, which guides weight updates and changes internal representations

Appendix · What should a dissimilarity satisfy?

Similarity ratings come from people, not from a formula. A network's RDM inherits its properties from the measure dd; a rater is free to answer anything. Before we recover a space from such numbers, what must they satisfy? For tiger, penguin, gorilla (i,k,li, k, l):

  • The tiger against itself: no difference. dii=0d_{ii}=0 (zero diagonal)
  • Tiger against penguin, penguin against tiger: the same number. dik=dkid_{ik}=d_{ki} (symmetric). A rater need not agree (Tversky 1977: a camel is judged more like a horse than a horse like a camel)
  • No difference below zero. dik≥0d_{ik}\ge0 (nonnegative)
  • Tiger to gorilla is never more than tiger to penguin plus penguin to gorilla: dil≤dik+dkld_{il}\le d_{ik}+d_{kl} (triangle inequality). Nothing stops a rater from breaking it

The first three make D\mathbf{D} a dissimilarity matrix. The fourth is what points in a space obey: without it, no map reproduces the entries exactly. Which measures satisfy which: the distance table in the lecture. Whether people do: the assignment bonus, and the Tversky slides later in this appendix

Appendix · Running MDS in three lines

from sklearn.manifold import MDS

mds = MDS(n_components=2,             # m = 2, because we want to look at it
          metric="precomputed",          # feed D directly, NOT the raw R
          n_init=8, normalized_stress="auto", random_state=0)
Y = mds.fit_transform(D)       # (N, 2) coordinates - one row per stimulus
print(Y.shape, mds.stress_)    # residual mismatch: lower is better

# non-metric (ordinal) MDS - fits the RANK ORDER of the dissimilarities
mds_nm = MDS(n_components=2, metric="precomputed", metric_mds=False,
             n_init=8, normalized_stress="auto", random_state=0)
Y_nm = mds_nm.fit_transform(D)
  • metric="precomputed" (dissimilarity= before scikit-learn 1.8); without it sklearn treats D as raw data and computes Euclidean distances from it
  • D must be the square N×NN \times N matrix: squareform(d_vec) first if you kept the pdist vector
  • n_init=8: stress-based MDS has local minima, so restart and keep the best

Appendix · Fitting a curve is a regression

  • Each point is one pair of images: its map distance dd and its rated similarity ss
  • Fit = choose the curve's parameters (aa, bb) so the curve passes as close as possible to the points: smallest sum of squared residuals s−s^s-\hat s. Same idea as fitting a line, the curve just bends
  • R2R^2 = the share of the spread in ss the curve accounts for: 11 perfect, 00 no better than the flat mean sˉ\bar s

R2=1−∑(s−s^)2∑(s−sˉ)2R^2=1-\frac{\sum(s-\hat s)^2}{\sum(s-\bar s)^2}

  • Fit the three curves to the same points, compare their R2R^2; the highest wins, the line is the baseline

Scatter of forty distance–similarity pairs with a fitted exponential curve through them, a vertical residual segment from every point to the curve, a dotted horizontal line at the mean similarity, and R squared 0.90 in the title

Illustrative points, not assignment data · Assignment 1 §2g: np.polyfit for the line, curve_fit for the two curves

Appendix · Human reliability sets the comparison scale

Split-half reliability: correlate two independent sets of ratings for the same pairs. That agreement is the noise ceiling: how well the measurement predicts itself, the bar no model should beat.

Comparison Statistic
Rating half 1 versus rating half 2 Reliability of the human measurement
Network distances versus human dissimilarities Representational alignment
  • Assignment 1 uses Spearman rank correlation for both quantities. Pearson rr, Spearman ρ\rho and regression R2R^2 are different statistics
  • Peterson's two rater batches agree at Spearman ρ=0.82\rho = 0.82 on our 120 images. Report a model as a fraction of the ceiling: late layer 0.420.42 / ceiling 0.82≈0.50.82 \approx 0.5
  • The averaged ratings are more reliable than either half, so that fraction is an upper bound on what the model has reached

Appendix · Hierarchical clustering — the idea

Left: the two-dimensional MDS map of six animals, tiger, gorilla, monkey, eagle, penguin, frog, thumbnails beside the points. Right: the dendrogram built from the same six-by-six human dissimilarity matrix with average linkage, thumbnails as leaves; gorilla and monkey join at 1.8, eagle and penguin at 5.2, the frog joins last

  • Group the stimuli recursively and draw it as a dendrogram: leaves are stimuli, the height of a join is the dissimilarity at which two groups merged
  • Agglomerative (bottom-up): start with NN singletons, merge the closest pair, repeat. What the Assignment 1 bonus uses
  • Input is again just D\mathbf{D}: no coordinates, so judged similarity works as well as firing rates
The six animals of the MDS slides, human ratings · average linkage

Appendix · Agglomerative clustering and linkage

  1. Compute all pairwise distances: the matrix D\mathbf{D}.
  2. Merge the two closest clusters.
  3. Update the distances from the new cluster to all others, using a linkage rule.
  4. Repeat until one cluster remains.

In practice: scipy.cluster.hierarchy.linkage(squareform(D), method="average").

Step 3 is the whole design decision:

  • Single: distance to the closest member. Chains.
  • Complete: distance to the farthest member. Compact, equal-size clusters.
  • Average: mean over all cross-pairs. The usual compromise.
  • Ward: merge the pair that increases within-cluster variance least. Needs coordinates, not just D\mathbf{D}.

Appendix · Linkage changes the answer

Three dendrograms of the same 120 animals from the human dissimilarity matrix, single, complete and average linkage, each leaf marked with its taxon colour: single linkage grows one long chain, complete linkage makes compact groups, average linkage sits between

  • Same 120 animals, same human ratings, three linkage rules: three different trees. The colours under the leaves are the eight taxa
  • Single linkage chains: one long thin cluster forms by hopping neighbor to neighbor. Complete linkage forces compact groups, which can split an elongated one.
All 120 Assignment 1 animals, human ratings · Peterson, Abbott & Griffiths Cognitive Science 2018

Appendix · How many clusters? Cut the tree

Left: the average-linkage dendrogram of the 120 animals with three dashed cut heights labelled 2, 4 and 8 clusters. Right: the MDS map of the same animals coloured by the clusters each cut produces; the 8-cluster cut nearly matches the eight taxa

  • A dendrogram is every clustering at once; you get groups by cutting it at a height. Cut low, many small clusters; cut high, few large ones. The method does not choose for you
  • Cut at 8: nearly the eight taxa (adjusted Rand index 0.890.89). Cut at 2 or 4: not. kk is a choice you defend
  • Choosing where to cut the tree is parallel to choosing mm dimensions in MDS: both methods hand you a free parameter, and both are routinely reported as if it had been discovered rather than chosen
Same 120 animals and ratings · average linkage · map: metric MDS

Appendix · Clustering always returns a tree

Dendrogram of over a thousand patients clustered on fMRI connectivity, cut by a dashed line into four colored groups proposed as depression biotypes

the dashed line is the cut; four colors are the four proposed groups

  • Over a thousand patients with depression, clustered on resting-state fMRI connectivity, cut into four "biotypes", each one claimed to predict who responds to which treatment.
  • Same algorithm as the last three slides. The arithmetic is not in question.
  • The biotypes did not replicate. Reanalysis found no evidence of genuine cluster structure in the data, and a tree comes back from any matrix, including one with no groups in it.
  • Whether the branches name something real is a separate claim, needing separate evidence.

Appendix · MDS and clustering are different hypotheses

MDS Hierarchical clustering
Output continuous coordinates Y\mathbf{Y} a nested tree
Structure assumed a smooth space — items vary along dimensions discrete categories — items belong to groups
Fits well perceptual continua (hue, pitch, size) taxonomies, category structure
Free parameter number of dimensions mm linkage rule + where to cut
Fails on superordinate terms, deep hierarchies continuous variation

Run both. Where they disagree is the interesting question.

Appendix · Human similarity breaks all three rules

Every method so far assumes dissimilarity is a distance. Tversky, with data: human judgments violate all three axioms.

  • Symmetry, dik=dkid_{ik}=d_{ki}: a camel is judged more similar to a horse than a horse to a camel
  • Triangle inequality, dil≤dik+dkld_{il}\le d_{ik}+d_{kl}: lamp ≈ moon (luminous), moon ≈ ball (round), lamp ≉ ball
  • Minimality, dii=0≤dikd_{ii}=0 \le d_{ik}: some different Morse pairs were judged "same" more often than an identical pair
  • "Psychological space" is a useful approximation, not a fact · Assignment 1 bonus · next two slides

Appendix · Violations of minimality and symmetry

Minimality, dik≥dii=0d_{ik} \ge d_{ii} = 0

  • Rothkopf's Morse-code study: participants judged some different pairs "same" more reliably than they judged an identical pair "same".
  • Self-similarity is not constant across items: familiar signals are recognized as themselves more reliably than unfamiliar ones.

Symmetry, dik=dkid_{ik} = d_{ki}

  • A camel is more similar to a horse than a horse is to a camel.
  • An ellipse is more similar to a circle than a circle is to an ellipse.
  • The less typical / less salient item is judged more similar to the prototype than the reverse. Direction of comparison matters.
Rothkopf 1957 (Morse-code confusions) · Tversky Psychol. Rev. 1977

Appendix · Triangle inequality and hierarchy, violated

Triangle of three objects — lamp, moon, ball — where each neighboring pair shares a feature but the far pair shares none

lamp ~ moon (both luminous) · moon ~ ball (both round) · lamp ≁ ball

  • Triangle inequality: dik+dkl≥dild_{ik} + d_{kl} \ge d_{il}. Two items can each be similar to a third on different features, and not at all similar to each other.
  • Nearest-neighbor restriction: in an mm-dimensional Euclidean space, a point can be the nearest neighbor of only a bounded number of others. Conceptual hierarchies break that bound: fruit is the nearest neighbor of many fruits at once, but an MDS map can only put it near two or three.
Tversky Psychol. Rev. 1977 · Tversky & Hutchinson Psychol. Rev. 1986

Appendix · Comparing RDMs: what can go wrong

One correlation is a point estimate. Papers add three things; each has a catch.

  • Uncertainty. Bootstrap over stimuli (and over subjects when people made the RDM) for an interval on ρ\rho. One model beats another only if their intervals separate under the same resampling
  • Ceilings. Split-half is a proxy; toolboxes give a lower and an upper bound from how well each subject agrees with the group
  • Which pairs, which measure. ρ\rho depends on the stimulus set, and Spearman, Pearson, Kendall's τa\tau_a or cross-validated distances can rank models differently on the same data. A result that holds under one measure is a result about the measure
  • What it does not show. Matching RDMs means the two systems order the pairs alike, not that they compute alike

Appendix · Emotion has a geometry too

Map of 2,185 emotion videos, each a letter colored by its dominant category: 27 categories form regions joined by smooth gradients rather than gaps

each letter is one video, colored by its dominant category

  • 2,185 short videos rated by hundreds of viewers for what they felt. No stimulus dimensions, no recordings, judgments are the only input.
  • The structure that comes back needs 27 categories, not the textbook two, and they are joined by smooth gradients rather than gaps: anxiety shades into fear, fear into horror.

The 2-D layout here uses a nonlinear cousin of MDS, taught in a later lecture.

- Tuesday ended at the representational dissimilarity matrix: every pair of images compared, without saying what the comparison is - First half: the comparison, three measures and their properties - Second half: the reverse construction; given an RDM, recover the space the images live in

- Question 1 of Tuesday's minute paper referred to slides not reached in class and is credited to everyone; the answer is the first half of today - The questions are quoted from the minute papers, lightly shortened. The first is the gap Tuesday left: how a pair of rows becomes one number of the RDM - Channels, in a convolutional network: one channel is a sheet of units sharing the same weights, the same feature detector applied at every position; the sheet is a map of where that feature occurs. Each unit is one entry of x, one input dimension in the machine-learning sense. The feature carries no label; the lecture on convolutional networks examines what it detects. A recording channel is one electrode, an unrelated use of the word - Neural vectors: one entry per neuron, each entry a firing rate averaged over that image's presentations; the Assignment 1 monkey data are those averages - Real values: dividing 0–255 by 255 changes units, not information; every image is rescaled the same way, so which images resemble which does not change

- Minute papers asked for the linear algebra to slow down. The pace stays: the course has a lot to cover, and the material is front-loaded; the objects of Theme 1 are reused all semester - Vector operations and matrix products now, the eigendecomposition next week (PCA), gradients in Theme 2 (learning). Nothing beyond that is required for the assignments - Office hours Wednesday 1 PM in Carney 419 (email first so I expect you). TA office hours: vote on Ed; they will be scheduled once there is demand - Two handouts, on the course site under Recitations and on Canvas: linear algebra for the course (vectors, dot products, norms, matrices, the three measures, worked numbers) and programming for Assignment 1 (NumPy, the library calls, the checks). The notation reference defines every symbol in the order the course uses it - Assignment 1 Part 1 is the steepest part of the whole assignment; Parts 2 and 3 reuse what Part 1 builds

- The step Tuesday skipped: two rows of the response matrix, compared, fill one cell of the RDM - The response matrix is the Assignment 1 activation file: 120 images × 4,096 stored late-layer activations of ResNet-18 - The same picture holds for neurons (120 × 482), voxels and human ratings; only the measure d and the meaning of a row change

- Zero diagonal: an image compared with itself shows no difference, d_ii = 0. Symmetry: a comparison does not depend on which image is named first, d_ik = d_ki. Nonnegativity: nothing is less different than identical, d_ik ≥ 0. Any measure that computes a difference between two vectors gives all three; the three of today do - Ratings need not obey them: a rater can say tiger-to-penguin 3 and penguin-to-tiger 5, or rate a photograph against itself below 10. The usual repair averages the two directions and sets the diagonal to zero; the appendix shows where people depart from the constraints - D is images by images; entry (i, k) compares the response vectors of images i and k. Lower-case d is the dissimilarity measure; different measures keep different properties of the responses, and choosing d is a scientific decision - The matrix is the one you build in Assignment 1: 120 images ordered by eight animal categories, correlation distance on the trained network's 4,096 late-layer activations. Within-category mean 0.82, between 0.92; primates and carnivores form the clearest blocks. Colour scale clipped to the 2nd–98th percentile - Pairwise human ratings fit the same format with no response vector measured, so systems can be compared without equating their units or components

- Euclidean distance is the length of the difference vector, the straight-line separation between two points; the double bars denote vector length, also called the Euclidean norm - In two dimensions Pythagoras gives sqrt(55.6² + 21.4²) = 59.5 for tiger and elephant; in D dimensions the sum of squared component differences runs over all D components - Reversing the subtraction changes the direction of the difference but not its length - Euclidean distance counts a change of scale as a difference: doubling every entry of the tiger's vector (a brighter copy of the same image) moves its point to 2x, at distance ‖x‖ = 83.3 from the original, even though the pattern across components is unchanged; cosine similarity ignores this change - Component scales matter: changing a feature's units changes its contribution; multiplying every vector by the same positive constant multiplies all distances by that constant, and adding the same vector to every image leaves all distances unchanged

- Multiplying every pixel by a constant c multiplies both features by c (the mean and the standard deviation of the pixels both scale), so the point moves along the ray through the origin. The points are measured on the scaled images themselves: at 1.5x a few of the 784 resampled pixels saturate at white, so that point sits a hair below the ray (108.1, 58.2 instead of 109.0, 60.9) - Euclidean distance reports a change of brightness as a difference between images; the same happens for a population that fires harder or a layer with larger activations. The next measures compare directions instead

- The dot product of two response vectors is the sum of the entry-wise products; the identity with the cosine of the angle between them is what makes it a measure of alignment. Numbers: tiger (72.7, 40.6), elephant (128.3, 19.2), dot product 10,104, lengths 83.3 and 129.7, cos φ = 0.936, φ = 20.7° - Dividing the lengths out gives cosine similarity, the next slide; the projection picture returns when a neuron's weight vector plays the role of x^(1)

- The vector here is the image itself, 784 pixel values, not the two summary features of the previous slides. Correlation subtracts the mean over a vector's own entries; on the (mean intensity, contrast) plane those entries are the mean and the contrast, so centring there does not remove a shift of the mean. On the pixels the invariance is exact: x + c·1 minus its mean equals x minus its mean - Cosine similarity of 0.991 and 0.978 is close to 1 but not 1: cosine sees the offset, correlation does not. Euclidean distance grows linearly with the offset, 28·c for a 784-pixel image - The pixel values are not clipped: the tiger's brightest pixel is 217, so +60 takes it to 277, above the 255 a camera would record. With unclipped values the three centred images are identical and r is exactly 1; the displayed +60 image shows its brightest pixels saturated at white - The same pair of nuisances exists for neurons: a change of gain (the previous slide) and a change of baseline firing (this one); which measure you choose decides which of the two counts as a difference

- Cosine similarity is the dot product divided by both vector lengths, cos(phi): 1 when two vectors point the same way, 0 when they are perpendicular. Doubling every response leaves it unchanged (next slide, middle panel: Euclidean distance 13.8, cosine similarity 1). Two vectors on the same line through the origin have cosine similarity 1 and cosine dissimilarity 0; their Euclidean distance equals their difference in length - Pearson correlation is cosine similarity after each pattern has had its own mean subtracted, across its units. Adding a constant to every unit leaves it unchanged (next slide, right panel: cosine similarity drops to 0.978, correlation stays 1). Centering is done within each image across its units, not within each unit across images - These are similarities, large when patterns are alike. Distances are dissimilarities, large when patterns differ. To put every measure in one RDM we use dissimilarities, so cosine and correlation enter as 1 minus the similarity - Practical difference: cosine ignores overall gain (a brighter image, a more excitable neuron); correlation also ignores a shared offset (a baseline firing rate, a mean activation). Euclidean distance keeps both - A version of correlation applied to ranks is used later today, when two RDMs are compared. Cosine is undefined for the zero vector and correlation for a constant vector

- In the two-feature plot the eagle's point is nearer the tiger's (Euclidean 23.8 against 54.1), but the penguin's arrow is more nearly parallel to the tiger's (7.5 degrees against 11.9), so cosine dissimilarity ranks the penguin closer (0.009 against 0.021). Correlation cannot rank them: with only two components, centering leaves (a, −a), so every pair has r = ±1; correlation is a measure for long vectors - The same happens with the network's late-layer responses: Euclidean distance ranks the penguin closer to the tiger than the gorilla (122.8 against 125.7); cosine (0.649 against 0.635) and correlation (0.911 against 0.897) rank the gorilla closer - The differences are small, and with real responses the measure often decides close calls. Euclidean distance asks about position; cosine and correlation ask about direction alone - Assignment 1 builds all three measures and applies them to pixels and to early, middle and late activations; expect the rankings to disagree

- The count is automatic: open-access papers whose full text mentions representational dissimilarity matrices or representational similarity analysis, the sentences around the matrix classified by keyword; a paper counts once per measure, so the shares exceed 100 percent - The estimate is rough in two directions: a measure named elsewhere in the paper is missed, and "Euclidean" near an RDM can refer to something other than the RDM's own measure; the ranking is robust, the percentages are not

- Top row, the 120 animals of Assignment 1. Pixels: correlation distance on the 28 × 28 grey version of each image; within-category mean 0.985, between 0.995, no category structure. People: 10 minus the mean rating over ten raters; within 3.7, between 8.6, the strongest block structure on the slide, with birds, primates and reptiles the tightest. ResNet-18 late layer, 4,096 stored units, correlation distance: within 0.82, between 0.92 - Bottom row, the 1,224 object images of Bao and colleagues (Nature 2020), six categories. Pixels: within 0.445, between 0.466. Monkey IT, 482 neurons, correlation distance on the trimmed firing rates: within 0.91, between 1.02, faces and animals the clearest blocks. The same ResNet-18 layer on these images: within 0.60, between 0.75, with faces the tightest block - Each row is one stimulus set seen by three systems; the two rows are different image sets, so their matrices are not comparable with each other, only within a row - Every panel obeys the same three constraints, zero diagonal, symmetry, nonnegativity; that is all the second half of today needs from a matrix

- The second half inverts the first: start from the matrix and find points whose distances reproduce it. Only the table is used, never the responses - The matrix is the human dissimilarity D = 10 − mean rating for six of the 120 images; the map is the two-dimensional configuration from metric MDS (scikit-learn), with three of the fifteen map distances printed on their segments - Six points on a plane have twelve free coordinates (fewer after removing rotation and translation) against fifteen distances to match, and ratings are not exact distances, so every pair is off by a little; stress measures that mismatch

- The idea of a psychological space is old: objects, colours, sounds and words as points, with nearby points alike, so that generalization, confusion and categorization are facts about distance. Shepard built a career on it. No behavioural method reads the coordinates of a point in that space - What can be measured is what a distance would produce: a judgment of how alike two things are. Psychophysics has many designs for this (direct ratings, confusions, odd-one-out choices, sorting); all end as one number per pair, a dissimilarity matrix. MDS goes from that table back to points - Peterson, Abbott & Griffiths (2018) collected ratings on a 0 to 10 scale for every pair of 120 animal photographs, ten raters per pair, 7,140 pairs in all. The numbers on the slide are their mean ratings: the two tigers A and B were judged maximally similar (10.0), and each tiger about equally dissimilar from the monkey (2.9 and 2.2). Tuesday's opening slide used the same two tigers with a third tiger in place of the monkey - Behaviour is a representation too: the pattern of ratings describes how the images are organised in people's heads without saying anything about neurons - Assignment 1 uses these ratings and asks which measurement, pixels, IT neurons or a network layer, comes closest to reproducing them; other formats such as odd-one-out triplets give similar structure

- MDS recovers a space in which the images are points whose separations reproduce the RDM entries. Points on a page satisfy all four requirements, so an RDM that violates one cannot be reproduced exactly by any arrangement of points; the recovered map distorts it. For a network's RDM this depends on the measure d; for a human matrix nothing guarantees it (appendix) - The rows: nonnegativity, no pair of images is less than zero apart; identity, an image is at distance zero only from itself, which cosine breaks because a brighter copy of the tiger is at cosine dissimilarity zero from the original; symmetry, tiger to elephant equals elephant to tiger; triangle, no route through a third image is shorter than the direct comparison, which cosine violates on 2.3 % of pixel triples in our images - Only Euclidean distance is a metric: it is the length of a difference vector. Cosine and correlation discard gain (correlation also offset), so identity fails, since x and 2x count as the same pattern, and so does the triangle inequality - They remain useful. Cosine dissimilarity is a monotone function of the angle between the vectors, and the angle is a proper distance (the arc between two points on a sphere), so cosine and correlation order pairs of images as a distance would; comparing RDMs uses that order. √(2(1 − cos φ)) is the Euclidean distance between the two length-normalized vectors, and it is a metric. Expect small distortions in a map recovered from cosine or correlation dissimilarities - A cross means the requirement can fail, not that every pair violates it; one counterexample disproves a metric, and no finite check proves one. Assignment 1 builds these checks, with a numerical guard for zero or constant vectors

- Stimulus i is the point y^(i), a vector with m entries (here m = 2); the map distance between two points carries a hat to keep it apart from the measured entry d_ik - Each arrow is a correspondence the map has to honour for all fifteen pairs: an entry in the table is a distance on the map

- MDS takes only the dissimilarity matrix D and returns one point per stimulus in m dimensions, placed so that map distances reproduce the measured dissimilarities; the raw responses, the channels and the stimuli are never used - Analogy: given a table of driving distances between cities, draw the map

- Ekman (1954) had observers judge the similarity of 14 monochromatic lights from 434 to 674 nm; Shepard's (1962) MDS of those judgments returns a circle ordered by wavelength - The recovered structure is a closed curve: one dimension (position around the ring) drawn in two, which is why the fit needs two dimensions - The closure of the circle, violet next to red, is a psychological fact with no physical counterpart: the stimulus dimension is a line, the percept is a ring

- Physical analogy: each pair (i, k) is a spring with rest length d_ik, stress is the total potential energy, and the layout relaxes into the least-strained configuration - Lowering a number by moving the points a little at a time is the same downhill idea that trains the networks of Theme 2 - A local minimum is a layout that no small move improves, although a better layout exists elsewhere

- scikit-learn's MDS solver, run from one fixed random start and stopped after 0, 10 and 300 rounds; each frame is the solver's actual state - Moving each point by an amount computed from the mismatch of its distances is the optimisation idea of Theme 2

- After ten rounds the stress has fallen from 217 to 64: the coarse layout, which animal is near which, is already in place, and the remaining rounds shrink the small mismatches

- Convergence means no small move of any point reduces the stress; 22 is what remains because the fifteen ratings are not exactly the distances of any six points on a plane - Different random starts can converge to different local minima; scikit-learn's n_init runs several and keeps the best, which is what Assignment 1 does

- Raw stress has squared-distance units; dividing by the summed squared map distances removes the overall scale, so Stress-1 can be compared across analyses. In nonmetric MDS the fitted targets are disparities, a monotone transformation of the dissimilarities - Kruskal's (1964) rule of thumb (0.05 excellent, 0.10 fair, 0.20 poor) is not a significance test: item count, target dimension, ties and the optimiser all move the value, which is why Assignment 1 also refits on shuffled dissimilarities - Stress values in this lecture come from scikit-learn 1.8

- The common error is reading the horizontal axis of an MDS plot as if it were a principal component; PCA gives ordered, meaningful directions and MDS does not, which is the cleanest way to tell the two methods apart before the PCA lecture - On dimensionality, the answer is the elbow in the stress-against-m curve together with a stated choice

- Shepard's 1962 result: ordinal information alone, given enough stimuli, pins down a metric configuration almost uniquely; that is why non-metric MDS became the standard tool in psychology - Which falling curve is recovered is itself a finding; Shepard's law, two slides on, is the most famous answer. Shepard 1962; Kruskal 1964

- Fifty-one published similarity matrices, collected by Michael Lee (UC Irvine); every fit was computed in advance with the same scikit-learn call as Assignment 1 - Ekman's colours: stress 0.29 in one dimension, 0.03 in two, and the two-dimensional solution is the colour circle, 434 nm next to 674 nm - Rothkopf's Morse code: one dimension already orders the signals by length (E and T at one end, the digits at the other); the second separates dots from dashes - The stress curve is the tool for choosing dimensions; look for the elbow, not for a threshold; the kinship terms and Kruschke's rectangles show clear elbows, the Romney word lists do not - The dendrogram is the second hypothesis about the same table; the clusters it returns need not match the neighbourhoods of the map

- MDS is fit to the ordinal structure only, so the exponential shape of generalization against distance is an empirical outcome - The constant c is a sensitivity parameter that varies across tasks; the exponential form does not - The law was only ever tested on simple, low-dimensional stimuli; whether it survives real-world stimuli, and why the exponential arises at all, are open questions; Tenenbaum and Griffiths (*Behavioral and Brain Sciences* 2001) give the Bayesian account of the exponential. Shepard, *Science* 1987

- Assignment 1 does this: three fits to the same pairs of MDS distance and rated similarity, each reported with R², the share of the spread in similarity the curve accounts for (1 perfect, 0 no better than the mean) - The line is the null: constant-rate decay, no law worth stating. Any bending curve beats it, so the real comparison is a second bending curve with the same number of parameters - The Gaussian a·exp(−b·d²) decays and flattens like the exponential but is flat at the origin; Shepard's claim is that generalisation is steepest at zero distance - The curves on the slide are analytic with a = b = 1, not fits to assignment data. Shepard 1987 - Preview of Assignment 1: in the psychological space recovered from the ratings, the exponential is the best of the three curves. Against distances in a network layer the same ratings give a different answer, because the x-axis changed: distance in a 4,096-unit layer is bent relative to psychological distance. Shepard's law is a claim about the psychological space

- Two systems, the same stimuli in the same order, two matrices of the same size. Both show blocks in the same places, primates and carnivores clearest, and the network's are weaker. This section puts a number on that resemblance - The units differ: a rating scale on the left, a correlation distance on the right. Whatever number is chosen has to survive that

- The change of level: in the first half a measure compared two response vectors and produced one entry of an RDM; now a measure compares two RDMs and produces one number for a pair of systems - Nothing new is needed: the upper triangle of an RDM is a vector of N(N−1)/2 entries, and Euclidean, cosine and correlation apply to any two vectors of the same length - Rank correlation is the default because the two RDMs have unrelated units. Pearson asks a stricter question, whether the dissimilarities are linearly related, and gives a different number on the same pair of matrices (0.57 against 0.42) - Assignment 1 computes this for the human RDM against a trained and an untrained network

- The same idea underlies non-metric MDS (only the order of the distances matters) and the comparison of two RDMs on the next slides (only the order of the pairs matters); the middle panel is what Spearman actually computes - Warping: two systems whose dissimilarities are related by any rising curve, exp(0.35x) here, have Spearman 1 and Pearson below 1; the swap in the right panel is the kind of disagreement Spearman is sensitive to

- The operation has a name, representational similarity analysis (RSA); model comparison and ranking, Brain-Score, noise-ceiling inference, significance testing and the zoo of alternative measures come in a later lecture - This is the operation that produced the number relating the human RDM and the network RDM over the same 120 photographs - The diagonal is N pairs on which the two matrices agree perfectly by construction, so including it drags rho upward; the lower triangle doubles the apparent sample size, invisible until someone asks how likely the result is by chance - Spearman is used because the matrices share no common scale, and because Shepard's law is why rank information is the trustworthy part of a similarity measurement - Two 120 × 120 matrices in, one number out; Assignment 1 does this operation, the behavioural RDM against each layer of each network. Kriegeskorte, Mur & Bandettini 2008; Kriegeskorte & Kievit 2013

- Two systems can agree on every pair while measuring different things about the images. Two networks trained on different objectives can produce nearly the same RDM over a stimulus set that does not distinguish them; the RDM is a summary, and systems that differ underneath can match it - So "the model matches IT" is a weaker claim than it sounds - The split-half reliability of the human ratings, two slides on, gives Assignment 1 its comparison scale; it is a proxy, not a hard ceiling. Formal inference and model-ranking uncertainty come in the later model–brain lecture. Kriegeskorte & Kievit, *TICS* 2013

- Kriegeskorte, Mur, Ruff et al. (Neuron 2008): 92 colour object photographs, monkey IT single units (Kiani et al. 2007 data) and human IT fMRI voxels, both matrices computed with correlation distance and shown as percentiles - Both matrices show the animate/inanimate division and the face and body sub-blocks; the paper reports a linear correlation of 0.49 between the two matrices, the operation of the previous slides applied to two brains - Figure: Kriegeskorte et al., Neuron 2008, fig. 1, reproduced for teaching with the colour bar moved to the side

- The RDM correlation applied five times, each system against the same human RDM over the same 120 photographs; you can recompute every number from the Assignment 1 data - Raw pixels carry none of the human ordering, each stage of the trained network carries more, and the untrained copy of the best stage carries none, so the alignment is a property of the learned weights - The dashed line is the noise ceiling: split-half reliability, the two rater batches Peterson released correlated with each other, 0.82 on the Spearman scale. The averaged ratings are more reliable than either half, so "half of the ceiling" is an upper bound on what the model has reached - Assignment 1 repeats this with your own RDMs; it also asks a second, Shepard-like question: does the network's similarity fall off with distance in the human space the way human similarity does

- Same operation as the previous slide, five more models; features are the pooled penultimate layer (512 to 2,048 numbers per image). On the same pooled features ResNet-18 reaches 0.50, so the deeper ResNets add 0.02, not the 0.10 the bars suggest against the 4,096 sampled units of Assignment 1 - ResNet-50 and ResNet-152 (ImageNet V2 weights) both reach 0.52; the supervised ViT-B/16 (a vision transformer) and ConvNeXt-Base, more accurate on ImageNet than any ResNet, fall to 0.26 and 0.20 - DINOv2 (ViT-S/14, self-supervised on 142 million images, no labels) reaches 0.61, about three quarters of the 0.82 ceiling; it is the model whose features Assignment 1 also provides - This is the pattern of Linsley et al. 2026: past a point, ImageNet accuracy and human-likeness come apart; the training objective, not the size, is what moves alignment

- Parts 1 and 2 are fully covered as of today; Part 3 needs Tuesday's lecture (PCA) - Each notebook is self-contained, so you can work ahead; trying a part before its lecture is a good way to learn it

- The RDM is the common currency that makes brains, behaviour and models commensurable - MDS, Shepard's law, Tversky's violations, clustering and RSA are operations on, or checks of, the RDM

- None of these is required reading. Shepard 1987 and 1980 are the two classics behind today's lecture; the two Transmitter essays are short and set up two questions of later lectures (how to compare representations, what dimensionality means); the last three are recent studies that use MDS and RDM correlation - The research trail asks for a paper you chose and can explain; if several of you converge on the same one, branch from its references or from the papers that cite it

- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 3) - Question 1 is the one Tuesday's minute paper asked before the lecture reached it - Credit is for a thoughtful attempt, not for being correct; recurring questions are summarized anonymously and addressed on Ed or at the start of the next lecture

- Subtracting vectors gives the displacement from one image to the other: x^(2) − x^(1) = (128.3 − 72.7, 19.2 − 40.6) = (55.6, −21.4) for the tiger, image 1, and the elephant, image 2 - The difference is itself a vector, one entry per feature: drawn from the tiger's point, it ends at the elephant's, and its length is the distance between the two representations

- N enters the output quadratically, as N(N−1)/2 pairs, while the number of channels n does not enter at all: n numbers per image in, one number per pair out - Assignment 1 asks you to write these distances by hand, then checks them against this SciPy call

- All 4,096 stored late-layer (layer4) ResNet-18 activations for the tiger, without thresholding or normalization: 1,843 (45.0%) are exactly zero and the maximum is about 13.57; activations are model values, not firing rates or probabilities - Each ResNet-18 block ends with a rectified linear unit (ReLU) after the residual addition, and the layer4 activations are recorded after it; ReLU maps nonpositive inputs to zero and passes positive inputs unchanged (−2 → 0, 3 → 3), which is why so many entries are exactly zero, and a stored zero does not reveal its preactivation - ReLU is nonlinear because it is not additive: ReLU(−2 + 3) = 1 but ReLU(−2) + ReLU(3) = 3; it is piecewise linear, not one linear function - Nonlinearity is what gives depth its power: a stack of linear maps collapses to a single linear map, W2(W1 x) = (W2 W1) x, and with biases to a single affine map, whereas nonlinear units let successive layers build input-dependent representations. Max pooling is nonlinear too - A zero for one image is not an inactive feature: none of the 4,096 features is zero for all 120 assignment images

- Comparing the tiger's and gorilla's full late-layer vectors on the same 4,096 features, exactly 879 entries are zero in both images. ReLU makes such entries common; firing rates and BOLD values are almost never exactly equal - What goes wrong: correlation subtracts each vector's mean before comparing. Hundreds of shared zeros pull both means down, and every zero then counts as a below-average entry that the two images share, which manufactures agreement: r = 0.103 with all 4,096 entries, −0.011 on the 3,217 entries that are not zero in both. Correlation dissimilarity moves from 0.897 to 1.011. Euclidean distance (125.66) and cosine dissimilarity (0.635) do not move, because a shared zero adds nothing to a squared difference, a dot product or a norm - Correlation depends on which units are included. Dropping "units silent for this pair" gives every pair its own unit set, and the RDM no longer compares like with like - What to do: decide the unit set once, over the whole image set (for instance, drop units that are zero for every image; here there are none, so keep all 4,096), apply that same set to every pair, and state the choice with the result. With ReLU activations, cosine is the gain-free measure that is immune to this particular choice; correlation is fine provided the unit set is fixed. Assignment 1 examines retaining versus excluding inactive units

- The weights come from training: the network was shown labelled photographs and its weights adjusted, step by step, to reduce the classification loss. The RDM is a property of those weights, not of the architecture - The full 120 × 120 RDM (correlation distance on the 4,096 late-layer activations), computed twice: with random weights and with ImageNet-trained weights; images ordered by taxon, so a category is a block on the diagonal - Random weights: every off-diagonal entry is close to 0.15 (range 0.03 to 0.41), within and between categories alike; the untrained network gives almost the same pattern for every image - Trained: the mean distance rises to 0.90, and same-category pairs (0.82) are closer than different-category pairs (0.92), the block structure; amphibians (0.72) and rodents (0.76) are the tightest categories - Each panel has its own colour scale: on the trained panel's scale the random-weights matrix would be one pale square; on its own scale it is noise with no block structure.

- Every network RDM in this lecture was computed with weights that came from training; the RDM is not a property of the architecture alone. Supervised training pairs each image with a category label; the network outputs class probabilities, and a loss scores the prediction against the true label; cross-entropy, the usual classification loss, penalizes low probability on the correct class rather than counting misclassifications - A learning algorithm repeatedly updates the weights on batches of examples to lower the average loss; the outcome is a learned set of weights, not a guaranteed global optimum, and performance is measured on held-out images, images not used in training - Changing the weights changes the internal responses to the same image, and therefore the distances between image representations; the objective concerns labels, not human judgments or neural responses - Once trained, the weights are fixed while responses are extracted; presenting an image to a fixed network does not train it - The diagram is schematic: the four photographs are Assignment 1 images with their labels, the frog is the example currently passing through, and the probability bars are invented, not a measured prediction

- The change of source matters: a network's RDM inherits zero diagonal, symmetry and nonnegativity from the measure that made it (all three of ours satisfy them), and the triangle inequality when the measure is a distance. A human rating is a number a person chose; averaging over ten raters and symmetrising is what makes the Peterson matrix usable, and nothing in the procedure guarantees the triangle inequality - Tversky (1977) showed with data that human judgments violate symmetry, minimality and the triangle inequality; the appendix has the three demonstrations, and the Assignment 1 bonus tests the axioms on our own matrices - Why it matters: the second half of the lecture starts from an RDM and tries to recover the space behind it, one point per image. Points on a page obey all four properties automatically, so a matrix that violates one cannot be reproduced exactly; the mismatch is absorbed by stress

- Trap 1: forgetting `metric="precomputed"` gives a silently wrong answer rather than an error, because scikit-learn then computes Euclidean distances from D as if it were raw data; Assignment 1 hands you the MDS call ready-made for this reason - Trap 2: what `mds.stress_` reports depends on the call; with scikit-learn 1.8 and `normalized_stress="auto"`, the metric call prints raw stress and the nonmetric call prints Stress-1, so compare like with like

- Least squares is the whole method: the fitted curve is the member of the family (all a·exp(−b·d), say) with the smallest sum of squared vertical residuals; curve_fit searches a and b for that minimum, polyfit solves it directly for a line - R² compares that residual sum with the spread around the mean; it can be negative when a curve does worse than the flat line at the mean, and it is not a probability - Assignment 1 fits the raw pairs, all 7,140 of them. Published R² values are often computed on binned means, which flatters any curve; when you compare with a paper, check which was done

- Assignment 1 needs the ceiling operation: agreement with humans requires a reference for how consistently humans agree with one another, and the split-half correlation is that reference - The split-half value is a conservative proxy, not a hard maximum: under independent noise the half-sample denominator is too small, which inflates the reported fraction, hence "upper bound" - Do not turn a rank correlation into variance explained; Assignment 1 uses ResNet-18 and Spearman correlation on correlation distances, while Peterson's paper reports VGG features and a regression R², so the numbers are not comparable - The 0.42 and 0.82 are computed on the Assignment 1 data: the two released rater batches, and ResNet-18's late layer. Peterson et al. 2018

- In a dendrogram the horizontal axis is not meaningful: leaf order is arbitrary up to flipping any join, just as an MDS configuration is arbitrary up to rotation - Height is meaningful, since it is the distance at which two groups merged; the same warning as for MDS axes, in a different guise

- For the Assignment 1 bonus: SciPy's `linkage` expects the condensed form, so pass `squareform(D, checks=False)`, not the square matrix - Passing the square matrix is a classic twenty-minute debugging trap, because it is silently treated as N observations in N dimensions

- Every linkage rule has a knob, and the knob encodes an assumption about what a cluster is: single linkage chains, complete linkage forces compact groups - The same lesson as the choice of dissimilarity measure: the choice decides the answer

- The adjusted Rand index compares two partitions: 1 when they are identical, about 0 for chance agreement. 0.89 means the eight-cluster cut nearly recovers the eight taxa

- Drysdale et al. (*Nature Medicine* 2017) proposed neurophysiological subtypes of depression tied to treatment response, with a standard pipeline; the paper has thousands of citations - Hierarchical clustering has no null hypothesis: 1,000 points from a single Gaussian blob return a clean-looking dendrogram, because merging is all the algorithm can do - "We found four clusters" is not a finding; "the four clusters survive resampling, hold up in a held-out sample, or beat a null with no group structure" is - Dinga et al. (2019) reran the pipeline on the original data and new samples: the cluster structure was not statistically distinguishable from no clusters, and the treatment-prediction advantage did not hold up; later multi-site work was similarly unsupportive - The Assignment 1 bonus meets the same failure; like the number of dimensions m, the number of clusters is a modelling choice. Figure: Drysdale et al., *Nature Medicine* 2017, fig. 1e, Springer Nature, all rights reserved, used with critical commentary in a non-commercial lecture

- The choice of visualisation is itself a hypothesis about the domain, a smooth space for MDS and discrete categories for clustering - Shepard, "Multidimensional scaling, tree-fitting, and clustering", *Science* 1980, makes this argument; optional reading

- Tversky's argument sets the domain in which MDS is valid - Euclidean RDMs satisfy the metric axioms, but cosine and correlation dissimilarities need not, even when computed from response vectors; human judgments add further failures, including asymmetry - Both kinds of dissimilarity are computed this semester, so the distinction matters. Tversky, *Psychological Review* 1977

- The asymmetry is easy to reproduce: rate how similar North Korea is to China, then China to North Korea; the ratings differ - Tversky's explanation is featural, not spatial: similarity is a weighted contrast of common and distinctive features, and the features of the first-named item are weighted more heavily. Rothkopf 1957; Tversky 1977

- An MDS map of fruit names puts "fruit" in the centre, yet it can be the nearest neighbour of only two or three items, because a Euclidean space bounds how many points can share a nearest neighbour - A tree handles a superordinate term naturally and a flat map does not, which is why clustering follows MDS. Tversky 1977; Tversky & Hutchinson 1986

- Not part of Assignment 1; papers you may pick for the research trail use these terms - Bootstrapping over stimuli asks whether the result would survive a different sample of images from the same pool; bootstrapping over subjects asks the same about the people. Nili et al. 2014 (the RSA toolbox) and Schütt et al. 2023 (rsatoolbox, model-comparison inference with both bootstraps) are the standard references - Ceilings: a group RDM predicts each subject's RDM only so well; the upper bound uses the group including the subject, the lower bound excludes them; Nili et al. 2014 - Cross-validated (crossnobis) distances remove the positive bias that noise adds to every distance; Walther et al. 2016. Soni et al. 2024 show that the ranking of models against brain data changes with the similarity measure; the next lecture, on the geometry of neural representations, defines CKA; shape metrics come in the later model–brain lecture - Geometry is evidence about a representation, not a demonstration of shared mechanism; Kriegeskorte & Kievit 2013

- This is the first case with no physical stimulus dimension to fall back on: colour had wavelength, emotion has nothing - The 2-D layout in the figure is t-SNE, not MDS; what transfers is the recipe (judgments, pairwise structure, map), not the solver - The field's default account has two dimensions, valence and arousal; Cowen & Keltner need about 27 categories with continuous boundaries, more categories than the dimensional view wants and fuzzier edges than the basic-emotion view wants - The open weakness is whether the 27th dimension is real or an artefact of how many response options were offered. Cowen & Keltner, *PNAS* 2017, fig. 2A