MDS: recovering psychological spaces

Reader-friendly version of the 67 lecture slides: the lecture content in normal flow, with figure descriptions and data tables where the figure carries data. Open the presentation.

Slide 1

Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 2 · Theme 1: Representational spaces

MDS: recovering psychological spaces

Thursday, September 17 · Fall 2026

Return to contents

Slide 2

Last time, and today

Tuesday: a representation is a list of numbers, a point in a space; pixels, voxels, neurons and network units all give one; averaging over repetitions; the response matrix 𝐗\mathbf{X}; the RDM 𝐃\mathbf{D}, one comparison per pair of images.

Today, part 1: what the comparison is. Euclidean distance, cosine, correlation; which of them are distances; the RDMs of pixels, people, neurons and a network.

Today, part 2: from the RDM back to a space. The psychological space: each stimulus a point, the distance between two points how dissimilar they seem to a person. Recovered from judgments alone by multidimensional scaling.

Return to contents

Slide 3

Your questions from Tuesday

Four questions came up often enough to answer here:

Return to contents

Slide 4

"The linear algebra is going too fast"

Return to contents

Slide 5

Comparing two responses

Return to contents

Slide 6

From the response matrix to the RDM

Left: the response matrix, 120 rows, one image each, 4,096 columns of unit activations, drawn as a tall strip of values with two rows highlighted; an arrow labelled with the dissimilarity measure d leads to the right: a 120 by 120 matrix with the single entry for that pair of rows highlighted

Left: the response matrix, 120 rows, one image each, 4,096 columns of unit activations, drawn as a tall strip of values with two rows highlighted; an arrow labelled with the dissimilarity measure d leads to the right: a 120 by 120 matrix with the single entry for that pair of rows highlighted

Open full-size figure
  • Response matrix 𝐗\mathbf{X}: NN rows (images), DD columns (units); here N=120N = 120, D=4,096D = 4{,}096
  • Two rows, 𝐱(i)\mathbf{x}^{(i)} and 𝐱(k)\mathbf{x}^{(k)}, compared with the measure dd: one number, dikd_{ik}
  • Every pair, stored by row and column: the RDM 𝐃\mathbf{D}, N×NN \times N
  • DD numbers per image go in; one number per pair comes out

Return to contents

Slide 7

An RDM compares every pair of images

Representational dissimilarity matrix (RDM): every pairwise comparison.

𝐃∈ℝN×N,dik=d(𝐱(i),𝐱(k))\mathbf{D}\in\mathbb{R}^{N\times N},\qquad d_{ik}=d\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)

  • Rows and columns: images, here ordered by category. Each entry: one comparison
  • dd: the dissimilarity measure. It decides which differences count
  • What makes it a dissimilarity matrix: zero diagonal, dii=0d_{ii}=0; symmetric, dik=dkid_{ik}=d_{ki}; nonnegative, dik≥0d_{ik}\ge0. A fourth property, the one points in a space obey, comes later today
A 120 by 120 representational dissimilarity matrix of the Assignment 1 images from the trained ResNet-18 late layer, correlation distance, images ordered by animal category with category boundaries drawn; light blocks along the diagonal show that primates, carnivores and hoofed mammals are more alike within their category than across

A 120 by 120 representational dissimilarity matrix of the Assignment 1 images from the trained ResNet-18 late layer, correlation distance, images ordered by animal category with category boundaries drawn; light blocks along the diagonal show that primates, carnivores and hoofed mammals are more alike within their category than across

Open full-size figure

Return to contents

Slide 8

Euclidean distance

Length of the difference vector:

deuc(𝐱(i),𝐱(k))=∥𝐱(k)−𝐱(i)∥=∑j=1D(xj(k)−xj(i))2d_{\mathrm{euc}}\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\bigl\|\mathbf{x}^{(k)}-\mathbf{x}^{(i)}\bigr\|=\sqrt{\sum_{j=1}^{D}\bigl(x_j^{(k)}-x_j^{(i)}\bigr)^2}

Tiger and elephant, 𝐱(2)−𝐱(1)=(55.6, −21.4)\mathbf{x}^{(2)}-\mathbf{x}^{(1)}=(55.6,\,-21.4):

deuc=55.62+21.42=59.5d_{\mathrm{euc}}=\sqrt{55.6^2+21.4^2}=59.5

  • Distance asks how far apart two points are; nothing else
The two-feature plot, mean intensity against contrast, with the tiger's and the elephant's points; the segment between them is labelled with its Euclidean length, 59.5

The two-feature plot, mean intensity against contrast, with the tiger's and the elephant's points; the segment between them is labelled with its Euclidean length, 59.5

Open full-size figure

Return to contents

Slide 9

Same tiger, three brightnesses

  • Scale every pixel by 0.50.5, 11, 1.51.5: the same tiger, darker or brighter. The three points lie on one line through the origin
  • Euclidean distance between the darkest and the brightest: ∥1.5 𝐱(1)−0.5 𝐱(1)∥=∥𝐱(1)∥=83.3\|1.5\,\mathbf{x}^{(1)}-0.5\,\mathbf{x}^{(1)}\|=\|\mathbf{x}^{(1)}\|=83.3, more than tiger to elephant (59.559.5)
  • What the three share is their direction; what differs is their length. Direction = the pattern. Length = overall strength, the gain (brightness here; firing rate or activation elsewhere)
  • A measure of the pattern should compare directions and ignore the gain
The feature-space plot with the tiger at half, normal and one-and-a-half brightness, three thumbnails whose points lie on one dotted line through the origin, the normal tiger marked by the blue vector x superscript 1

The feature-space plot with the tiger at half, normal and one-and-a-half brightness, three thumbnails whose points lie on one dotted line through the origin, the normal tiger marked by the blue vector x superscript 1

Open full-size figure

Return to contents

Slide 10

The dot product measures alignment

For two samples, 𝐱(1)\mathbf{x}^{(1)} (tiger) and 𝐱(2)\mathbf{x}^{(2)} (elephant):

𝐱(1)⊤𝐱(2)=∑j=1Dxj(1)xj(2)=∥𝐱(1)∥ ∥𝐱(2)∥cos⁡ϕ\mathbf{x}^{(1)\top}\mathbf{x}^{(2)}=\sum_{j=1}^{D}x^{(1)}_j x^{(2)}_j=\|\mathbf{x}^{(1)}\|\,\|\mathbf{x}^{(2)}\|\cos\phi

  • ϕ\phi is the angle between the two vectors: 0°0° when they point the same way, 90°90° when perpendicular
  • The dot product grows with both lengths and with the alignment cos⁡ϕ\cos\phi
  • Tiger and elephant: 𝐱(1)⊤𝐱(2)=10,104=83.3×129.7×cos⁡ϕ\mathbf{x}^{(1)\top}\mathbf{x}^{(2)}=10{,}104 = 83.3\times129.7\times\cos\phi, so cos⁡ϕ=0.94\cos\phi=0.94, ϕ=21°\phi=21°
  • ∥𝐱(2)∥cos⁡ϕ\|\mathbf{x}^{(2)}\|\cos\phi is the projection of 𝐱(2)\mathbf{x}^{(2)} onto the direction of 𝐱(1)\mathbf{x}^{(1)}, how far 𝐱(2)\mathbf{x}^{(2)} reaches along it: 121.4121.4
The feature-space plot with the tiger vector x superscript 1 and the elephant vector x superscript 2, the angle phi between them marked at the origin, and the projection of the elephant vector onto the tiger's direction shown as a magenta segment with a dashed perpendicular

The feature-space plot with the tiger vector x superscript 1 and the elephant vector x superscript 2, the angle phi between them marked at the origin, and the projection of the elephant vector onto the tiger's direction shown as a magenta segment with a dashed perpendicular

Open full-size figure

Return to contents

Slide 11

Same tiger, plus a constant

  • Add 3030, then 6060 to every pixel (the 784 numbers of the 28 × 28 image): a washed-out tiger. Euclidean distance 840840 and 1,6801{,}680 (28×3028\times30, 28×6028\times60: this is the 784-pixel space); cos⁡ϕ=0.991\cos\phi = 0.991, 0.9780.978. Both report a change; the pattern did not change
  • Subtract each image's own mean from its pixels: the three become the same image, entry for entry
  • Cosine of the centred vectors is 11 for every pair: that is correlation, next slide
  • A measure of the pattern should ignore a gain (last slide) and an offset (this one): a brighter image, a shift in baseline firing rate
Top row: the tiger as a 28 by 28 grey image, then with 30 and with 60 added to every pixel, progressively washed out, means 73, 103 and 133. Bottom row: each image minus its own mean, three identical images with mean 0

Top row: the tiger as a 28 by 28 grey image, then with 30 and with 60 added to every pixel, progressively washed out, means 73, 103 and 133. Bottom row: each image minus its own mean, three identical images with mean 0

Open full-size figure

Return to contents

Slide 12

Two similarity measures: cosine and correlation

scos⁡(𝐱(i),𝐱(k))=cos⁡ϕ=𝐱(i)⊤𝐱(k)∥𝐱(i)∥ ∥𝐱(k)∥r(𝐱(i),𝐱(k))=𝐱c(i)⊤𝐱c(k)∥𝐱c(i)∥ ∥𝐱c(k)∥,𝐱c=𝐱−xˉ 𝟏s_{\cos}\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\cos\phi=\frac{\mathbf{x}^{(i)\top}\mathbf{x}^{(k)}}{\|\mathbf{x}^{(i)}\|\,\|\mathbf{x}^{(k)}\|}\qquad\qquad r\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)=\frac{\mathbf{x}_c^{(i)\top}\mathbf{x}_c^{(k)}}{\|\mathbf{x}_c^{(i)}\|\,\|\mathbf{x}_c^{(k)}\|},\quad \mathbf{x}_c=\mathbf{x}-\bar{x}\,\mathbf{1}

Return to contents

Slide 13

What each measure ignores

Three bar charts of a twelve-unit response pattern: the pattern itself, every unit doubled, and a constant added to every unit; under each, marks show which measures call it the same as the original: Euclidean distance none, cosine similarity only the doubled pattern, correlation both

Three bar charts of a twelve-unit response pattern: the pattern itself, every unit doubled, and a constant added to every unit; under each, marks show which measures call it the same as the original: Euclidean distance none, cosine similarity only the doubled pattern, correlation both

Open full-size figure
Assignment 1 §1b: the three functions, then all three matrices

Return to contents

Slide 14

Which is closer to the tiger? It depends on the measure

Two features (plot): eagle or penguin?

Which is closer to the tiger? It depends on the measure — table
Pair Euclidean Cosine* Correlation*
tiger–eagle 23.8 0.021 0†
tiger–penguin 54.1 0.009 0†

4,096 network responses: penguin or gorilla?

Which is closer to the tiger? It depends on the measure — table
Pair Euclidean Cosine* Correlation*
tiger–penguin 122.8 0.649 0.911
tiger–gorilla 125.7 0.635 0.897

* dissimilarities, 1−cos⁡ϕ1-\cos\phi and 1−r1-r · † with two components, r=±1r=\pm1 always: correlation needs more than two

The two-feature plot with three arrows from the origin: the tiger x superscript 1, the eagle x superscript 2, and the penguin x superscript 3. The eagle's point is nearer the tiger's, but the penguin's arrow is more nearly parallel to the tiger's

The two-feature plot with three arrows from the origin: the tiger x superscript 1, the eagle x superscript 2, and the penguin x superscript 3. The eagle's point is nearer the tiger's, but the penguin's arrow is more nearly parallel to the tiger's

Open full-size figure

Return to contents

Slide 15

Which measure do papers actually use? All of them

Horizontal bar chart over 192 published RSA papers from 2021 to 2026: correlation distance 49 percent, Euclidean 38, Mahalanobis or crossnobis 8, cosine 12, measure not stated 20 percent; the title says 24 percent use two or more measures and that RDMs are compared with Pearson in 48 percent, Spearman in 40, Kendall in 7

Horizontal bar chart over 192 published RSA papers from 2021 to 2026: correlation distance 49 percent, Euclidean 38, Mahalanobis or crossnobis 8, cosine 12, measure not stated 20 percent; the title says 24 percent use two or more measures and that RDMs are compared with Pearson in 48 percent, Spearman in 40, Kendall in 7

Open full-size figure
192 open-access papers that compute an RDM, 2021–2026, Europe PMC full text; the measure counted automatically · an estimate

Return to contents

Slide 16

Six RDMs, one format

Six dissimilarity matrices in two rows, images ordered by category with boundaries drawn, each on its own colour scale. Top row, the 120 animal photographs: pixels show no block structure, human judgments show strong light blocks on the diagonal, one per category, the ResNet-18 late layer shows partial blocks. Bottom row, the 1,224 object images of Bao and colleagues: pixels show faint structure, monkey IT neurons show clear blocks for faces and animals, the ResNet-18 late layer shows a strong face block

Six dissimilarity matrices in two rows, images ordered by category with boundaries drawn, each on its own colour scale. Top row, the 120 animal photographs: pixels show no block structure, human judgments show strong light blocks on the diagonal, one per category, the ResNet-18 late layer shows partial blocks. Bottom row, the 1,224 object images of Bao and colleagues: pixels show faint structure, monkey IT neurons show clear blocks for faces and animals, the ResNet-18 late layer shows a strong face block

Open full-size figure
  • Two stimulus sets, five kinds of measurement, one format: an N×NN\times N matrix, the stimuli on both axes, one number per pair
  • Light blocks on the diagonal: the categories a system treats as alike. None in pixels; strong in people and in IT neurons; partial in the network
  • Units differ (a rating, a correlation distance); the format does not. The shared format is what lets two systems be compared
Assignment 1 data · people: Peterson, Abbott & Griffiths Cognitive Science 2018 · monkey IT: Bao et al. Nature 2020 · each panel on its own colour scale

Return to contents

Slide 17

From the matrix back to a space

Return to contents

Slide 18

From the matrix to a map

Left: the 6 by 6 human dissimilarity matrix of a tiger, a gorilla, a monkey, an eagle, a penguin and a frog, entries printed. Right: the two-dimensional map that MDS returns, a dot per animal with its thumbnail beside it, three map distances written on their segments: tiger to monkey 5.57, gorilla to monkey 1.15, eagle to penguin 3.72

Left: the 6 by 6 human dissimilarity matrix of a tiger, a gorilla, a monkey, an eagle, a penguin and a frog, entries printed. Right: the two-dimensional map that MDS returns, a dot per animal with its thumbnail beside it, three map distances written on their segments: tiger to monkey 5.57, gorilla to monkey 1.15, eagle to penguin 3.72

Open full-size figure
Assignment 1 data, ratings from Peterson, Abbott & Griffiths Cognitive Science 2018 · metric MDS, m = 2

Return to contents

Slide 19

Human similarity judgments

A long tradition in psychology: what we know of objects is arranged as points in a space, near when alike. That psychological space cannot be measured directly. What can be measured is a judgment. Peterson and colleagues asked people: how similar are these two images, from 0 to 10? Ten raters for every pair of 120 animal photographs.

Two tiger photographs A and B and a monkey photograph C, from the Assignment 1 image set

Two tiger photographs A and B and a monkey photograph C, from the Assignment 1 image set

Open full-size figure

Mean rating   A–B: 10.0  ·  A–C: 2.9  ·  B–C: 2.2  ·  one number per pair: a dissimilarity matrix, from which the space is to be recovered

Ratings: Peterson, Abbott & Griffiths Cognitive Science 2018, the data used in Assignment 1

Return to contents

Slide 20

Which measures are distances in the full sense?

Will it work? Yes, if the matrix is, or is well approximated by, a table of distances between points in that space. What a table of distances must satisfy, and which of our measures satisfy it:

Which measures are distances in the full sense? — table
Requirement Mathematical statement Euclidean deucd_{\mathrm{euc}} Cosine dcos⁡d_{\cos} Correlation dcorrd_{\mathrm{corr}}
Nonnegativity · no negative distances dik≥0d_{ik}\ge0 ✓ Yes ✓ Yes ✓ Yes
Identity · zero only for an image and itself dik=0  ⟺  𝐱(i)=𝐱(k)d_{ik}=0\iff\mathbf{x}^{(i)}=\mathbf{x}^{(k)} ✓ Yes ✗ No ✗ No
Symmetry · same both ways dik=dkid_{ik}=d_{ki} ✓ Yes ✓ Yes ✓ Yes
Triangle · no shortcut via a third image dil≤dik+dkld_{il}\le d_{ik}+d_{kl} ✓ Yes ✗ No ✗ No
A distance in the full sense (a metric)? All four ✓ Yes ✗ No ✗ No

✓ holds throughout the domain; ✗ can fail even when the measure is defined. On our images: 2.3 % of pixel triples break the triangle inequality under cosine, none in the trained late layer, and a brighter copy of an image sits at cosine distance 0 from it. You test this yourself: Assignment 1 bonus B1 · worked examples in the appendix.

Return to contents

Slide 21

Reading the map: the notation

The six-animal dissimilarity matrix on the left with three entries outlined, tiger–monkey, gorilla–monkey and eagle–penguin, each joined by an arrow to the segment between the corresponding two points on the map on the right, where every point is labelled y with its stimulus number and each segment is labelled d-hat with the pair's indices

The six-animal dissimilarity matrix on the left with three entries outlined, tiger–monkey, gorilla–monkey and eagle–penguin, each joined by an arrow to the segment between the corresponding two points on the map on the right, where every point is labelled y with its stimulus number and each segment is labelled d-hat with the pair's indices

Open full-size figure

Return to contents

Slide 22

Multidimensional scaling: the problem

Given a dissimilarity matrix 𝐃\textcolor{#3F7A6B}{\mathbf{D}} over NN stimuli, with entries dik\textcolor{#3F7A6B}{d_{ik}}.

Find coordinates, one point per stimulus,

𝐘∈ℝN×m,𝐲(i)∈ℝm,    m≪N\textcolor{#B5396B}{\mathbf{Y}} \in \mathbb{R}^{N\times m}, \qquad \textcolor{#B5396B}{\mathbf{y}^{(i)}} \in \mathbb{R}^{m}, \;\; m \ll N

such that the map distances d^ik\hat d_{ik} reproduce the measured dissimilarities dikd_{ik}:

d^ik  =  ∥𝐲(i)−𝐲(k)∥  ≈  dikfor all i<k\hat d_{ik} \;=\; \big\lVert \textcolor{#B5396B}{\mathbf{y}^{(i)}} - \textcolor{#B5396B}{\mathbf{y}^{(k)}} \big\rVert \;\approx\; \textcolor{#3F7A6B}{d_{ik}} \quad \text{for all } i<k

  • Usually m=2m=2 or 33, to look at it
  • Not required: 𝐗\mathbf{X}, the channels, a model of the stimuli. 𝐃\mathbf{D} is the only input.
  • When 𝐃\mathbf{D} comes from judgments, 𝐘\mathbf{Y} is the psychological space: each stimulus a point, judged dissimilarity the distance between points
  • Solved by lowering a mismatch, one small step at a time: next slides

Return to contents

Slide 23

Worked example: the colour circle falls out

Left: Ekman's 14 by 14 matrix of judged dissimilarity between monochromatic lights from 434 to 674 nanometres, with a wavelength colour strip along each edge. Right: the two-dimensional non-metric MDS solution, each light drawn in its own colour: a circle, violet next to red

Left: Ekman's 14 by 14 matrix of judged dissimilarity between monochromatic lights from 434 to 674 nanometres, with a wavelength colour strip along each edge. Right: the two-dimensional non-metric MDS solution, each light drawn in its own colour: a circle, violet next to red

Open full-size figure
Data: Ekman J. Psychol. 1954, via Michael Lee's collection · MDS, m = 2 · analysis after Shepard Psychometrika 1962

Return to contents

Slide 24

Stress-based MDS — the optimization view

Real dissimilarities are noisy; often no exact Euclidean configuration exists. Minimize the mismatch instead:

Stress(𝐘)  =  ∑i<k(d^ik−dik)2,d^ik=∥𝐲(i)−𝐲(k)∥\mathrm{Stress}(\textcolor{#B5396B}{\mathbf{Y}}) \;=\; \sum_{i<k}\Big( \hat d_{ik} - \textcolor{#3F7A6B}{d_{ik}} \Big)^{2}, \qquad \hat d_{ik}=\big\lVert \textcolor{#B5396B}{\mathbf{y}^{(i)}} - \textcolor{#B5396B}{\mathbf{y}^{(k)}} \big\rVert

  • Stress: zero for a perfect map, growing with every mismatched pair
  • Stress depends on the whole layout 𝐘\mathbf{Y}; the map is whatever makes it smallest
  • Start from any layout, move each point a little in the direction that lowers the stress, repeat until it stops falling: what sklearn does
  • How the direction is found: Theme 2. For now the solver is a black box
  • Local minima: different starts give different maps, so run several and keep the lowest stress (n_init)
Kruskal Psychometrika 1964 · usually reported in normalized form, as Kruskal stress-1

Return to contents

Slide 27

Watch it move

The six animals at convergence, the map MDS returns for the six-animal ratings: gorilla and monkey close together, eagle and penguin close together; stress 22

The six animals at convergence, the map MDS returns for the six-animal ratings: gorilla and monkey close together, eagle and penguin close together; stress 22

Open full-size figure
  • Converged: stress 2222, and no move lowers it further. This is the six-animal map from the start of this section
  • A different random start ends in the same map, possibly rotated or mirrored, or in a slightly worse one: run several starts, keep the lowest stress
Metric MDS (scikit-learn), six animals, human ratings · same axes on all three slides

Return to contents

Slide 28

Normalized stress: one number for the fit

Stress ⁣ ⁣− ⁣ ⁣1=∑i<k(d^ik−dik)2∑i<kd^ik 2\mathrm{Stress\!\! -\!\!1}=\sqrt{\frac{\sum_{i<k}(\hat d_{ik}-d_{ik})^2}{\sum_{i<k}\hat d_{ik}^{\,2}}}

Return to contents

Slide 29

Reading an MDS solution — three warnings

If a story about your map depends on which way is "up", it is a story about the plot.

Return to contents

Slide 30

Metric → non-metric: keep only the order

Two scatter plots of rated similarity, 0 to 10, against distance in the recovered map for the 120 animal ratings. Left, metric MDS: the fitted relation is a straight falling line, stress-1 0.31. Right, non-metric MDS: the fitted relation is a falling step curve, stress-1 0.23; the points bend because ratings pile up near 0

Two scatter plots of rated similarity, 0 to 10, against distance in the recovered map for the 120 animal ratings. Left, metric MDS: the fitted relation is a straight falling line, stress-1 0.31. Right, non-metric MDS: the fitted relation is a falling step curve, stress-1 0.23; the points bend because ratings pile up near 0

Open full-size figure

Return to contents

Slide 31

Explore: fifty classic similarity datasets

Data: Michael Lee's similarity-data collection (Ekman 1954, Rothkopf 1957, Rosenberg & Kim 1975, Romney et al. 1993, …) · open full screen

Return to contents

Slide 32

Shepard's universal law of generalization

Twelve small panels of generalization plotted against distance in psychological space — sizes, hues, phonemes, Morse code, human and pigeon data — each tracing the same exponential decay

Twelve small panels of generalization plotted against distance in psychological space — sizes, hues, phonemes, Morse code, human and pigeon data — each tracing the same exponential decay

Open full-size figure
  • Recover the psychological space by non-metric MDS (Shepard 1962, Kruskal 1964) from confusions or ratings; then plot measured generalization against distance in that space. Twelve datasets, one curve:

sik  =  e−c dik\textcolor{#3F7A6B}{s_{ik}} \;=\; e^{-\textcolor{#B5396B}{c}\, \textcolor{#3E6D8E}{d_{ik}}}

  • Sizes, lightnesses, spectral hues, vowels, consonants, Morse code, free-form shapes. Humans and pigeons.
  • Similarity is an exponential function of psychological distance, and the law is not circular: non-metric MDS fits rank order only, so the exponential shape was never built in.

Return to contents

Slide 33

Testing the law: line, exponential, Gaussian

Assignment 1 §2g fits three curves to the same (distance, similarity) pairs:

line: s=m d+cexponential: s=a e−bdGaussian: s=a e−bd2\text{line: } s = m\,d + c \qquad\quad \text{exponential: } s = a\,e^{-bd} \qquad\quad \text{Gaussian: } s = a\,e^{-bd^2}

  • The line is the baseline: constant-rate decay, no law to state
  • Any bend beats a line, so the fair rival is the Gaussian: same two parameters, but flat at d=0d=0 where the exponential is steepest
  • The test is at the origin: which shape do the pairs follow there?
  • Preview: in the space recovered from people the exponential wins; against a network's distances the picture changes, because a network's distance is a different ruler from psychological distance
Exponential and Gaussian similarity curves with the same value at zero but different initial slopes

Exponential and Gaussian similarity curves with the same value at zero but different initial slopes

Open full-size figure
Shepard Science 1987 · Assignment 1 §2g

Return to contents

Slide 34

Comparing two RDMs

Return to contents

Slide 35

Two matrices, one question

Two 120 by 120 dissimilarity matrices of the same animals, ordered by category with category boundaries drawn: human judgments on the left with strong light blocks along the diagonal, the ResNet-18 late layer on the right with weaker blocks in the same places

Two 120 by 120 dissimilarity matrices of the same animals, ordered by category with category boundaries drawn: human judgments on the left with strong light blocks along the diagonal, the ResNet-18 late layer on the right with weaker blocks in the same places

Open full-size figure
Assignment 1 data · ratings Peterson, Abbott & Griffiths Cognitive Science 2018 · ResNet-18 late layer, correlation distance · each on its own colour scale

Return to contents

Slide 36

An RDM is itself a vector

Two 120 by 120 dissimilarity matrices of the same animals side by side, human judgments and the ResNet-18 late layer, both with light blocks along the diagonal, weaker for the network

Two 120 by 120 dissimilarity matrices of the same animals side by side, human judgments and the ResNet-18 late layer, both with light blocks along the diagonal, weaker for the network

Open full-size figure
  • Each upper triangle laid out as a list: N(N−1)/2N(N-1)/2 numbers, one per pair. For N=120N = 120: 7,140 per system
  • Two lists of the same length: every measure from the first half of the lecture applies, one level up. Euclidean, cosine, correlation
  • No shared scale (ratings on 0–10; correlation distances on 0–2), so compare ranks: which pairs each system calls close
Numbers: Assignment 1 data · ResNet-18 late layer, correlation-distance RDM · human RDM D = 10 − S

Return to contents

Slide 37

Spearman's ρ: agree on the order, not the values

Three scatter plots of eight pairs' dissimilarities in two systems. Left: the second system's values are a bent but rising function of the first's, Pearson 0.95, Spearman 1. Middle: the same eight pairs plotted as ranks, a perfect straight line. Right: two neighbouring pairs swapped, Pearson 0.94, Spearman 0.98

Three scatter plots of eight pairs' dissimilarities in two systems. Left: the second system's values are a bent but rising function of the first's, Pearson 0.95, Spearman 1. Middle: the same eight pairs plotted as ranks, a perfect straight line. Right: two neighbouring pairs swapped, Pearson 0.94, Spearman 0.98

Open full-size figure
Toy values

Return to contents

Slide 38

Correlating two RDMs

Two systems, the same NN stimuli, two matrices: 𝐃human\textcolor{#3F7A6B}{\mathbf{D}^{\text{human}}} from the judgments, 𝐃model\textcolor{#3F7A6B}{\mathbf{D}^{\text{model}}} from a layer. Flatten each upper triangle into a list of N(N−1)/2N(N-1)/2 numbers and correlate the two lists:

ρ  =  corrrank(triu(𝐃human),  triu(𝐃model))\rho \;=\; \mathrm{corr}_{\text{rank}}\Big(\mathrm{triu}\big(\textcolor{#3F7A6B}{\mathbf{D}^{\text{human}}}\big),\; \mathrm{triu}\big(\textcolor{#3F7A6B}{\mathbf{D}^{\text{model}}}\big)\Big)

In practice: iu = np.triu_indices(N, k=1), then spearmanr(D1[iu], D2[iu]).

The operation has a name: representational similarity analysis, RSA

  • Upper triangle only. The diagonal is NN zeros in both matrices, agreement you did not measure; the lower half copies the upper. Either one inflates agreement or the sample size
  • Rank correlation (Spearman): the two systems share no scale, a distance between voxel patterns and one between unit activations being different kinds of number. The raw values are not comparable; their order is
  • The two matrices of a moment ago: ρ=0.42\rho = 0.42. Pearson rr on the same lists gives 0.570.57, a different question

Return to contents

Slide 39

What that number does and does not say

Return to contents

Slide 40

Why anyone cares: two species, one structure

Two 92-by-92 dissimilarity matrices, monkey IT from single-unit recordings and human IT from fMRI, both showing the same blue animate block and warm inanimate block with face and body sub-blocks; a vertical colour bar on the right runs from blue, similar, to red, dissimilar

Two 92-by-92 dissimilarity matrices, monkey IT from single-unit recordings and human IT from fMRI, both showing the same blue animate block and warm inanimate block with face and body sub-blocks; a vertical colour bar on the right runs from blue, similar, to red, dissimilar

Open full-size figure
  • Same 92 images; a monkey (IT single units) and a human (IT fMRI voxels): two measurements, two 92×9292\times92 matrices, one block structure
  • r = 0.49 between the two matrices (Pearson, as the paper reports it), computed as on the previous slides: the upper triangles, correlated
Kriegeskorte et al. Neuron 2008, fig. 1 · colour: percentile of 1 − r

Return to contents

Slide 41

Which system is most human-like? One wins

Bar chart of Spearman correlation between the human RDM and five candidate RDMs on the 120 animal images: pixels 0.01, early layer 0.10, middle layer 0.17, late layer 0.42, late layer with random weights 0.03; a dashed line at 0.82 marks the agreement between two independent groups of raters

Bar chart of Spearman correlation between the human RDM and five candidate RDMs on the 120 animal images: pixels 0.01, early layer 0.10, middle layer 0.17, late layer 0.42, late layer with random weights 0.03; a dashed line at 0.82 marks the agreement between two independent groups of raters

Open full-size figure
  • Pixels: 0.01. Nothing in the raw image orders the pairs the way people do
  • Depth helps: early 0.10 → middle 0.17 → late 0.42. The late layer wins
  • Same architecture, random weights: 0.03. The structure came from training, not from the wiring
  • Two independent rater groups agree at 0.82, the noise ceiling: the late layer reaches about half of it
Assignment 1 §2i numbers · pixels not in the assignment · ratings Peterson, Abbott & Griffiths Cognitive Science 2018

Return to contents

Slide 42

Do bigger models get closer? Some do

Bar chart of Spearman correlation with the human RDM for five models using pooled penultimate features: ResNet-50 0.52, ResNet-152 0.52, supervised ViT-B/16 0.26, ConvNeXt-B 0.20, self-supervised DINOv2 ViT-S 0.61 in green; dashed ceiling at 0.82

Bar chart of Spearman correlation with the human RDM for five models using pooled penultimate features: ResNet-50 0.52, ResNet-152 0.52, supervised ViT-B/16 0.26, ConvNeXt-B 0.20, self-supervised DINOv2 ViT-S 0.61 in green; dashed ceiling at 0.82

Open full-size figure
  • Deeper ResNets: 0.52. Depth alone buys little
  • Two newer, more accurate ImageNet models: 0.26 and 0.20. Bigger is not closer
  • A model trained without labels (DINOv2): 0.61, three quarters of the ceiling. How a model was trained matters more than how big it is
  • Transformers, self-supervised learning, why accuracy and human-likeness come apart: later in the course
Same 120 images and human RDM · pooled penultimate layer, correlation distance, Spearman · torchvision; DINOv2 ViT-S/14

Return to contents

Slide 43

Assignment 1: where you stand

Return to contents

Slide 44

Recap

Return to contents

Slide 45

Further reading for the research trail

Seven seeds, none required. Branch from any of them; pick a paper your neighbour has not picked.

Return to contents

Slide 46

Minute paper

Canvas submission: Canvas → Minute papers → Minute Paper 3 (access code read out in class)

Write three brief points in your own words:

  1. One sentence: when can two response vectors have cosine similarity 1 (cosine dissimilarity 0) and yet a large Euclidean distance?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Return to contents

Slide 47

Appendix · A difference defines a direction

Subtraction: the displacement from image 1 (tiger) to image 2 (elephant):

𝐱(2)−𝐱(1)=(128.3−72.719.2−40.6)=(55.6−21.4)\mathbf{x}^{(2)}-\mathbf{x}^{(1)}=\begin{pmatrix}128.3-72.7\\19.2-40.6\end{pmatrix}=\begin{pmatrix}55.6\\-21.4\end{pmatrix}

𝐱(2)−𝐱(1)\mathbf{x}^{(2)}-\mathbf{x}^{(1)} is itself a vector: from the tiger's point to the elephant's.

The two-feature plot, mean intensity against contrast, with three arrows: x superscript 1 to the tiger's point, x superscript 2 to the elephant's point, and their difference drawn from the tiger's point to the elephant's

The two-feature plot, mean intensity against contrast, with three arrows: x superscript 1 to the tiger's point, x superscript 2 to the elephant's point, and their difference drawn from the tiger's point to the elephant's

Open full-size figure

Return to contents

Slide 48

Appendix · Computing an RDM from a response matrix

import numpy as np from scipy.spatial.distance import pdist, squareform R = np.load("responses.npy") # (N stimuli, n channels) print(R.shape) # (120, 4096) units or (120, 482) neurons # pdist returns the upper triangle: N(N-1)/2 numbers, the pairs i < k d_vec = pdist(R, metric="correlation") # 1 - Pearson r, pattern to pattern D = squareform(d_vec) # (N, N): symmetric, zero diagonal assert np.allclose(D, D.T) and np.allclose(np.diag(D), 0)

Return to contents

Slide 49

Appendix · Why network responses contain zeros

ReLU (rectified linear unit): input zjz_j, output aj=max⁡(0,zj)a_j=\max(0,z_j).

The tiger image with all 4096 stored late ResNet-18 activations, 1843 exact zeros marked below the baseline, and the ReLU response curve

The tiger image with all 4096 stored late ResNet-18 activations, 1843 exact zeros marked below the baseline, and the ReLU response curve

Open full-size figure

Return to contents

Slide 50

Appendix · Shared zeros can change correlation

The tiger and gorilla photographs beside their full measured ResNet-18 activation vectors, with 879 shared-zero entries marked at matching positions

The tiger and gorilla photographs beside their full measured ResNet-18 activation vectors, with 879 shared-zero entries marked at matching positions

Open full-size figure

Return to contents

Slide 51

Appendix · What training does to the RDM

Two 120-by-120 correlation-distance RDMs of the late ResNet-18 layer, images ordered by category with dark lines at the category boundaries, each on its own colour scale: with random weights the entries are noise with no block structure; with ImageNet-trained weights lighter blocks sit on the diagonal, one per category

Two 120-by-120 correlation-distance RDMs of the late ResNet-18 layer, images ordered by category with dark lines at the category boundaries, each on its own colour scale: with random weights the entries are noise with no block structure; with ImageNet-trained weights lighter blocks sit on the diagonal, one per category

Open full-size figure
  • Same 120 images, same architecture, two sets of weights
  • Random weights: every pair alike, mean 1−r=0.151-r = 0.15, the same within and between categories. No blocks, whatever the colour scale
  • Trained: pairs far apart, mean 0.900.90; within a category 0.820.82, between 0.920.92. The blocks on the diagonal are the categories
  • The RDM is a property of the weights, not of the architecture
Correlation distance on the 4,096 stored late-layer activations · random weights, seed 1291 · each panel on its own colour scale

Return to contents

Slide 52

Appendix · Training changes the representation

Supervised training: four labelled training photographs (frog, eagle, elephant, penguin) pass one at a time through a network to produce class predictions; comparing the prediction with the true label gives a loss, which guides weight updates and changes internal representations

Supervised training: four labelled training photographs (frog, eagle, elephant, penguin) pass one at a time through a network to produce class predictions; comparing the prediction with the true label gives a loss, which guides weight updates and changes internal representations

Open full-size figure

Return to contents

Slide 53

Appendix · What should a dissimilarity satisfy?

Similarity ratings come from people, not from a formula. A network's RDM inherits its properties from the measure dd; a rater is free to answer anything. Before we recover a space from such numbers, what must they satisfy? For tiger, penguin, gorilla (i,k,li, k, l):

The first three make 𝐃\mathbf{D} a dissimilarity matrix. The fourth is what points in a space obey: without it, no map reproduces the entries exactly. Which measures satisfy which: the distance table in the lecture. Whether people do: the assignment bonus, and the Tversky slides later in this appendix

Return to contents

Slide 54

Appendix · Running MDS in three lines

from sklearn.manifold import MDS mds = MDS(n_components=2, # m = 2, because we want to look at it metric="precomputed", # feed D directly, NOT the raw R n_init=8, normalized_stress="auto", random_state=0) Y = mds.fit_transform(D) # (N, 2) coordinates - one row per stimulus print(Y.shape, mds.stress_) # residual mismatch: lower is better # non-metric (ordinal) MDS - fits the RANK ORDER of the dissimilarities mds_nm = MDS(n_components=2, metric="precomputed", metric_mds=False, n_init=8, normalized_stress="auto", random_state=0) Y_nm = mds_nm.fit_transform(D)

Return to contents

Slide 55

Appendix · Fitting a curve is a regression

  • Each point is one pair of images: its map distance dd and its rated similarity ss
  • Fit = choose the curve's parameters (aa, bb) so the curve passes as close as possible to the points: smallest sum of squared residuals s−s^s-\hat s. Same idea as fitting a line, the curve just bends
  • R2R^2 = the share of the spread in ss the curve accounts for: 11 perfect, 00 no better than the flat mean sˉ\bar s

R2=1−∑(s−s^)2∑(s−sˉ)2R^2=1-\frac{\sum(s-\hat s)^2}{\sum(s-\bar s)^2}

  • Fit the three curves to the same points, compare their R2R^2; the highest wins, the line is the baseline
Scatter of forty distance–similarity pairs with a fitted exponential curve through them, a vertical residual segment from every point to the curve, a dotted horizontal line at the mean similarity, and R squared 0.90 in the title

Scatter of forty distance–similarity pairs with a fitted exponential curve through them, a vertical residual segment from every point to the curve, a dotted horizontal line at the mean similarity, and R squared 0.90 in the title

Open full-size figure
Illustrative points, not assignment data · Assignment 1 §2g: np.polyfit for the line, curve_fit for the two curves

Return to contents

Slide 56

Appendix · Human reliability sets the comparison scale

Split-half reliability: correlate two independent sets of ratings for the same pairs. That agreement is the noise ceiling: how well the measurement predicts itself, the bar no model should beat.

Appendix · Human reliability sets the comparison scale — table
Comparison Statistic
Rating half 1 versus rating half 2 Reliability of the human measurement
Network distances versus human dissimilarities Representational alignment

Return to contents

Slide 57

Appendix · Hierarchical clustering — the idea

Left: the two-dimensional MDS map of six animals, tiger, gorilla, monkey, eagle, penguin, frog, thumbnails beside the points. Right: the dendrogram built from the same six-by-six human dissimilarity matrix with average linkage, thumbnails as leaves; gorilla and monkey join at 1.8, eagle and penguin at 5.2, the frog joins last

Left: the two-dimensional MDS map of six animals, tiger, gorilla, monkey, eagle, penguin, frog, thumbnails beside the points. Right: the dendrogram built from the same six-by-six human dissimilarity matrix with average linkage, thumbnails as leaves; gorilla and monkey join at 1.8, eagle and penguin at 5.2, the frog joins last

Open full-size figure
  • Group the stimuli recursively and draw it as a dendrogram: leaves are stimuli, the height of a join is the dissimilarity at which two groups merged
  • Agglomerative (bottom-up): start with NN singletons, merge the closest pair, repeat. What the Assignment 1 bonus uses
  • Input is again just 𝐃\mathbf{D}: no coordinates, so judged similarity works as well as firing rates
The six animals of the MDS slides, human ratings · average linkage

Return to contents

Slide 58

Appendix · Agglomerative clustering and linkage

  1. Compute all pairwise distances: the matrix 𝐃\mathbf{D}.
  2. Merge the two closest clusters.
  3. Update the distances from the new cluster to all others, using a linkage rule.
  4. Repeat until one cluster remains.

In practice: scipy.cluster.hierarchy.linkage(squareform(D), method="average").

Step 3 is the whole design decision:

  • Single: distance to the closest member. Chains.
  • Complete: distance to the farthest member. Compact, equal-size clusters.
  • Average: mean over all cross-pairs. The usual compromise.
  • Ward: merge the pair that increases within-cluster variance least. Needs coordinates, not just 𝐃\mathbf{D}.

Return to contents

Slide 59

Appendix · Linkage changes the answer

Three dendrograms of the same 120 animals from the human dissimilarity matrix, single, complete and average linkage, each leaf marked with its taxon colour: single linkage grows one long chain, complete linkage makes compact groups, average linkage sits between

Three dendrograms of the same 120 animals from the human dissimilarity matrix, single, complete and average linkage, each leaf marked with its taxon colour: single linkage grows one long chain, complete linkage makes compact groups, average linkage sits between

Open full-size figure
All 120 Assignment 1 animals, human ratings · Peterson, Abbott & Griffiths Cognitive Science 2018

Return to contents

Slide 60

Appendix · How many clusters? Cut the tree

Left: the average-linkage dendrogram of the 120 animals with three dashed cut heights labelled 2, 4 and 8 clusters. Right: the MDS map of the same animals coloured by the clusters each cut produces; the 8-cluster cut nearly matches the eight taxa

Left: the average-linkage dendrogram of the 120 animals with three dashed cut heights labelled 2, 4 and 8 clusters. Right: the MDS map of the same animals coloured by the clusters each cut produces; the 8-cluster cut nearly matches the eight taxa

Open full-size figure
Same 120 animals and ratings · average linkage · map: metric MDS

Return to contents

Slide 61

Appendix · Clustering always returns a tree

Dendrogram of over a thousand patients clustered on fMRI connectivity, cut by a dashed line into four colored groups proposed as depression biotypes

Dendrogram of over a thousand patients clustered on fMRI connectivity, cut by a dashed line into four colored groups proposed as depression biotypes

Open full-size figure

the dashed line is the cut; four colors are the four proposed groups

  • Over a thousand patients with depression, clustered on resting-state fMRI connectivity, cut into four "biotypes", each one claimed to predict who responds to which treatment.
  • Same algorithm as the last three slides. The arithmetic is not in question.
  • The biotypes did not replicate. Reanalysis found no evidence of genuine cluster structure in the data, and a tree comes back from any matrix, including one with no groups in it.
  • Whether the branches name something real is a separate claim, needing separate evidence.

Return to contents

Slide 62

Appendix · MDS and clustering are different hypotheses

Appendix · MDS and clustering are different hypotheses — table
MDS Hierarchical clustering
Output continuous coordinates 𝐘\mathbf{Y} a nested tree
Structure assumed a smooth space — items vary along dimensions discrete categories — items belong to groups
Fits well perceptual continua (hue, pitch, size) taxonomies, category structure
Free parameter number of dimensions mm linkage rule + where to cut
Fails on superordinate terms, deep hierarchies continuous variation

Run both. Where they disagree is the interesting question.

Return to contents

Slide 63

Appendix · Human similarity breaks all three rules

Every method so far assumes dissimilarity is a distance. Tversky, with data: human judgments violate all three axioms.

Return to contents

Slide 64

Appendix · Violations of minimality and symmetry

Minimality, dik≥dii=0d_{ik} \ge d_{ii} = 0

  • Rothkopf's Morse-code study: participants judged some different pairs "same" more reliably than they judged an identical pair "same".
  • Self-similarity is not constant across items: familiar signals are recognized as themselves more reliably than unfamiliar ones.

Symmetry, dik=dkid_{ik} = d_{ki}

  • A camel is more similar to a horse than a horse is to a camel.
  • An ellipse is more similar to a circle than a circle is to an ellipse.
  • The less typical / less salient item is judged more similar to the prototype than the reverse. Direction of comparison matters.
Rothkopf 1957 (Morse-code confusions) · Tversky Psychol. Rev. 1977

Return to contents

Slide 65

Appendix · Triangle inequality and hierarchy, violated

Triangle of three objects — lamp, moon, ball — where each neighboring pair shares a feature but the far pair shares none

Triangle of three objects — lamp, moon, ball — where each neighboring pair shares a feature but the far pair shares none

Open full-size figure

lamp ~ moon (both luminous) · moon ~ ball (both round) · lamp ≁ ball

  • Triangle inequality: dik+dkl≥dild_{ik} + d_{kl} \ge d_{il}. Two items can each be similar to a third on different features, and not at all similar to each other.
  • Nearest-neighbor restriction: in an mm-dimensional Euclidean space, a point can be the nearest neighbor of only a bounded number of others. Conceptual hierarchies break that bound: fruit is the nearest neighbor of many fruits at once, but an MDS map can only put it near two or three.
Tversky Psychol. Rev. 1977 · Tversky & Hutchinson Psychol. Rev. 1986

Return to contents

Slide 66

Appendix · Comparing RDMs: what can go wrong

One correlation is a point estimate. Papers add three things; each has a catch.

Return to contents

Slide 67

Appendix · Emotion has a geometry too

Map of 2,185 emotion videos, each a letter colored by its dominant category: 27 categories form regions joined by smooth gradients rather than gaps

Map of 2,185 emotion videos, each a letter colored by its dominant category: 27 categories form regions joined by smooth gradients rather than gaps

Open full-size figure

each letter is one video, colored by its dominant category

  • 2,185 short videos rated by hundreds of viewers for what they felt. No stimulus dimensions, no recordings, judgments are the only input.
  • The structure that comes back needs 27 categories, not the textbook two, and they are joined by smooth gradients rather than gaps: anxiety shades into fear, fear into horror.

The 2-D layout here uses a nonlinear cousin of MDS, taught in a later lecture.

Return to contents