Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 1 · Theme 1: Representational spaces

Representations, spaces & metrics

Tuesday, September 15 · Fall 2026

Reader-friendly version — lecture content with figure descriptions and data tables

Before we start

  • Minute paper 1: thank you
    • Two wishes: build and train models yourselves; understand where brains and AI agree and differ. Both start today
    • Two worries: time, and rusty or absent Python. The recitations exist for that
  • Recitation tonight, Tuesday 6:30–7:50 PM (room on Canvas)
    • TA-led and optional
    • Tonight: Python bootcamp
    • Strongly encouraged if you have not programmed much, or not in Python
  • Slides and notes on Canvas
    • No need to take notes
    • Reader-friendly version linked from the title slide

What does a brain or network represent?

Three tigers. Which two are most alike? Two systems, two answers. What does each look at?

Three tiger photographs A, B and C from the assignment image set. Below them, two rows of checks and crosses: one system pairs A with B; another pairs B with C. A and B are near mirror images of one another; B and C are both walking with the head lowered
Natural photographs from Peterson, Abbott & Griffiths Cognitive Science 2018 · ImageNet photographs; the image set you will use in Assignment 1

A representation stands for something

“This is not a pipe.”

  • A depiction is not the object
  • It stands for it: a representation

Magritte’s painting depicts a smoking pipe above the French inscription Ceci n’est pas une pipe, meaning This is not a pipe

The representation in an image

Color images: three values per pixel

Vectorization: the 28 × 28 array is read row by row into a single list of 784 numbers, a vector. Distances and averages are defined on vectors.

Grayscale: one value per pixel

Vectorization: the 28 × 28 array is read row by row into a single list of 784 numbers, a vector. Distances and averages are defined on vectors.

One image, 784 measured values

Vectorization: the 28 × 28 array is read row by row into a single list of 784 numbers, a vector. Distances and averages are defined on vectors.

The representation in a brain

An image in human visual cortex

  • fMRI: blood-oxygenation signal per voxel
    • ~2 mm cube, thousands of neurons
  • Localizer scan: finds the regions
  • Each voxel's response to the dog: one entry of the vector
Top: cortical surfaces with early visual cortex in red and the lateral occipital complex in green, found by a localizer scan. Bottom: a dog photograph beside the measured responses of all 210 early visual and 152 LOC voxels of participant CSI1 to that photograph. Surface colors identify regions, not responses.

Visual selectivity, heard as spikes

Hubel & Wiesel · cat V1 · Historical context: Wiesel, iBiology

One electrode at a time, or hundreds of sites at once

A single metal microelectrode, a thin needle, mounted in a holder that advances it into the brain

A single microelectrode: one site

The Neuropixels probe: a schematic of the shank tip with two staggered columns of square recording sites 20 micrometres apart on a 70-micrometre-wide shank; a scanning electron micrograph of the tip; and the whole device, a 1 centimetre shank rising into a base, flex cable and headstage

Neuropixels: 960 sites on one shank

  • Bao et al. 2020, the monkey data of this course: a tungsten microelectrode advanced into IT, one neuron at a time; 482 neurons over many sessions
  • Neuropixels (Jun et al. 2017): 960 sites, 384 read at once, 20 µm apart
    • One neuron appears on several sites; that pattern is what spike sorting uses
    • One insertion: hundreds of neurons, several areas
  • Same kind of number either way: a firing rate per neuron
Microelectrode photograph: Alpha Omega Engineering · Jun et al. Nature 2017 · Fig. 1 · Steinmetz et al. Science 2021 (Neuropixels 2.0: 5,120 sites)

Recording neural populations

Montage of neural recording technologies: (a) a rodent head in cross-section showing scalp EEG screws, an ECoG grid on the cortical surface, and a microelectrode entering the tissue; (b) the signal chain and example traces from EEG, ECoG, local field potential, and extracellular action potentials; (c) a flexible ECoG grid beside a coin; (d) a Utah array of one hundred silicon needles; (e) a silicon probe shank with dense recording sites; (f) shanks carrying micro-LEDs; (g) a micro-ECoG surface array over blood vessels; (h) a probe with on-board CMOS electronics; (i) a fluidic probe with electrodes and channels; (j) a small three-dimensional array with flexible interconnect
a  Scales of access: scalp EEG → ECoG on the cortical surface → microelectrodes in the tissue
b  Signal chain: EEG and ECoG average many neurons; only electrodes in the tissue resolve spikes
c  ECoG grid — a flexible surface array (coin for scale)
d  Utah array — 100 silicon needles, 400 µm apart; used in human BCI
e  Silicon probe — dozens of recording sites along one shank (Neuropixels lineage)
f  Optoelectrode — micro-LEDs on the shank for optogenetics
g  Micro-ECoG — surface array fine enough for single cells
h  CMOS probe — amplifiers built into the probe base
i  Fluidic probe — electrodes plus channels for drug delivery
j  3D active array — flexible interconnect

From image to response: present an image

  • Image shown for 250 ms (Bao et al. 2020); signal travels retina → inferotemporal cortex (IT)

From image to response: record and sort spikes

  • Spike: a brief electrical pulse from a neuron. An electrode measures it; spike sorting assigns each spike to one neuron

From image to response: estimate a firing rate

  • Count spikes in a window, divide by its length: 7 / 0.160 s = 43.75 spikes/s. One number per neuron, per image, per trial

One image, several trials: average the rates

  • The same image is shown several times; each trial gives a different count. The response we keep is the mean rate across trials
Six repeated presentations of the tiger: six spike counts in the 60 to 220 ms window (5, 8, 6, 9, 7, 7), six rates from 31.25 to 56.25 spikes per second, and their mean, 43.75 spikes per second
Illustrative counts · Bao et al. Nature 2020 averaged 4–8 presentations per image

An object in a neural population

  • Macaque inferotemporal (IT) cortex: high-level visual cortex for objects
  • One image → one firing rate per neuron → a pattern across 482 neurons

Left: Bao Fig. 4c coronal MRI showing functional IT patches, and a cropped Extended Data 8a recording-site detail with a 5 mm scale bar. Right: the framed cat stimulus above a bar plot of firing rates in spikes per second for all 482 units, 479 available values and three unavailable entries marked below the axis

Bao et al. Nature 2020 · IT patches: Fig. 4c · recording sites: Extended Data Fig. 8a · responses: the Assignment 1 data

The representation in an artificial neural network

Visual cortex and convolutional networks

Macaque visual areas above a hierarchical CNN with stacks of spatial feature maps; dashed green arrows mark proposed model–brain correspondences

From AlexNet to ResNet

ImageNet top-5 accuracy by year: AlexNet 83.6 percent with 8 layers in 2012, ZFNet 88.3 with 8 layers in 2013, GoogLeNet 93.3 with 22 layers and VGG-16 92.7 with 16 layers in 2014, ResNet-152 96.4 with 152 layers in 2015
A mosaic of hundreds of small ImageNet photographs of animals, objects, vehicles and scenes
  • ImageNet: 14 M photographs, 22 k nouns
  • Benchmark: 1.28 M training images, 1,000 categories
    • Top-5: label among the five best guesses
  • Deeper every year. These models feed Assignment 1
  • Browse: the 1,000 categories · LENS, what a ResNet-50 learned, concept by concept.

Early, middle and late representations

  • Layer: a processing stage · activation: a unit's output, one feature · channel: a map of activations
  • ResNet-18 (Assignment 1), six most active channels per stage: “edges and stripes” → “parts and textures” → “where the tiger is”. Rough descriptions, not what the channels compute
The tiger photograph and three rows of six activation maps from ResNet-18: early-stage 56 by 56 maps that trace edges and stripes, middle-stage 14 by 14 maps that pick out parts and textures, and late-stage 7 by 7 maps that light up over the tiger's body
He et al. CVPR 2016 · LENS (Serre lab; what a ResNet-50 learned, class by class — tiger)

An object in ResNet-18

Representations as measurements

A representation: a pattern that carries information about a stimulus, a property or a state. Each measurement so far gives one such pattern as a list of numbers.

Notation

Symbol Meaning
x\mathbf{x} (bold lowercase) a vector: the whole list of numbers for one object
xjx_j (italic, subscript) its jjth component, a single number
x(i)\mathbf{x}^{(i)}, x(i,r)\mathbf{x}^{(i,r)} the vector for image ii; the same image on repetition rr
X\mathbf{X} (bold uppercase) a matrix: many vectors stacked, one image per row
xijx_{ij} its entry in row ii, column jj: feature jj of image ii
x⊤\mathbf{x}^{\top} the transpose: the same numbers laid out as a row instead of a column
∥x∥\lVert\mathbf{x}\rVert the length of the vector
DD, NN the number of measurements per object; the number of objects

Two measurements, two coordinates

x1x_1: mean intensity · x2x_2: RMS contrast (standard deviation of the intensities)

The tiger photograph used in the pixel explorer

One vector, two components:

x=(x1, x2)≈(72.7, 40.6)for the tiger.\mathbf{x}=(x_1,\,x_2)\approx(72.7,\,40.6)\quad\text{for the tiger.}

The two axes define a space; the tiger is one point in it.

Square plot of RMS contrast against mean intensity; the tiger's point at 72.7 and 40.6 is highlighted with its thumbnail beside it; five other images appear as grey dots

Two measurements, two coordinates

x1x_1: mean intensity · x2x_2: RMS contrast (standard deviation of the intensities)

The gorilla photograph used in the pixel explorer

One vector, two components:

x=(x1, x2)≈(88.4, 44.8)for the gorilla.\mathbf{x}=(x_1,\,x_2)\approx(88.4,\,44.8)\quad\text{for the gorilla.}

A different image, a different point.

The same plot with the gorilla's point at 88.4 and 44.8 highlighted and its thumbnail beside it

Two measurements, two coordinates

x1x_1: mean intensity · x2x_2: RMS contrast (standard deviation of the intensities)

One vector, two components:

x=(x1, x2)\mathbf{x}=(x_1,\,x_2)

Every image is a point in the same space. Comparing representations therefore reduces to geometry: distances and angles between points.

The same plot with all six images shown as thumbnails at their points: tiger, gorilla, eagle, frog, penguin and elephant

A vector is an arrow from the origin

Same two numbers: a point, or an arrow from the origin to it.

x(1)≈(72.740.6)(the tiger)\mathbf{x}^{(1)}\approx\begin{pmatrix}72.7\\40.6\end{pmatrix}\quad\text{(the tiger)}

  • Vectors are columns; inline (x1, x2)(x_1,\,x_2) is shorthand
  • Superscript (1)(1): image number 1
  • Arrows: what lets us add and subtract vectors, measure similarity, dissimilarity and distance, and so compare representations
The same square plot of RMS contrast against mean intensity with a single arrow drawn from the origin to the tiger's point at 72.7 and 40.6, labelled x superscript 1; the other five images are grey dots

Extending beyond two dimensions

x(i)=(x1(i)x2(i)⋮xD(i))∈RD\mathbf{x}^{(i)}=\begin{pmatrix}x_1^{(i)}\\x_2^{(i)}\\\vdots\\x_D^{(i)}\end{pmatrix}\in\mathbb{R}^{D}

  • x(i)\mathbf{x}^{(i)}: representation of image ii, one number per measurement
  • DD: the dimension, the number of measurements
    • RD\mathbb{R}^{D}: all such lists
  • DD measurements define a DD-dimensional space; same object, different measurement → different space
Measurement DD
Two image features 2
Resampled grayscale image 784
Recorded IT population 482
Human early visual cortex (voxels) 210
Network layer (sampled units) 4,096

How reliable is one response?

Brain responses are noisy

  • Same image three times → three different vectors
  • Correlation rr: agreement, 1 = identical, 0 = unrelated (defined next lecture)
  • r12=r_{12}= 0.17, r13=r_{13}= −0.01, r23=r_{23}= 0.03; median over 112 repeated images 0.14
  • Averaging keeps what repeats
The bullfrog photograph shown three times to participant CSI1 Four rows of bar plots over 210 early-visual-cortex voxels: three single presentations of the same photograph to participant CSI1, each visibly different, and below them in red the mean of the three
Data: Chang et al. Scientific Data 2019 · BOLD5000 CSI1, image n01641577_1229 (bullfrog), early visual region, average of two post-stimulus time points

Electrophysiology is noisy too; averaging works

  • Monkey IT, 168 sites, same image 51 times
  • Two single trials: r=r= 0.17
  • Split-half correlation: mean of 25 trials vs mean of the other 25, r=r= 0.69
  • Every response vector in this course is such a mean
The stimulus: a rendered lioness on a natural background, one of the 3,200 images of Majaj et al. 2015 Four rows of bar plots over 168 IT sites: three single presentations of the same image, each different, and below them in red the mean over all 51 presentations, much smaller and smoother
Data: Majaj, Hong, Solomon & DiCarlo J. Neurosci. 2015, public Brain-Score assembly; responses are normalised per site, 70–170 ms window

Average repeated responses to one image

Repetition rr of RR gives x(i,r)∈RD\mathbf{x}^{(i,r)}\in\mathbb{R}^{D}. Three neurons, three repetitions (illustrative spikes/s):

x(i,1)=(251),x(i,2)=(432),x(i,3)=(343)⟹x(i)=1R∑r=1Rx(i,r)=(342)\mathbf{x}^{(i,1)}=\begin{pmatrix}2\\5\\1\end{pmatrix},\qquad \mathbf{x}^{(i,2)}=\begin{pmatrix}4\\3\\2\end{pmatrix},\qquad \mathbf{x}^{(i,3)}=\begin{pmatrix}3\\4\\3\end{pmatrix} \qquad\Longrightarrow\qquad \mathbf{x}^{(i)}=\frac{1}{R}\sum_{r=1}^{R}\mathbf{x}^{(i,r)}=\begin{pmatrix}3\\4\\2\end{pmatrix}

  • Mean component by component: xj(i)=1R∑rxj(i,r)x_j^{(i)}=\tfrac{1}{R}\sum_r x_j^{(i,r)}
    • Neurons never averaged into one another
  • x(i)\mathbf{x}^{(i)} denotes image ii in what follows
  • Assignment 1 gives you the averages, not the trials: one vector of 482 mean rates per image (1,224 × 482)

Column vectors and row vectors

Column: DD rows, one column. Transpose ⊤^{\top}: the same numbers as a row.

x(i)=(342)∈R3×1→ ⊤ (x(i))⊤=(342)∈R1×3\mathbf{x}^{(i)}=\begin{pmatrix}3\\4\\2\end{pmatrix}\in\mathbb{R}^{3\times 1} \qquad\xrightarrow{\ \top\ }\qquad \big(\mathbf{x}^{(i)}\big)^{\top}=\begin{pmatrix}3&4&2\end{pmatrix}\in\mathbb{R}^{1\times 3}

  • Same numbers, same order; only the orientation changes
  • Rows: how we stack many images

Many images form a response matrix

Stack NN row vectors: one matrix for the whole image set.

X=((x(1))⊤(x(2))⊤⋮(x(N))⊤)∈RN×D\mathbf{X}=\begin{pmatrix}\big(\mathbf{x}^{(1)}\big)^{\top}\\\big(\mathbf{x}^{(2)}\big)^{\top}\\\vdots\\\big(\mathbf{x}^{(N)}\big)^{\top}\end{pmatrix}\in\mathbb{R}^{N\times D}

  • Row ii: image ii. Column jj: feature jj
  • xij=xj(i)x_{ij}=x_j^{(i)}: feature jj of image ii, one cell
  • Here: N=6N=6 images, 12 of D=4,096D=4{,}096 activations shown

X=(x11x12⋯x1Dx21x22⋯x2D⋮⋮⋮xN1xN2⋯xND)\mathbf{X}=\begin{pmatrix}x_{11}&x_{12}&\cdots&x_{1D}\\x_{21}&x_{22}&\cdots&x_{2D}\\\vdots&\vdots&&\vdots\\x_{N1}&x_{N2}&\cdots&x_{ND}\end{pmatrix}

Heatmap of six named animal images by twelve late-layer ResNet-18 activations; each cell is one entry x i j

Comparing representations across systems

Pixels, voxels, neurons and network units cannot be compared entry by entry: the vectors differ in length and in units. Their pairwise dissimilarities can. The RDM records them.

An RDM compares every pair of images

Representational dissimilarity matrix (RDM): every pairwise comparison.

D∈RN×N,dik=d(x(i),x(k))\mathbf{D}\in\mathbb{R}^{N\times N},\qquad d_{ik}=d\bigl(\mathbf{x}^{(i)},\mathbf{x}^{(k)}\bigr)

  • Rows and columns: images, here ordered by category. Each entry: one comparison
  • dd: the dissimilarity measure. It decides which differences count
  • Entries are not arbitrary: dii=0d_{ii}=0, dik=dkid_{ik}=d_{ki}, and for a distance, no shortcut beats the direct route (next lecture)
A 120 by 120 representational dissimilarity matrix of the Assignment 1 images from the trained ResNet-18 late layer, correlation distance, images ordered by animal category with category boundaries drawn; light blocks along the diagonal show that primates, carnivores and hoofed mammals are more alike within their category than across

Want to see inside a network before we get there?

Two lectures on convolutional networks and one on training are coming. Previews:

Not required. Ten minutes well spent: the explainer and the playground.

Further reading for the research trail

Six starting points, none required. Your trail paper can come from any lecture; these are seeds, not a list to choose from. Peterson et al. 2018 and Bao et al. 2020 are the assignment papers: fine to read, but branch out.

How to read for the trail

  • Skim first: abstract, figures, last paragraph of the discussion. Then decide whether to read
  • Ask what would change the conclusion
    • Is a control missing?
    • Would another network, another animal, another image set give the same result?
    • Is the result trivial given the method?
  • Expect it to be slow at first. It gets faster; the questions become habits
  • Then post: four to six sentences on Ed, one reply to a classmate

Minute paper

Canvas submission: Canvas → Minute papers → Minute Paper 2 (access code read out in class)

Write three brief points in your own words:

  1. One sentence: when can two response vectors have cosine similarity 1 (cosine dissimilarity 0) and yet a large Euclidean distance?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Reference: computing an RDM

import numpy as np
from scipy.spatial.distance import pdist, squareform

X = np.array([[1., 1.],      # pattern A
              [4., 4.],      # pattern B
              [1., -1.]])    # pattern C

for metric in ("euclidean", "cosine"):
    D = squareform(pdist(X, metric=metric))
    print(metric, np.round(D, 2))

pdist: the unique pairs. squareform: the symmetric matrix.

Appendix · Scaling image intensity

Scaling: x↦c x\mathbf{x}\mapsto c\,\mathbf{x}. Same 784 entries, only their size changes.

Appendix · Averaging and blending images

Blending: z=(1−λ) x(1)+λ x(2)\mathbf{z}=(1-\lambda)\,\mathbf{x}^{(1)}+\lambda\,\mathbf{x}^{(2)}, entry by entry. Done to labels too, this is mixup.

- From minute paper 1: many of you want to build and train models yourselves, and many want to understand how brains and AI relate; the two most common worries were time and programming. Lectures own the concepts; recitations own the code - Recitations meet on Tuesdays, 6:30–7:50 PM; the room is posted on Canvas. The first sessions are a Python bootcamp, optional but strongly encouraged if Python is new to you - Slides and presenter notes are posted on Canvas after each lecture; a reader-friendly version with figure descriptions is linked from the title slide

- These photographs are examples of the natural images you will use in Assignment 1: the 120 animal photographs, drawn from ImageNet, of Peterson, Abbott and Griffiths (Cognitive Science 2018), for which human similarity ratings exist - One system pairs A with B; another pairs B with C. A and B are near mirror images of one another (the outline); B and C share the same body orientation. The two are hypothetical classifiers that could reach the same overall accuracy while attending to very different features of the images - Similarity is therefore not a property of the pixels alone; it depends on the system doing the comparing, whether a human observer, a neural population or a network, and on what that system attends to - Similarity judgments and measurements of this kind are a standard tool in cognitive science and in AI for characterizing the representations of different systems and comparing them; this week introduces the tool

- What is a representation? Without going far into philosophy, Magritte's painting is the classical starting point - Magritte's painted pipe cannot be smoked: an object and its depiction are different things - The tiger exists as an animal, as a photograph of the animal, and as the array of numbers that encodes the photograph; the three are related but not interchangeable - A representation, in this course, is something that stands for an object without being it. The painting is the simplest case; the numbers an image, a brain or a network hold about the tiger are the cases we study - René Magritte, The Treachery of Images, 1929, LACMA (© C. Herscovici / Artists Rights Society, New York; photo © Museum Associates/LACMA)

- Three representations of the same image will be examined in turn: in the image itself, in the activity of a brain viewing it, and in the activations of an artificial neural network processing it - The image is the simplest case. A digital image is an array of numbers and nothing more. It is legible to us only because a display converts those numbers into light: at every location, three numbers set the amounts of red, green and blue emitted by the projector. The next three slides examine that array

- A photograph is already a set of measurements: a color image stores three values per pixel, one per channel (red, green, blue); at 224 × 224 pixels that is 224 × 224 × 3 = 150,528 values - Eight bits give 2^8 = 256 levels, so unsigned 8-bit values run from 0 to 255; RGB commonly uses 8 bits per channel, 24 bits per pixel - The 0–255 scale is an encoding convention: dividing by 255 rescales the same levels to [0, 1] and adds no information - Networks compute in floating point, because learning needs a large range and small steps; a trained network can later be rounded back to integers (quantized) to run smaller and faster

- A grayscale image keeps one value per pixel and discards color; the 224 × 224 source is resampled to 28 × 28, so each 8 × 8 source block becomes one pixel. Both steps lose information - Why 28 × 28: 784 values can be written out and plotted in full on one slide, which 150,528 cannot, and the resampling shows that resolution is a choice made by the measurement - Vectorization is the name of the operation: the 28 × 28 array is read left to right along each row and the rows are concatenated into one list of 784 numbers. Reversing the order restores the image, so this step loses nothing - Why do it: distances, averages and projections are defined on vectors, and the same formulas then apply to a list of pixels, of voxels, of neurons or of network units. Software still stores and processes the image as an array; the vector is the mathematical view

- The 784 grayscale values of the 28 × 28 image, in row order, plotted as a profile: the horizontal axis is a position in the vector, not time - The profile is the same image in a different format; every measurement in this lecture will be plotted in this same form, one bar per number

- A photograph carries its representation with it; a brain does not. Its response to the same image must be recorded, and each recording instrument observes the brain at a different scale - The next slides go from the human whole-brain scan to single neurons in the monkey, and end with what one image looks like across a population of neurons

- Functional MRI measures the blood-oxygenation-level-dependent (BOLD) signal, an indirect vascular measure that unfolds over seconds. It is non-invasive, which is why it is the main method in humans - Each voxel is a 2 mm cube of tissue holding many thousands of neurons, so one entry of the vector summarizes a neighbourhood - A localizer scan defines the regions before the experiment: early visual cortex (red) responds to any image, scrambled or not; the lateral occipital complex, LOC (green), responds more to objects than to scrambled images. Surface colours identify regions, not response strength; the surface illustration is from Gao et al. (Communications Biology 2025) - The two vectors are participant CSI1's measured responses to this exact ImageNet photograph, a Newfoundland dog, in the BOLD5000 dataset (Chang et al., Scientific Data 2019): all 210 early-visual and all 152 LOC voxels, in their original order - How the numbers were obtained: the participant viewed the photograph for 1 s while whole-brain volumes were acquired every 2 s; each voxel's value is the mean of the third and fourth volumes after onset (4–8 s, when the BOLD response peaks), expressed as a deviation from that voxel's average, so negative entries are below-average signal, not negative activity

- Historical film (about three minutes) of Hubel and Wiesel recording a single cell in cat primary visual cortex, with spikes made audible as clicks - The neuron fires only when the bar of light is at a particular position and orientation: stimulus location, stimulus orientation and the neuron's response are separated in the recording - One cat V1 neuron recorded with one electrode, the classical method; the monkey recordings later in the lecture come from inferotemporal cortex, several stages further along. The original footage is low resolution, and the reading version provides a text description of the audio

- Left: a single metal microelectrode in its holder, the classical instrument. It is advanced through the tissue until it isolates one neuron; the monkey recordings of this course (Bao et al., Nature 2020) were made this way with tungsten microelectrodes, 482 neurons collected over many sessions - Right: much of modern animal electrophysiology runs on Neuropixels probes (Jun et al., Nature 2017): one shank with 960 sites records hundreds of neurons across several brain areas in a single session, with amplification and digitisation on the probe itself - Sites are 20 µm apart, so one neuron's spike appears on several sites; that spatial pattern, with the waveform shape, is what spike sorting uses to separate neurons - The two instruments differ by three orders of magnitude in how many neurons they reach per session, but they produce the same kind of measurement: a firing rate per neuron, one entry of a response vector

- fMRI is non-invasive. This figure surveys the invasive recording methods, ordered by how close the sensor gets to the neurons - Three scales of access: scalp EEG averages millions of neurons; ECoG grids on the cortical surface see a smaller population; only microelectrodes in the tissue resolve the electrical pulses (spikes) of individual neurons - A Utah array (d) or a silicon probe (e) records many sites at once, which is why modern datasets reach thousands of neurons; the monkey recordings used in this lecture and in Assignment 1 (Bao et al., Nature 2020) were instead made with single tungsten microelectrodes, one site at a time, and their 482-neuron vector was assembled across many sessions - An electrode records extracellular voltage; attributing each pulse to a neuron (spike sorting) uses waveform shape and, on multi-site probes, the spatial pattern across sites, so contacts, channels and neurons are different counts - The head in panel a is a rodent; the same three scales exist in humans. Scalp EEG is non-invasive; ECoG and intracortical recordings are invasive, so in humans they are made only when there is a clinical need, for example in epilepsy patients monitored before surgery - No single instrument is best; probe geometry, species, brain region and recording stability decide the choice

- An image is shown for 250 ms; the signal enters at the retina and passes through the lateral geniculate nucleus (LGN), V1, V2, V4 and inferotemporal cortex (IT), whose posterior, central and anterior parts are PIT, CIT and AIT - The pathway is simplified: parallel routes and feedback connections between areas are omitted, and no timing is implied - Anatomy from DiCarlo, Zoccolan and Rust (Neuron 2012), Fig. 3A; the tiger is an Assignment 1 image, and the recording on the next two steps illustrates the method, not a measured response to this photograph

- An electrode advanced into IT sits outside the cells and records small voltage changes in the surrounding fluid; a spike (action potential) appears as a brief pulse - Spike sorting groups pulses by waveform shape to separate one neuron from its neighbours; each tick marks one spike of the sorted unit and its position marks when it occurred - Spike amplitude identifies which neuron fired, not how strongly it is responding - Bao et al. used tungsten microelectrodes advanced one at a time and kept only well-isolated units; each image was shown several times for 250 ms with a grey interval between presentations

- A firing rate is the spike count in a window after image onset divided by the window length; here the window is 60 to 220 ms, allowing for the delay before visual cortex responds - 7 spikes in 0.160 s gives 43.75 spikes per second: one number for this neuron, this image, this trial - Repeated presentations give different counts, so the response used for an image is the average over trials - A rate is one entry of a response vector, as one pixel intensity is one entry of the image vector

- Neurons do not give the same count twice: the same image, the same neuron and the same window yield different numbers of spikes on different trials - The rate recorded for an image is therefore the average across its presentations; in Bao et al. (Nature 2020) each image was shown 4–8 times and the vector entries are those means - The counts are illustrative. One mean per neuron per image is computed before any of the comparisons below

- Recording sites were placed in inferotemporal (IT) patches located by fMRI (Bao et al., Nature 2020, Fig. 4c; three example electrode sites in Extended Data Fig. 8a with a 5 mm scale bar); the two views are separate and do not show all 482 cell locations - Every well-isolated neuron encountered was recorded, with no selectivity cutoff; stimuli were 51 objects at 24 views (1,224 images), each shown for 250 ms with 150 ms grey intervals, in random order, 4–8 repetitions - Each bar is one neuron's firing rate for the cat, in spikes/s: spikes counted 60–220 ms after image onset, divided by 0.160 s, averaged over repetitions; the largest rate for the cat is 80.75 spikes/s (axis to 100), and the largest over the whole image set is 177 spikes/s - 479 of the 482 entries are available; positions 429, 430 and 457 are missing measurements, not zero activity

- The third place to examine a representation is an artificial neural network. Unlike the brain's, its response to an image is computed, and every unit can be read out exactly - The next slides introduce the network used in Assignment 1 and its history, then look at its representation of the same tiger at an early, a middle and a late stage

- We have measured the brain's response to an image; the second system in this course is an artificial network, and this figure places the two side by side - The upper pathway is the biological one whose responses we just measured; the lower pathway is a candidate model, a convolutional neural network, whose responses are computed rather than recorded - In the model, a layer is a processing stage, an activation is a unit's numerical output, each sheet is a spatial feature map, a stack holds several channels, and a convolution applies the same learned filter at every position; LN denotes linear–nonlinear operations (filtering, thresholding, pooling, normalization) - The green dashed arrows propose correspondences to be tested with neural data, not anatomical equivalence; the feedback and recurrent arrows mark what the feedforward model omits - RGC, retinal ganglion cells; LGN, lateral geniculate nucleus; PIT/CIT/AIT, posterior/central/anterior IT; the network is schematic; the 100-ms label comes from the source figure (Bao's images were shown for 250 ms)

- ImageNet classification assigns an image to one of 1,000 categories; top-5 accuracy counts a prediction correct if the true category is among the five highest-ranked; the dates are benchmark years - The points are selected test-set results, most of them ensembles, so the curve is not a controlled experiment on depth; the AlexNet point is its five-model result (83.6%), not the seven-model extra-data entry (84.7%) - AlexNet (Krizhevsky et al. 2012): GPU training, rectified linear units, dropout; VGG (Simonyan & Zisserman 2015): stacked 3 × 3 convolutions; GoogLeNet (Szegedy et al. 2015): Inception modules with parallel filter sizes and 1 × 1 projections; ResNet (He et al. 2016): residual connections that add each block's input to its learned transformation, which made much deeper networks trainable - Better optimization, learning-rate schedules, initialization, normalization and regularization (penalties that stop a network from fitting noise) arrived alongside; no single change explains the curve - ResNet-18, the network used in this lecture and in Assignment 1, is a small ResNet, not the ensemble plotted; higher accuracy is not by itself evidence of biological similarity

- ResNet-18 responses to the tiger: the six most active channels per stage, not entire layers - "Edges and stripes", "parts and textures", "where the tiger is" are rough readings of what the maps look like, not what the channels compute; a channel has no label, and what it responds to is established with the explanation tools introduced later in the course - Early, middle and late are layer1, layer3 and layer4, whose full outputs have 64 × 56 × 56, 256 × 14 × 14 and 512 × 7 × 7 values; one unit is one channel at one spatial position - "Feature" has two uses: in machine learning a feature is one input dimension, so every activation is a feature and a channel is a map of them; in everyday use a feature is what a channel detects, such as stripes - Display scales are shared within a stage but may differ across stages - Assignment 1 compares representations at these three stages; selecting and flattening activations records the layer without altering the network, and the internal computations (convolutions, residual blocks, training) come in later lectures - LENS (Serre lab) shows, for a ResNet-50, which image features each unit and each class depends on; the tiger page previews the explanation tools we return to later in the course

- An ImageNet-trained ResNet-18 receives the same cat image used in the IT example; a layer's output has C × H × W values, and in layer4 each of 512 channels is a 7 × 7 map, so one channel unrolls to 49 values and the whole layer to 512 × 49 = 25,088 - Flattening only reorders values; averaging each channel over space would instead give 512 values - Assignment 1 supplies a fixed subsample of 4,096 of the flattened activations (not 4,096 channels, not the whole layer); each unit is one channel at one location - The print version shows late-layer channel 376, the channel with the largest spatial mean for this input - The explorer does not show the network's class output. These grayscale cut-outs on a grey field are far from the photographs the network was trained on, and its labels for them are unreliable: a first example of an input outside the training distribution, a theme we return to - ResNet: He et al. (CVPR 2016), https://doi.org/10.1109/CVPR.2016.90; stimulus from Bao et al. (Nature 2020), https://doi.org/10.1038/s41586-020-2350-5

- Three sources and four measurements turn an object into a list of numbers of fixed length: 784 pixel intensities, 210 voxel responses, 482 firing rates, 4,096 activations - A vector describes a representation for analysis. The same operations apply to every such vector; what the result means depends on what was measured - The mathematics of such lists is the same whatever they measure. Notation comes first, so that a single formula serves pixels, voxels, neurons and units

- Bold lowercase x is a vector, the whole list of numbers for one object; italic x with subscript j is its jth component, one number; the bold face is typographic and changes nothing about the numbers - Superscript (i) names the image, and (i, r) names one repetition of that image; subscript j always selects one pixel, voxel, neuron or unit, in the same order every time - Bold uppercase X is a matrix, many vectors stacked one image per row; the superscript T is the transpose, which lays a column out as a row; double bars denote the length of a vector - D is the number of measurements per object (784 pixels, 210 voxels, 482 neurons, 4,096 units) and N is the number of objects; these symbols are used in the same way in every lecture and in the assignments

- Two features computed from the 784 grayscale pixel values of the earlier slides: mean intensity, the overall brightness, and RMS contrast, the standard deviation of the pixel intensities around that mean (roughly, how far a typical pixel sits from the mean intensity), in the same units - Values are from the 28 × 28 grayscale tiger on the 0–255 scale, rounded to one decimal; the axes are named x1 and x2 rather than x and y so the notation extends to more than two measurements - Two measurements define a space: one axis per measurement, and every image measured this way is a point in it. This is the sense of "space" in the lecture title, a representational space, and the plot is its simplest instance - The two summaries discard most of the information in the 784 pixels; their virtue is that the space can be drawn

- The same two features computed for the gorilla give (88.4, 44.8): slightly brighter than the tiger and with slightly more contrast, so its point lies up and to the right - Every image measured the same way lands in the same plane; the axes do not change from image to image, only the point does

- All six images of the pixel explorer placed by the same two measurements: the elephant is the brightest and least contrasted, the penguin the brightest with high contrast, the tiger the darkest - Once every image is a point in one space, a comparison between two representations is a comparison between two points, which is a geometric question: how far apart they are, and in which direction. The dissimilarity measures of the next lecture are answers to that question - Two features were chosen so that the space can be drawn; with 784 pixels, 482 neurons or 4,096 units the space cannot be drawn, but the same operations apply

- A vector records the two coordinates; drawn as an arrow from the origin, its tip is the point the coordinates locate, so the same pair of numbers describes both point and vector - Column vectors are the course convention; an inline list of coordinates is shorthand for the same column. The superscript in parentheses numbers the image, here the tiger as image 1 - The arrow picture is what makes the operations that follow geometric: subtracting two vectors gives an arrow between two points, the length of an arrow is a distance, and the angle between two arrows is a similarity. Comparing representations, the theme of the lecture, is done with these operations

- Adding a measurement adds a coordinate, and D coordinates give a vector in R^D whether or not every axis can be drawn - Replacing 784 pixels with intensity and contrast is a reduction, not a reshape; all pixels (784), 210 human voxels or 482 IT neurons each define a different space - Dimensionality is the number of measurements, not the number of objects: ten objects with 482 responses each are ten points in a 482-dimensional space - When each measurement is one neuron, this is the neural state space (Jazayeri & Ostojic 2021): one axis per neuron, and each image's response is one point in it. The same holds for network units or pixels - Negative processed BOLD values are deviations of a processed signal, not negative firing rates; voxel data from Chang et al. (Scientific Data 2019), https://doi.org/10.1038/s41597-019-0052-3 - Spaces with hundreds of axes cannot be plotted directly; the next lecture introduces a method that recovers a low-dimensional space from the distances among points instead

- We have a vector per image. Before comparing vectors across images, ask how stable one vector is: show the same image again and see whether the same numbers come back - Two datasets, human fMRI and monkey electrophysiology, give the same answer, and the remedy is the same: average over repetitions

- BOLD5000 participant CSI1 saw this photograph, an ImageNet bullfrog, three times; each row is the processed response of the same 210 early-visual voxels on one presentation. BOLD5000 repeats only 112 of its 4,916 images, three times each, so a split-half estimate is not possible here - Correlation is used here as a first measure of agreement between two response patterns: 1 when they are identical up to scale and offset, 0 when unrelated; its formula comes next lecture with the other similarity and dissimilarity measures. Correlating the three presentations pairwise gives r_12 = 0.17, r_13 = −0.01 and r_23 = 0.03 (subscripts name the two trials), and each trial correlates with the mean of the other two at 0.10, 0.14 and 0.01. This image is typical: across all 112 repeated images, the median correlation between two single trials is 0.14 and between a trial and the mean of the other two 0.20. The part of a single trial that repeats across presentations is small - The trial-specific variation comes from physiological and scanner sources; single-neuron recordings show the same trial-to-trial variability, as the next slide quantifies with many more repetitions - Averaging the presentations keeps what they share, the image-driven part, and shrinks the rest by roughly the square root of the number of repetitions - This is why datasets repeat stimuli: Bao et al. showed each image four to eight times, and BOLD5000 repeated 112 of its images; the response vector attached to an image is this average

- Majaj et al. (J. Neurosci. 2015) recorded 168 sites in monkey inferotemporal cortex with chronic arrays while showing 3,200 images about fifty times each; values are spike counts in a 70 to 170 ms window after onset, normalised per site - The stimulus is a rendered lioness on a natural background, one of the 3,200 images of the study. Two single presentations of it correlate at only 0.17, no better than the fMRI trials. Averaging the first 25 presentations and the last 25 separately, the two means correlate at 0.69. This is a split-half correlation, the standard way to measure how reliable an averaged response is; it will return when models are compared with neural data, because no model can be expected to predict a response better than the response predicts itself - Every response vector used in this course is a trial mean. This dataset allows 51 presentations per image; the Bao et al. (2020) recordings in Assignment 1 averaged 4 to 8, so their vectors are noisier - Electrophysiology yields far more data than fMRI, with many trials per image and a cleaner signal after averaging, but trial by trial a single neuron is as variable as a single voxel

- Each presentation of the same image yields a different response vector because neurons are noisy; averaging over repetitions estimates the typical response - The average is taken separately for each neuron, so the mean vector has the same length as each single-trial vector; the population of neurons is never collapsed to one number - x^(i,r) is one trial of image i and x^(i) the trial average; subscript j selects one neuron, in the same order every time - The numbers are a teaching example. The Assignment 1 monkey data are already averaged: the matrix you receive is 1,224 images by 482 units, each entry the mean rate over that image's 4–8 presentations; the single trials are not provided. A network in evaluation mode is deterministic, so one forward pass is its "average"

- A response vector is written as a column of shape D × 1; its transpose, marked by an upright superscript T, lays the same numbers out as a 1 × D row, and transposing twice returns the column - Superscript (i) names the image; the T marks an operation, not an image index or an exponent - Rows are the form in which many images are stacked into one matrix, so the column convention for a single vector x and the one-image-per-row convention for the matrix X coexist

- Transposing each image's column vector to a row and stacking the rows gives a matrix X of N images by D features. The same matrix written entry by entry has x_ij in row i, column j: the j-th feature of image i, which is the same number as x_j^(i) - Each cell of the heatmap is one such entry; moving along a row changes the feature, moving down a column changes the image - Some rows already look alike and some do not. Penguin and gorilla share the strong response in column 2 and agree elsewhere (correlation 0.55 on these 12 units); frog and elephant have no strong unit in common (correlation −0.54). Reading similarity off rows by eye is what the dissimilarity measures below make precise - For repeated neural measurements each row holds the trial-averaged response vector. The heatmap shows late-layer ResNet-18 activations for six animal images, 12 of the 4,096 stored features; colour encodes activation magnitude, not category - The same format holds voxel responses or EEG features as columns once the measurement and its time window are fixed

- Response vectors from different systems cannot be compared entry by entry: a pixel is not a voxel and a neuron is not a network unit, and the vectors have different lengths - What can be compared is how each system arranges the same set of images: which pairs it treats as alike and which as different. A matrix of pairwise dissimilarities records that arrangement for any system, in the same format - The next lecture recovers a space from this matrix; later lectures compare the matrices of two systems, which is how models are tested against brains and behaviour

- X is images by features; the representational dissimilarity matrix D is images by images, with entry (i, k) comparing the response vector of image i with that of image k. D is not another activation matrix - Lower-case d is the dissimilarity measure, the function that compares two vectors; different measures keep different properties of the responses, and choosing d is a scientific decision - The entries obey constraints: the diagonal is zero, the matrix is symmetric, and when d is a distance the direct comparison of two images can never exceed the sum of two comparisons through a third image (the triangle inequality, examined next lecture) - The matrix shown is the one you build in Assignment 1: all 120 images, ordered by their eight animal categories, correlation distance on the trained network's 4,096 late-layer activations. Light blocks along the diagonal mean that images of one category are more alike than images of different categories (within-category mean 0.82, between 0.91); primates and carnivores form the clearest blocks. The colour scale is clipped to the 2nd–98th percentile so the structure is visible - Pairwise human ratings fit the same format without any response vector being measured, which is what allows relationships to be compared across systems without equating their units or components

- Optional resources for seeing a network operate, none required - 3Blue1Brown, But what is a neural network? (19 min): a fully connected network on handwritten digits with weights and activations drawn out; Harley's visualizer: the same live in the browser, for a multilayer perceptron and a convolutional network - CNN Explainer (Georgia Tech): the arithmetic inside each convolution, ReLU (a unit that passes positive inputs and outputs zero otherwise) and pooling on the actual numbers, with your own uploaded image - Yosinski, Deep Visualization Toolbox (4 min): a webcam feed through AlexNet with every layer on screen - TensorFlow Playground: tiny networks on 2-D toy data, the best place to watch training reshape a representation

- Six papers connected to this lecture, none required. The research trail runs across several lectures: your paper can come from any lecture so far, and it need not be one of these six. Details and dates on Canvas - DiCarlo, Zoccolan and Rust (2012), How does the brain solve visual object recognition?, Neuron 73:415–434: the ventral pathway and the argument that object identity is read from a population of IT neurons, the setting of our recording slides - Kriegeskorte and Kievit (2013), Representational geometry, Trends in Cognitive Sciences 17:401–412: a broad survey; no derivations needed - Yamins and DiCarlo (2016), Using goal-driven deep learning models to understand sensory cortex, Nature Neuroscience 19:356–365: the source of the visual-cortex-and-CNN figure, and the argument that task training is what makes a network brain-like - Geirhos et al. (2019), ImageNet-trained CNNs are biased towards texture, ICLR: cue-conflict images show what a network's similarity structure keys on - Bowers et al. (2023), Deep problems with neural network models of human vision, Behavioral and Brain Sciences 46:e385: a critical target article, published with commentaries - Peterson, Abbott and Griffiths (2018, Cognitive Science) and Bao, She, McGill and Tsao (2020, Nature) are the sources of the Assignment 1 ratings and recordings; reading them helps the assignment, but the trail is for going beyond it

- Reading a paper for the trail is not reading a textbook: the aim is to find the one question the paper does not settle, not to absorb everything. The trail spans several lectures, so take your time choosing - Skim before reading: the abstract, the figures and their captions, and the last paragraph of the discussion tell you what was claimed and whether you care. Only then read the methods - Three questions that almost always produce something worth posting: is there a control the authors should have run; would the result survive a different network, animal, dataset or dissimilarity measure; and does the result follow trivially from the method, so that it would have appeared with random data - The first paper takes long and feels hard; that is normal. By the third or fourth the questions come on their own - The post itself is short: four to six sentences plus one reply to a classmate, on Ed; the dates are on Canvas

- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 2) - Three brief points in your own words: one short answer to the question of the day, one thing you did not understand, one idea you found interesting; a few sentences each is enough - Credit is for a thoughtful attempt, not for being correct; recurring questions are summarized anonymously and addressed on Ed or at the start of the next lecture

- X is objects by components and D is objects by objects; pdist computes the unique pairs and squareform arranges them into a symmetric matrix - Expected output for A = (1, 1), B = (4, 4), C = (1, −1): Euclidean A–B = 4.24, A–C = 2.00; cosine A–B = 0, A–C = 1 - Correlation is omitted because A and B are constant across their components, so it is undefined; SciPy's parameter is named metric even for functions that are not mathematical metrics (https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.distance.pdist.html)

- Multiplying the 28 × 28 tiger by a scalar c acts on every one of its 784 entries (an intensity of 27 becomes 13.5 at c = 0.5); the vector keeps fractional values, neither rounded nor clipped, even though a display quantizes them - For c > 0 the vector keeps its direction and its length is multiplied by c; at c = 0 it is the zero vector, whose direction is undefined - The control runs from 0 to 1 because a bounded display clips values above its maximum, so for c > 1 the displayed array would no longer be exactly c times the original - Intensity scaling is one form of data augmentation: transforming training examples while keeping the label, provided the relevant information stays visible - A network's internal responses do not in general scale by the same factor as its input

- Two 28 × 28 images are vectors with 784 entries in the same order, so a weighted average z = (1 − λ)x + λy combines corresponding entries: λ = 0 gives x, λ = 1 gives y, λ = 0.5 gives the arithmetic mean (x + y)/2, e.g. (18 + 59)/2 = 38.5 for the first tiger and gorilla pixel - Nonnegative weights summing to one keep intensities within the original range; fractional values stay in the vector and are rounded only for display - The result is a cross-fade between pixel arrays, not a morph between animal shapes; mixup (Zhang et al., ICLR 2018) trains networks on such blends of images and their labels - Averaging repeated trials of one stimulus estimates a mean response, whereas averaging two different images constructs a new input. A nonlinear network does not in general blend its internal responses the same way