A representation: a pattern that carries information about a stimulus, a property or a state. Each measurement so far gives one such pattern as a list of numbers.
| Symbol | Meaning |
|---|---|
| (bold lowercase) | a vector: the whole list of numbers for one object |
| (italic, subscript) | its th component, a single number |
| , | the vector for image ; the same image on repetition |
| (bold uppercase) | a matrix: many vectors stacked, one image per row |
| its entry in row , column : feature of image | |
| the transpose: the same numbers laid out as a row instead of a column | |
| the length of the vector | |
| , | the number of measurements per object; the number of objects |
: mean intensity · : RMS contrast (standard deviation of the intensities)
One vector, two components:
The two axes define a space; the tiger is one point in it.
: mean intensity · : RMS contrast (standard deviation of the intensities)
One vector, two components:
A different image, a different point.
: mean intensity · : RMS contrast (standard deviation of the intensities)
One vector, two components:
Every image is a point in the same space. Comparing representations therefore reduces to geometry: distances and angles between points.
Same two numbers: a point, or an arrow from the origin to it.
| Measurement | |
|---|---|
| Two image features | 2 |
| Resampled grayscale image | 784 |
| Recorded IT population | 482 |
| Human early visual cortex (voxels) | 210 |
| Network layer (sampled units) | 4,096 |
Repetition of gives . Three neurons, three repetitions (illustrative spikes/s):
Column: rows, one column. Transpose : the same numbers as a row.
Stack row vectors: one matrix for the whole image set.
Pixels, voxels, neurons and network units cannot be compared entry by entry: the vectors differ in length and in units. Their pairwise dissimilarities can. The RDM records them.
Representational dissimilarity matrix (RDM): every pairwise comparison.
Two lectures on convolutional networks and one on training are coming. Previews:
Not required. Ten minutes well spent: the explainer and the playground.
Six starting points, none required. Your trail paper can come from any lecture; these are seeds, not a list to choose from. Peterson et al. 2018 and Bao et al. 2020 are the assignment papers: fine to read, but branch out.
Canvas submission: Canvas → Minute papers → Minute Paper 2 (access code read out in class)
Write three brief points in your own words:
Credit for a thoughtful attempt, not for being correct.
import numpy as np
from scipy.spatial.distance import pdist, squareform
X = np.array([[1., 1.], # pattern A
[4., 4.], # pattern B
[1., -1.]]) # pattern C
for metric in ("euclidean", "cosine"):
D = squareform(pdist(X, metric=metric))
print(metric, np.round(D, 2))
pdist: the unique pairs. squareform: the symmetric matrix.
Scaling: . Same 784 entries, only their size changes.
Blending: , entry by entry. Done to labels too, this is mixup.

- From minute paper 1: many of you want to build and train models yourselves, and many want to understand how brains and AI relate; the two most common worries were time and programming. Lectures own the concepts; recitations own the code - Recitations meet on Tuesdays, 6:30–7:50 PM; the room is posted on Canvas. The first sessions are a Python bootcamp, optional but strongly encouraged if Python is new to you - Slides and presenter notes are posted on Canvas after each lecture; a reader-friendly version with figure descriptions is linked from the title slide
- These photographs are examples of the natural images you will use in Assignment 1: the 120 animal photographs, drawn from ImageNet, of Peterson, Abbott and Griffiths (Cognitive Science 2018), for which human similarity ratings exist - One system pairs A with B; another pairs B with C. A and B are near mirror images of one another (the outline); B and C share the same body orientation. The two are hypothetical classifiers that could reach the same overall accuracy while attending to very different features of the images - Similarity is therefore not a property of the pixels alone; it depends on the system doing the comparing, whether a human observer, a neural population or a network, and on what that system attends to - Similarity judgments and measurements of this kind are a standard tool in cognitive science and in AI for characterizing the representations of different systems and comparing them; this week introduces the tool
- What is a representation? Without going far into philosophy, Magritte's painting is the classical starting point - Magritte's painted pipe cannot be smoked: an object and its depiction are different things - The tiger exists as an animal, as a photograph of the animal, and as the array of numbers that encodes the photograph; the three are related but not interchangeable - A representation, in this course, is something that stands for an object without being it. The painting is the simplest case; the numbers an image, a brain or a network hold about the tiger are the cases we study - René Magritte, The Treachery of Images, 1929, LACMA (© C. Herscovici / Artists Rights Society, New York; photo © Museum Associates/LACMA)
- Three representations of the same image will be examined in turn: in the image itself, in the activity of a brain viewing it, and in the activations of an artificial neural network processing it - The image is the simplest case. A digital image is an array of numbers and nothing more. It is legible to us only because a display converts those numbers into light: at every location, three numbers set the amounts of red, green and blue emitted by the projector. The next three slides examine that array
- A photograph is already a set of measurements: a color image stores three values per pixel, one per channel (red, green, blue); at 224 × 224 pixels that is 224 × 224 × 3 = 150,528 values - Eight bits give 2^8 = 256 levels, so unsigned 8-bit values run from 0 to 255; RGB commonly uses 8 bits per channel, 24 bits per pixel - The 0–255 scale is an encoding convention: dividing by 255 rescales the same levels to [0, 1] and adds no information - Networks compute in floating point, because learning needs a large range and small steps; a trained network can later be rounded back to integers (quantized) to run smaller and faster
- A grayscale image keeps one value per pixel and discards color; the 224 × 224 source is resampled to 28 × 28, so each 8 × 8 source block becomes one pixel. Both steps lose information - Why 28 × 28: 784 values can be written out and plotted in full on one slide, which 150,528 cannot, and the resampling shows that resolution is a choice made by the measurement - Vectorization is the name of the operation: the 28 × 28 array is read left to right along each row and the rows are concatenated into one list of 784 numbers. Reversing the order restores the image, so this step loses nothing - Why do it: distances, averages and projections are defined on vectors, and the same formulas then apply to a list of pixels, of voxels, of neurons or of network units. Software still stores and processes the image as an array; the vector is the mathematical view
- The 784 grayscale values of the 28 × 28 image, in row order, plotted as a profile: the horizontal axis is a position in the vector, not time - The profile is the same image in a different format; every measurement in this lecture will be plotted in this same form, one bar per number
- A photograph carries its representation with it; a brain does not. Its response to the same image must be recorded, and each recording instrument observes the brain at a different scale - The next slides go from the human whole-brain scan to single neurons in the monkey, and end with what one image looks like across a population of neurons
- Functional MRI measures the blood-oxygenation-level-dependent (BOLD) signal, an indirect vascular measure that unfolds over seconds. It is non-invasive, which is why it is the main method in humans - Each voxel is a 2 mm cube of tissue holding many thousands of neurons, so one entry of the vector summarizes a neighbourhood - A localizer scan defines the regions before the experiment: early visual cortex (red) responds to any image, scrambled or not; the lateral occipital complex, LOC (green), responds more to objects than to scrambled images. Surface colours identify regions, not response strength; the surface illustration is from Gao et al. (Communications Biology 2025) - The two vectors are participant CSI1's measured responses to this exact ImageNet photograph, a Newfoundland dog, in the BOLD5000 dataset (Chang et al., Scientific Data 2019): all 210 early-visual and all 152 LOC voxels, in their original order - How the numbers were obtained: the participant viewed the photograph for 1 s while whole-brain volumes were acquired every 2 s; each voxel's value is the mean of the third and fourth volumes after onset (4–8 s, when the BOLD response peaks), expressed as a deviation from that voxel's average, so negative entries are below-average signal, not negative activity
- Historical film (about three minutes) of Hubel and Wiesel recording a single cell in cat primary visual cortex, with spikes made audible as clicks - The neuron fires only when the bar of light is at a particular position and orientation: stimulus location, stimulus orientation and the neuron's response are separated in the recording - One cat V1 neuron recorded with one electrode, the classical method; the monkey recordings later in the lecture come from inferotemporal cortex, several stages further along. The original footage is low resolution, and the reading version provides a text description of the audio
- Left: a single metal microelectrode in its holder, the classical instrument. It is advanced through the tissue until it isolates one neuron; the monkey recordings of this course (Bao et al., Nature 2020) were made this way with tungsten microelectrodes, 482 neurons collected over many sessions - Right: much of modern animal electrophysiology runs on Neuropixels probes (Jun et al., Nature 2017): one shank with 960 sites records hundreds of neurons across several brain areas in a single session, with amplification and digitisation on the probe itself - Sites are 20 µm apart, so one neuron's spike appears on several sites; that spatial pattern, with the waveform shape, is what spike sorting uses to separate neurons - The two instruments differ by three orders of magnitude in how many neurons they reach per session, but they produce the same kind of measurement: a firing rate per neuron, one entry of a response vector
- fMRI is non-invasive. This figure surveys the invasive recording methods, ordered by how close the sensor gets to the neurons - Three scales of access: scalp EEG averages millions of neurons; ECoG grids on the cortical surface see a smaller population; only microelectrodes in the tissue resolve the electrical pulses (spikes) of individual neurons - A Utah array (d) or a silicon probe (e) records many sites at once, which is why modern datasets reach thousands of neurons; the monkey recordings used in this lecture and in Assignment 1 (Bao et al., Nature 2020) were instead made with single tungsten microelectrodes, one site at a time, and their 482-neuron vector was assembled across many sessions - An electrode records extracellular voltage; attributing each pulse to a neuron (spike sorting) uses waveform shape and, on multi-site probes, the spatial pattern across sites, so contacts, channels and neurons are different counts - The head in panel a is a rodent; the same three scales exist in humans. Scalp EEG is non-invasive; ECoG and intracortical recordings are invasive, so in humans they are made only when there is a clinical need, for example in epilepsy patients monitored before surgery - No single instrument is best; probe geometry, species, brain region and recording stability decide the choice
- An image is shown for 250 ms; the signal enters at the retina and passes through the lateral geniculate nucleus (LGN), V1, V2, V4 and inferotemporal cortex (IT), whose posterior, central and anterior parts are PIT, CIT and AIT - The pathway is simplified: parallel routes and feedback connections between areas are omitted, and no timing is implied - Anatomy from DiCarlo, Zoccolan and Rust (Neuron 2012), Fig. 3A; the tiger is an Assignment 1 image, and the recording on the next two steps illustrates the method, not a measured response to this photograph
- An electrode advanced into IT sits outside the cells and records small voltage changes in the surrounding fluid; a spike (action potential) appears as a brief pulse - Spike sorting groups pulses by waveform shape to separate one neuron from its neighbours; each tick marks one spike of the sorted unit and its position marks when it occurred - Spike amplitude identifies which neuron fired, not how strongly it is responding - Bao et al. used tungsten microelectrodes advanced one at a time and kept only well-isolated units; each image was shown several times for 250 ms with a grey interval between presentations
- A firing rate is the spike count in a window after image onset divided by the window length; here the window is 60 to 220 ms, allowing for the delay before visual cortex responds - 7 spikes in 0.160 s gives 43.75 spikes per second: one number for this neuron, this image, this trial - Repeated presentations give different counts, so the response used for an image is the average over trials - A rate is one entry of a response vector, as one pixel intensity is one entry of the image vector
- Neurons do not give the same count twice: the same image, the same neuron and the same window yield different numbers of spikes on different trials - The rate recorded for an image is therefore the average across its presentations; in Bao et al. (Nature 2020) each image was shown 4–8 times and the vector entries are those means - The counts are illustrative. One mean per neuron per image is computed before any of the comparisons below
- Recording sites were placed in inferotemporal (IT) patches located by fMRI (Bao et al., Nature 2020, Fig. 4c; three example electrode sites in Extended Data Fig. 8a with a 5 mm scale bar); the two views are separate and do not show all 482 cell locations - Every well-isolated neuron encountered was recorded, with no selectivity cutoff; stimuli were 51 objects at 24 views (1,224 images), each shown for 250 ms with 150 ms grey intervals, in random order, 4–8 repetitions - Each bar is one neuron's firing rate for the cat, in spikes/s: spikes counted 60–220 ms after image onset, divided by 0.160 s, averaged over repetitions; the largest rate for the cat is 80.75 spikes/s (axis to 100), and the largest over the whole image set is 177 spikes/s - 479 of the 482 entries are available; positions 429, 430 and 457 are missing measurements, not zero activity
- The third place to examine a representation is an artificial neural network. Unlike the brain's, its response to an image is computed, and every unit can be read out exactly - The next slides introduce the network used in Assignment 1 and its history, then look at its representation of the same tiger at an early, a middle and a late stage
- We have measured the brain's response to an image; the second system in this course is an artificial network, and this figure places the two side by side - The upper pathway is the biological one whose responses we just measured; the lower pathway is a candidate model, a convolutional neural network, whose responses are computed rather than recorded - In the model, a layer is a processing stage, an activation is a unit's numerical output, each sheet is a spatial feature map, a stack holds several channels, and a convolution applies the same learned filter at every position; LN denotes linear–nonlinear operations (filtering, thresholding, pooling, normalization) - The green dashed arrows propose correspondences to be tested with neural data, not anatomical equivalence; the feedback and recurrent arrows mark what the feedforward model omits - RGC, retinal ganglion cells; LGN, lateral geniculate nucleus; PIT/CIT/AIT, posterior/central/anterior IT; the network is schematic; the 100-ms label comes from the source figure (Bao's images were shown for 250 ms)
- ImageNet classification assigns an image to one of 1,000 categories; top-5 accuracy counts a prediction correct if the true category is among the five highest-ranked; the dates are benchmark years - The points are selected test-set results, most of them ensembles, so the curve is not a controlled experiment on depth; the AlexNet point is its five-model result (83.6%), not the seven-model extra-data entry (84.7%) - AlexNet (Krizhevsky et al. 2012): GPU training, rectified linear units, dropout; VGG (Simonyan & Zisserman 2015): stacked 3 × 3 convolutions; GoogLeNet (Szegedy et al. 2015): Inception modules with parallel filter sizes and 1 × 1 projections; ResNet (He et al. 2016): residual connections that add each block's input to its learned transformation, which made much deeper networks trainable - Better optimization, learning-rate schedules, initialization, normalization and regularization (penalties that stop a network from fitting noise) arrived alongside; no single change explains the curve - ResNet-18, the network used in this lecture and in Assignment 1, is a small ResNet, not the ensemble plotted; higher accuracy is not by itself evidence of biological similarity
- ResNet-18 responses to the tiger: the six most active channels per stage, not entire layers - "Edges and stripes", "parts and textures", "where the tiger is" are rough readings of what the maps look like, not what the channels compute; a channel has no label, and what it responds to is established with the explanation tools introduced later in the course - Early, middle and late are layer1, layer3 and layer4, whose full outputs have 64 × 56 × 56, 256 × 14 × 14 and 512 × 7 × 7 values; one unit is one channel at one spatial position - "Feature" has two uses: in machine learning a feature is one input dimension, so every activation is a feature and a channel is a map of them; in everyday use a feature is what a channel detects, such as stripes - Display scales are shared within a stage but may differ across stages - Assignment 1 compares representations at these three stages; selecting and flattening activations records the layer without altering the network, and the internal computations (convolutions, residual blocks, training) come in later lectures - LENS (Serre lab) shows, for a ResNet-50, which image features each unit and each class depends on; the tiger page previews the explanation tools we return to later in the course
- An ImageNet-trained ResNet-18 receives the same cat image used in the IT example; a layer's output has C × H × W values, and in layer4 each of 512 channels is a 7 × 7 map, so one channel unrolls to 49 values and the whole layer to 512 × 49 = 25,088 - Flattening only reorders values; averaging each channel over space would instead give 512 values - Assignment 1 supplies a fixed subsample of 4,096 of the flattened activations (not 4,096 channels, not the whole layer); each unit is one channel at one location - The print version shows late-layer channel 376, the channel with the largest spatial mean for this input - The explorer does not show the network's class output. These grayscale cut-outs on a grey field are far from the photographs the network was trained on, and its labels for them are unreliable: a first example of an input outside the training distribution, a theme we return to - ResNet: He et al. (CVPR 2016), https://doi.org/10.1109/CVPR.2016.90; stimulus from Bao et al. (Nature 2020), https://doi.org/10.1038/s41586-020-2350-5
- Three sources and four measurements turn an object into a list of numbers of fixed length: 784 pixel intensities, 210 voxel responses, 482 firing rates, 4,096 activations - A vector describes a representation for analysis. The same operations apply to every such vector; what the result means depends on what was measured - The mathematics of such lists is the same whatever they measure. Notation comes first, so that a single formula serves pixels, voxels, neurons and units
- Bold lowercase x is a vector, the whole list of numbers for one object; italic x with subscript j is its jth component, one number; the bold face is typographic and changes nothing about the numbers - Superscript (i) names the image, and (i, r) names one repetition of that image; subscript j always selects one pixel, voxel, neuron or unit, in the same order every time - Bold uppercase X is a matrix, many vectors stacked one image per row; the superscript T is the transpose, which lays a column out as a row; double bars denote the length of a vector - D is the number of measurements per object (784 pixels, 210 voxels, 482 neurons, 4,096 units) and N is the number of objects; these symbols are used in the same way in every lecture and in the assignments
- Two features computed from the 784 grayscale pixel values of the earlier slides: mean intensity, the overall brightness, and RMS contrast, the standard deviation of the pixel intensities around that mean (roughly, how far a typical pixel sits from the mean intensity), in the same units - Values are from the 28 × 28 grayscale tiger on the 0–255 scale, rounded to one decimal; the axes are named x1 and x2 rather than x and y so the notation extends to more than two measurements - Two measurements define a space: one axis per measurement, and every image measured this way is a point in it. This is the sense of "space" in the lecture title, a representational space, and the plot is its simplest instance - The two summaries discard most of the information in the 784 pixels; their virtue is that the space can be drawn
- The same two features computed for the gorilla give (88.4, 44.8): slightly brighter than the tiger and with slightly more contrast, so its point lies up and to the right - Every image measured the same way lands in the same plane; the axes do not change from image to image, only the point does
- All six images of the pixel explorer placed by the same two measurements: the elephant is the brightest and least contrasted, the penguin the brightest with high contrast, the tiger the darkest - Once every image is a point in one space, a comparison between two representations is a comparison between two points, which is a geometric question: how far apart they are, and in which direction. The dissimilarity measures of the next lecture are answers to that question - Two features were chosen so that the space can be drawn; with 784 pixels, 482 neurons or 4,096 units the space cannot be drawn, but the same operations apply
- A vector records the two coordinates; drawn as an arrow from the origin, its tip is the point the coordinates locate, so the same pair of numbers describes both point and vector - Column vectors are the course convention; an inline list of coordinates is shorthand for the same column. The superscript in parentheses numbers the image, here the tiger as image 1 - The arrow picture is what makes the operations that follow geometric: subtracting two vectors gives an arrow between two points, the length of an arrow is a distance, and the angle between two arrows is a similarity. Comparing representations, the theme of the lecture, is done with these operations
- Adding a measurement adds a coordinate, and D coordinates give a vector in R^D whether or not every axis can be drawn - Replacing 784 pixels with intensity and contrast is a reduction, not a reshape; all pixels (784), 210 human voxels or 482 IT neurons each define a different space - Dimensionality is the number of measurements, not the number of objects: ten objects with 482 responses each are ten points in a 482-dimensional space - When each measurement is one neuron, this is the neural state space (Jazayeri & Ostojic 2021): one axis per neuron, and each image's response is one point in it. The same holds for network units or pixels - Negative processed BOLD values are deviations of a processed signal, not negative firing rates; voxel data from Chang et al. (Scientific Data 2019), https://doi.org/10.1038/s41597-019-0052-3 - Spaces with hundreds of axes cannot be plotted directly; the next lecture introduces a method that recovers a low-dimensional space from the distances among points instead
- We have a vector per image. Before comparing vectors across images, ask how stable one vector is: show the same image again and see whether the same numbers come back - Two datasets, human fMRI and monkey electrophysiology, give the same answer, and the remedy is the same: average over repetitions
- BOLD5000 participant CSI1 saw this photograph, an ImageNet bullfrog, three times; each row is the processed response of the same 210 early-visual voxels on one presentation. BOLD5000 repeats only 112 of its 4,916 images, three times each, so a split-half estimate is not possible here - Correlation is used here as a first measure of agreement between two response patterns: 1 when they are identical up to scale and offset, 0 when unrelated; its formula comes next lecture with the other similarity and dissimilarity measures. Correlating the three presentations pairwise gives r_12 = 0.17, r_13 = −0.01 and r_23 = 0.03 (subscripts name the two trials), and each trial correlates with the mean of the other two at 0.10, 0.14 and 0.01. This image is typical: across all 112 repeated images, the median correlation between two single trials is 0.14 and between a trial and the mean of the other two 0.20. The part of a single trial that repeats across presentations is small - The trial-specific variation comes from physiological and scanner sources; single-neuron recordings show the same trial-to-trial variability, as the next slide quantifies with many more repetitions - Averaging the presentations keeps what they share, the image-driven part, and shrinks the rest by roughly the square root of the number of repetitions - This is why datasets repeat stimuli: Bao et al. showed each image four to eight times, and BOLD5000 repeated 112 of its images; the response vector attached to an image is this average
- Majaj et al. (J. Neurosci. 2015) recorded 168 sites in monkey inferotemporal cortex with chronic arrays while showing 3,200 images about fifty times each; values are spike counts in a 70 to 170 ms window after onset, normalised per site - The stimulus is a rendered lioness on a natural background, one of the 3,200 images of the study. Two single presentations of it correlate at only 0.17, no better than the fMRI trials. Averaging the first 25 presentations and the last 25 separately, the two means correlate at 0.69. This is a split-half correlation, the standard way to measure how reliable an averaged response is; it will return when models are compared with neural data, because no model can be expected to predict a response better than the response predicts itself - Every response vector used in this course is a trial mean. This dataset allows 51 presentations per image; the Bao et al. (2020) recordings in Assignment 1 averaged 4 to 8, so their vectors are noisier - Electrophysiology yields far more data than fMRI, with many trials per image and a cleaner signal after averaging, but trial by trial a single neuron is as variable as a single voxel
- Each presentation of the same image yields a different response vector because neurons are noisy; averaging over repetitions estimates the typical response - The average is taken separately for each neuron, so the mean vector has the same length as each single-trial vector; the population of neurons is never collapsed to one number - x^(i,r) is one trial of image i and x^(i) the trial average; subscript j selects one neuron, in the same order every time - The numbers are a teaching example. The Assignment 1 monkey data are already averaged: the matrix you receive is 1,224 images by 482 units, each entry the mean rate over that image's 4–8 presentations; the single trials are not provided. A network in evaluation mode is deterministic, so one forward pass is its "average"
- A response vector is written as a column of shape D × 1; its transpose, marked by an upright superscript T, lays the same numbers out as a 1 × D row, and transposing twice returns the column - Superscript (i) names the image; the T marks an operation, not an image index or an exponent - Rows are the form in which many images are stacked into one matrix, so the column convention for a single vector x and the one-image-per-row convention for the matrix X coexist
- Transposing each image's column vector to a row and stacking the rows gives a matrix X of N images by D features. The same matrix written entry by entry has x_ij in row i, column j: the j-th feature of image i, which is the same number as x_j^(i) - Each cell of the heatmap is one such entry; moving along a row changes the feature, moving down a column changes the image - Some rows already look alike and some do not. Penguin and gorilla share the strong response in column 2 and agree elsewhere (correlation 0.55 on these 12 units); frog and elephant have no strong unit in common (correlation −0.54). Reading similarity off rows by eye is what the dissimilarity measures below make precise - For repeated neural measurements each row holds the trial-averaged response vector. The heatmap shows late-layer ResNet-18 activations for six animal images, 12 of the 4,096 stored features; colour encodes activation magnitude, not category - The same format holds voxel responses or EEG features as columns once the measurement and its time window are fixed
- Response vectors from different systems cannot be compared entry by entry: a pixel is not a voxel and a neuron is not a network unit, and the vectors have different lengths - What can be compared is how each system arranges the same set of images: which pairs it treats as alike and which as different. A matrix of pairwise dissimilarities records that arrangement for any system, in the same format - The next lecture recovers a space from this matrix; later lectures compare the matrices of two systems, which is how models are tested against brains and behaviour
- X is images by features; the representational dissimilarity matrix D is images by images, with entry (i, k) comparing the response vector of image i with that of image k. D is not another activation matrix - Lower-case d is the dissimilarity measure, the function that compares two vectors; different measures keep different properties of the responses, and choosing d is a scientific decision - The entries obey constraints: the diagonal is zero, the matrix is symmetric, and when d is a distance the direct comparison of two images can never exceed the sum of two comparisons through a third image (the triangle inequality, examined next lecture) - The matrix shown is the one you build in Assignment 1: all 120 images, ordered by their eight animal categories, correlation distance on the trained network's 4,096 late-layer activations. Light blocks along the diagonal mean that images of one category are more alike than images of different categories (within-category mean 0.82, between 0.91); primates and carnivores form the clearest blocks. The colour scale is clipped to the 2nd–98th percentile so the structure is visible - Pairwise human ratings fit the same format without any response vector being measured, which is what allows relationships to be compared across systems without equating their units or components
- Optional resources for seeing a network operate, none required - 3Blue1Brown, But what is a neural network? (19 min): a fully connected network on handwritten digits with weights and activations drawn out; Harley's visualizer: the same live in the browser, for a multilayer perceptron and a convolutional network - CNN Explainer (Georgia Tech): the arithmetic inside each convolution, ReLU (a unit that passes positive inputs and outputs zero otherwise) and pooling on the actual numbers, with your own uploaded image - Yosinski, Deep Visualization Toolbox (4 min): a webcam feed through AlexNet with every layer on screen - TensorFlow Playground: tiny networks on 2-D toy data, the best place to watch training reshape a representation
- Six papers connected to this lecture, none required. The research trail runs across several lectures: your paper can come from any lecture so far, and it need not be one of these six. Details and dates on Canvas - DiCarlo, Zoccolan and Rust (2012), How does the brain solve visual object recognition?, Neuron 73:415–434: the ventral pathway and the argument that object identity is read from a population of IT neurons, the setting of our recording slides - Kriegeskorte and Kievit (2013), Representational geometry, Trends in Cognitive Sciences 17:401–412: a broad survey; no derivations needed - Yamins and DiCarlo (2016), Using goal-driven deep learning models to understand sensory cortex, Nature Neuroscience 19:356–365: the source of the visual-cortex-and-CNN figure, and the argument that task training is what makes a network brain-like - Geirhos et al. (2019), ImageNet-trained CNNs are biased towards texture, ICLR: cue-conflict images show what a network's similarity structure keys on - Bowers et al. (2023), Deep problems with neural network models of human vision, Behavioral and Brain Sciences 46:e385: a critical target article, published with commentaries - Peterson, Abbott and Griffiths (2018, Cognitive Science) and Bao, She, McGill and Tsao (2020, Nature) are the sources of the Assignment 1 ratings and recordings; reading them helps the assignment, but the trail is for going beyond it
- Reading a paper for the trail is not reading a textbook: the aim is to find the one question the paper does not settle, not to absorb everything. The trail spans several lectures, so take your time choosing - Skim before reading: the abstract, the figures and their captions, and the last paragraph of the discussion tell you what was claimed and whether you care. Only then read the methods - Three questions that almost always produce something worth posting: is there a control the authors should have run; would the result survive a different network, animal, dataset or dissimilarity measure; and does the result follow trivially from the method, so that it would have appeared with random data - The first paper takes long and feels hard; that is normal. By the third or fourth the questions come on their own - The post itself is short: four to six sentences plus one reply to a classmate, on Ed; the dates are on Canvas
- A short, low-stakes reflection written at the end of the lecture and submitted on Canvas (Minute papers → Minute Paper 2) - Three brief points in your own words: one short answer to the question of the day, one thing you did not understand, one idea you found interesting; a few sentences each is enough - Credit is for a thoughtful attempt, not for being correct; recurring questions are summarized anonymously and addressed on Ed or at the start of the next lecture
- X is objects by components and D is objects by objects; pdist computes the unique pairs and squareform arranges them into a symmetric matrix - Expected output for A = (1, 1), B = (4, 4), C = (1, −1): Euclidean A–B = 4.24, A–C = 2.00; cosine A–B = 0, A–C = 1 - Correlation is omitted because A and B are constant across their components, so it is undefined; SciPy's parameter is named metric even for functions that are not mathematical metrics (https://docs.scipy.org/doc/scipy/reference/generated/scipy.spatial.distance.pdist.html)
- Multiplying the 28 × 28 tiger by a scalar c acts on every one of its 784 entries (an intensity of 27 becomes 13.5 at c = 0.5); the vector keeps fractional values, neither rounded nor clipped, even though a display quantizes them - For c > 0 the vector keeps its direction and its length is multiplied by c; at c = 0 it is the zero vector, whose direction is undefined - The control runs from 0 to 1 because a bounded display clips values above its maximum, so for c > 1 the displayed array would no longer be exactly c times the original - Intensity scaling is one form of data augmentation: transforming training examples while keeping the label, provided the relevant information stays visible - A network's internal responses do not in general scale by the same factor as its input
- Two 28 × 28 images are vectors with 784 entries in the same order, so a weighted average z = (1 − λ)x + λy combines corresponding entries: λ = 0 gives x, λ = 1 gives y, λ = 0.5 gives the arithmetic mean (x + y)/2, e.g. (18 + 59)/2 = 38.5 for the first tiger and gorilla pixel - Nonnegative weights summing to one keep intensities within the original range; fractional values stay in the vector and are rounded only for display - The result is a cross-fade between pixel arrays, not a morph between animal shapes; mixup (Zhang et al., ICLR 2018) trains networks on such blends of images and their labels - Averaging repeated trials of one stimulus estimates a mean response, whereas averaging two different images constructs a new input. A nonlinear network does not in general blend its internal responses the same way