Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 3 · Theme 1 — Representational spaces

PCA & eigendecomposition

Tuesday, September 22 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

Last time: comparing representational spaces

Left: the 120 animal photographs as a dissimilarity matrix built from people's ratings, blocked by animal category. Right: the two-dimensional map non-metric MDS recovers from that matrix, each point one photograph, coloured by category, with primates, birds and reptiles falling in separate regions

  • RDMs describe which stimuli a system treats as similar or different.
  • Comparing RDMs lets us ask: do two systems organize the same stimuli similarly?
  • MDS recovers a psychological space whose distances approximate the judged dissimilarities.

How many dimensions describe the variation?

A response matrix linked to the RDM and MDS map from the previous lecture. Below it, the new questions: what varies together, and how many dimensions describe the variation (the embedding dimension)?

  • MDS summarized dissimilarities as coordinates in a psychological space. Now return to the response matrix: hundreds of measured dimensions—pixels, features or neurons—for each stimulus.
  • If many responses vary together, their variation may occupy fewer dimensions: a smaller embedding dimension, the number of flat (linear) directions that hold most of it. How many are needed, and what is lost when we retain only some of them?

Can fewer coordinates capture the variation?

Which directions capture most of the variation, and what do we lose when we retain only a lower-dimensional description?

  • Start with two pixels from the same photographs.
  • Find a direction, project onto it, and reconstruct.
  • Use the same operations to study images and neural responses.

Why look for a direction at all?

A photograph at 64 by 64 with two adjacent centre pixels outlined, one blue and one magenta; below, the values of those two pixels across all 120 photographs form an elongated scatter of dots

  • Take 120 photographs and read the brightness at the same two pixel positions in each.
  • One dot is one photograph: its two coordinates are those two pixel values.
  • The values tend to rise and fall together. Could one number describe much of that change?

Why look for a direction at all?

The same figure again, now with a long shared-brightness direction approximately corresponding to the two pixels' average and a short local-contrast direction approximately corresponding to their difference

  • Along the long arrow, both pixels brighten or darken together. Its weights are close to (1,1)/2(1,1)/\sqrt{2}, so it behaves like their average and accounts for 90% of the variation.
  • Along the short arrow, one pixel brightens as the other darkens. Its weights are close to (1,−1)/2(1,-1)/\sqrt{2}, so it compares the two pixels and accounts for the remaining 10%.
  • For this pixel pair, the two fitted directions are close to average and difference.

From the same dots to one distribution

The same two-pixel scatter of dots from the previous slide appears on the left, with a vertical line at the mean x1 value of 101.6; an arrow labeled read only x1 leads to a histogram of those same 120 x1 coordinates with the identical mean marked

  • These are the same 120 dots. For the moment, read each dot's horizontal x1x_1 coordinate and set x2x_2 aside.
  • The histogram is the distribution of those 120 values. Their mean is 101.6; their spread tells us how much this measurement varies across photographs.

From observed values to a Gaussian model

A complete symmetric Gaussian bell curve from three standard deviations below its mean to three above, with the mean marked, the interval of one standard deviation shaded, and approximately 68 percent written inside it

  • The previous histogram showed observed pixel values. This is an idealized reference model, not an assumption today's method requires.
  • A Gaussian is smooth and bell-shaped. Its mean μ\mu gives the center and its standard deviation σ\sigma gives the spread; about 68% lies between μ−σ\mu-\sigma and μ+σ\mu+\sigma.

The Gaussian model in two dimensions

Five hundred points sampled from a two-dimensional Gaussian, forming a tilted ellipse of points centered at the mean, with one- and two-standard-deviation elliptical contours

  • Each simulated point contains two measurements, (x1,x2)(x_1,x_2). These are draws from the model, not the 120 photographs.
  • The mean becomes a mean vector μ\boldsymbol{\mu}, the center of the points; the covariance matrix describes their width, elongation and tilt.
  • PCA will find the directions along which the ellipse is longest and shortest.

Where are the points, and how do they spread?

The two neighboring pixels plotted in their raw zero-to-255 units, the points sitting well away from the origin, with a cross labeled mu = (mu1, mu2), marking its mean at approximately 102, 104

For each pixel jj, average over the N=120N=120 images:

  • Mean: μj=1N∑ixj(i)\displaystyle \mu_j=\frac{1}{N}\sum_i x_j^{(i)}
    The center of the points is μ=(μ1,μ2)\boldsymbol{\mu}=(\mu_1,\mu_2).
  • Variance: σj2=1N∑i(xj(i)−μj)2\displaystyle \sigma_j^2=\frac{1}{N}\sum_i\bigl(x_j^{(i)}-\mu_j\bigr)^2
    Spread around the mean along that coordinate.
  • Standard deviation: σj=σj2\sigma_j=\sqrt{\sigma_j^2}
    Spread in pixel-intensity units.

Where are the points, and how do they spread?

The same points after subtracting their mean, now sitting at the origin with the axes crossing through the middle of the points

  • So subtract the mean from every point: x~(i)=x(i)−μ\tilde{\mathbf{x}}^{(i)} = \mathbf{x}^{(i)} - \boldsymbol{\mu}. This is centering, and it slides the points until their centre sits at the origin.
  • PCA describes variation around the mean.
  • A negative centered value means “darker than this pixel’s average,” not a negative raw intensity.

Centering, in code

mean = features.mean(axis=0)     # (D,) mean response across stimuli, for each coordinate
centered = features - mean        # (N, D) each stimulus response minus the mean response
  • features is the response matrix from the first lectures: one row per stimulus (images here), one column per coordinate.
  • The mean is a vector with one entry per coordinate. Subtracting it from every row leaves each column summing to zero.
  • Keep the mean. PCA works on differences from it, and reconstruction adds it back.

Covariance: do the deviations go together?

The same centered pixel points and coordinate frame as before, with axes x1 minus mu1 and x2 minus mu2. Signs in each quadrant show whether multiplying the two deviations gives a positive or negative result.

  • These are the same centered responses: x~1=x1−μ1\tilde{x}_1=x_1-\mu_1 and x~2=x2−μ2\tilde{x}_2=x_2-\mu_2. Same signs give a positive product; opposite signs give a negative product.
  • Covariance is their average product:
    cov⁡(x~1,x~2)=1N∑ix~1(i)x~2(i)\displaystyle \operatorname{cov}(\tilde{x}_1,\tilde{x}_2)=\frac{1}{N}\sum_i\tilde{x}_1^{(i)}\tilde{x}_2^{(i)}.
  • Pearson correlation is covariance measured in standard-deviation units:
    r=cov⁡(x~1,x~2)σ1σ2\displaystyle r=\frac{\operatorname{cov}(\tilde{x}_1,\tilde{x}_2)}{\sigma_1\sigma_2}.

The covariance matrix

Three sets of two-dimensional points side by side, each with its two-by-two covariance matrix printed underneath: a round set with equal spread in both directions, an ellipse stretched along the horizontal axis, and a tilted ellipse that runs diagonally

  • In C\mathbf{C}, the diagonal contains the variances. The same covariance appears on both sides of the diagonal because cov⁡(x1,x2)=cov⁡(x2,x1)\operatorname{cov}(x_1,x_2)=\operatorname{cov}(x_2,x_1).
  • Geometrically C\mathbf{C} summarizes the points’ spread and orientation: round, an ellipse, or an ellipse tilted because the two vary together.
  • The ellipse's long direction is its direction of greatest variance.

Projection onto a unit direction

The two neighboring pixels, centered, with one point marked as the centered vector x tilde, a magenta line through the origin along the unit direction w, and the construction of the projection

  • Which direction gives the best one-number summary?
  • A unit direction w\mathbf{w} is any direction, written as a vector of length 1.
  • The projection length of a centered point is z=w⊤x~=∥x~∥cos⁡ϕz = \mathbf{w}^\top\tilde{\mathbf{x}} = \|\tilde{\mathbf{x}}\|\cos\phi: one number, its coordinate along w\mathbf{w}.

Projection onto a unit direction

The two neighboring pixels, centered, with one point marked as the centered vector x tilde, a magenta line through the origin along the unit direction w, and the construction of the projection

  • Drop a perpendicular from x~\tilde{\mathbf{x}} onto the line of w\mathbf{w}.
  • The dashed segment is what the direction does not capture: the part of the point left over once its position along w\mathbf{w} is known.

Projection onto a unit direction

The two neighboring pixels, centered, with one point marked as the centered vector x tilde, a magenta line through the origin along the unit direction w, and the construction of the projection

  • The foot of that perpendicular is p=(w⊤x~) w=z w\mathbf{p}=(\mathbf{w}^\top\tilde{\mathbf{x}})\,\mathbf{w} = z\,\mathbf{w}: the part of x~\tilde{\mathbf{x}} that lies along w\mathbf{w}.
  • Here K=1K=1: two coordinates become one score, zz.
  • More generally, KK retained directions replace DD coordinates with KK scores.

A direction keeps part of the spread

The two neighboring pixels again, a direction drawn at 110 degrees, and a short red segment from every point to the line showing what that direction does not capture

  • Project every point onto some direction. The red segments are what is left over: the part each point keeps to itself.
  • At this angle the direction keeps 23% of the total spread and leaves 77% over.

Turn it until the leftovers are smallest

The same points with the direction turned to 44 degrees, lengthwise through the points, and the red segments now short

  • Turn the direction until those segments are as short as possible: here at 44°, lengthwise through the points.
  • Now it keeps 90% and leaves 10%.

Spin the direction yourself

  • Before turning the direction, predict which direction leaves the smallest error.
  • Projected variance and squared reconstruction error add to a fixed total. Maximizing one minimizes the other.

From turning the line to a calculation

The familiar centered pixel points with the best direction at 44 degrees, retaining 90 percent of the variance and leaving the smallest residual

  • For a unit direction w\mathbf{w}, each centered response has one score: z=w⊤x~z=\mathbf{w}^\top\tilde{\mathbf{x}}.
  • Across all stimuli, those scores have variance

Var⁡(z)=w⊤Cw.\operatorname{Var}(z)=\mathbf{w}^\top\mathbf{C}\mathbf{w}.

  • The best direction is the w\mathbf{w} that makes this variance largest. The covariance matrix C\mathbf{C} lets us find it without searching every angle.

The covariance matrix changes a direction

A unit direction w and the vector Cw produced when the covariance matrix acts on it; Cw has a different length and orientation

  • The blue vector w\mathbf{w} is one possible direction through the response space.
  • Multiply it by the covariance matrix: Cw\mathbf{C}\mathbf{w} is the red vector.
  • Usually, the red vector has a different length and points in a different direction.

Most directions turn

The unit circle with several directions and their transformed vectors; their tips trace an ellipse and almost every vector changes orientation

  • Apply C\mathbf{C} to every direction on the unit circle. The transformed tips trace an ellipse.
  • Almost every direction comes out turned as well as rescaled.

Eigenvectors are directions that do not turn

Two special directions shown before and after multiplication by the covariance matrix; each transformed vector remains on the same dashed line

  • For most vectors, Cw\mathbf{Cw} points in a different direction from w\mathbf{w}.
  • An eigenvector vk\mathbf{v}_k is special: C\mathbf{C} changes its length but leaves it on the same line.
  • The amount of scaling is its eigenvalue λk\lambda_k:

Cvk=λkvk.\mathbf{C}\mathbf{v}_k=\lambda_k\mathbf{v}_k.

An eigenvalue gives variance along its eigenvector

The same centered pixel points and the same 44-degree line as the best-direction slide, now with the eigenvector v1 drawn along it as an arrow and the variance along it given as 90 percent of the total

  • Put w=vk\mathbf{w}=\mathbf{v}_k into the variance formula, then replace Cvk\mathbf{C}\mathbf{v}_k by λkvk\lambda_k\mathbf{v}_k:

Var⁡(zk)=vk⊤Cvk=vk⊤(λkvk)=λk vk⊤vk=λk,\operatorname{Var}(z_k)=\mathbf{v}_k^\top\mathbf{C}\mathbf{v}_k=\mathbf{v}_k^\top(\lambda_k\mathbf{v}_k)=\lambda_k\,\mathbf{v}_k^\top\mathbf{v}_k=\lambda_k,

since vk\mathbf{v}_k has length 1.

  • The largest eigenvalue identifies the direction with the most variance. Here it is the same 44° direction found by turning the line.
  • Ranked by eigenvalue, these directions are the principal components. The first is PC1; computing them is principal component analysis (PCA).

PCA = eigendecomposition of C\mathbf{C}

Covariance (centered, NN samples):   C=1N∑ix~(i)x~(i)⊤\;\mathbf{C} = \dfrac{1}{N}\sum_{i} \tilde{\mathbf{x}}^{(i)}\tilde{\mathbf{x}}^{(i)\top}   (code divides by N−1N-1; same eigenvectors)

Eigendecomposition (orthonormal V\textcolor{#B5396B}{\mathbf{V}}):

C=V Λ V⊤,C vk=λk vk,λ1≥λ2≥⋯≥λD\mathbf{C} = \textcolor{#B5396B}{\mathbf{V}}\,\textcolor{#3F7A6B}{\boldsymbol{\Lambda}}\,\textcolor{#B5396B}{\mathbf{V}}^\top,\qquad \mathbf{C}\,\textcolor{#B5396B}{\mathbf{v}_k} = \textcolor{#3F7A6B}{\lambda_k}\,\textcolor{#B5396B}{\mathbf{v}_k},\qquad \textcolor{#3F7A6B}{\lambda_1}\ge\textcolor{#3F7A6B}{\lambda_2}\ge\cdots\ge\textcolor{#3F7A6B}{\lambda_D}

  • Principal directions = eigenvectors vk\textcolor{#B5396B}{\mathbf{v}_k} of C\mathbf{C}.
  • Variance along vk\textcolor{#B5396B}{\mathbf{v}_k} = λk\textcolor{#3F7A6B}{\lambda_k}.
  • PC1=v1\text{PC1} = \textcolor{#B5396B}{\mathbf{v}_1}, the eigenvector with the largest eigenvalue.

From a line to a plane

A three-dimensional set of points near a plane spanned by two principal directions; one point and its projection on the plane

  • In this picture, D=3D=3 measured coordinates and we retain K=2K=2 perpendicular directions.
  • The two directions define a plane; each centered point gets two scores, (z1,z2)(z_1,z_2).
  • Keeping all DD directions changes coordinates without losing information.
  • Keeping K<DK<D projects each point onto a lower-dimensional flat (linear) subspace: a line, a plane, or its analogue in more dimensions.

The new coordinates are the PC scores

The same three-dimensional centered points and retained plane; the highlighted point projects to a location with coordinates z1 and z2 in the principal-component basis

For one stimulus, take its dot product with each retained direction:

zk=vk⊤(x−μ).z_k=\mathbf{v}_k^\top(\mathbf{x}-\boldsymbol{\mu}).

  • vk\mathbf{v}_k gives weights on the original measurements.
  • The score zkz_k is this stimulus's coordinate along that direction.
  • Keeping KK directions gives the point KK scores.

The same points, in the new coordinates

A three-dimensional set of points lying close to a plane spanned by v1 and v2, with axes x-tilde 1, 2 and 3; one point off the plane and its projection The same 100 points drawn in two dimensions at their scores z1 and z2 along v1 and v2; the projected point is marked in green
  • Left: points in the original measured coordinates x~1,x~2,x~3\tilde{x}_1, \tilde{x}_2, \tilde{x}_3. Right: the same points at their scores (z1,z2)(z_1, z_2): the PCA embedding.
  • The green point is the same stimulus in both. Its distance from the plane, the dashed red residual, is what the 2-D map leaves out.

Reconstruction: the way back

The raw two-pixel points with their mean, one original point, and its reconstruction formed by adding the score-weighted first principal direction to the mean; the residual joins the reconstruction to the original

  • Return to the original measurement coordinates and start from the mean.
  • Add each retained direction, weighted by the image's score:

x^=μ+∑k=1Kzkvk.\hat{\mathbf{x}}=\boldsymbol{\mu}+\sum_{k=1}^{K}z_k\mathbf{v}_k.

  • The residual is the difference between the original and reconstructed measurements.

What do we lose when we keep fewer components?

Top: one face reconstructed with K equal to 1, 5, 20, 50, 150 and 400 components, from a blurry near-average to the exact original. Bottom: the residual fraction for this face at each K as dots, compared with the eigenvalue tail computed from the whole dataset

  • PCA on the pixels of 400 faces: keeping more components adds detail and reduces reconstruction error.
  • A residual fraction of 0 means exact reconstruction; larger values mean more of that face was left out.
  • Dots: this face. Curve: the discarded variance across all 400 faces, which need not equal one face's error.

How many components? The scree plot

Scree plot for the 400 Olivetti faces: the variance fraction carried by each principal component, and the cumulative fraction

  • These are the same 400 faces. The left panel shows the variance fraction carried by each PC.
  • The cumulative curve answers how many PCs are needed to retain a chosen fraction, such as 90%: the embedding dimension at that threshold.
  • A variance threshold sets a reconstruction budget. It does not establish performance on a task.

PCA in scikit-learn

from sklearn.decomposition import PCA

pca = PCA(n_components=K, svd_solver="full")
Z = pca.fit_transform(X)       # X: (N, D)  ->  Z: (N, K)
fraction = pca.explained_variance_ratio_
directions = pca.components_  # one principal direction per row
  • Score: a stimulus's coordinate along a principal direction.
  • fit_transform centers each feature across stimuli. It does not z-score them (divide each by its standard deviation).
  • SVD computes the PCA directions and variances from the centered data without explicitly forming its covariance matrix.

Can these directions help explain a neuron?

Three example IT cells from Bao and colleagues, each shown as a 24-view by 51-object response matrix with four views of the cell's preferred object picked out and shown underneath: an aeroplane, a swan and a basket

  • Inferotemporal (IT) cortex is part of the primate visual pathway involved in object recognition.
  • Each matrix shows one cell's responses: 51 objects × 24 views. A bright column means strong responses across views of one object.
  • Can a space derived from a network's responses help organize these preferences?
Bao et al. Nature 2020 · Fig. 3b, three of the example cells

Build an object space from network responses

Bao et al.'s schematic of the AlexNet object space: stimulus images arranged in the plane of the first two PCs, in four colored groups

  • Present the 1,224 images to AlexNet and collect responses from fc6, a late network layer.
  • Run PCA on those response vectors. In the paper, 50 components account for 85% of the variance.
  • The picture is a schematic of the first two PCs, arranged for legibility rather than a raw scatter plot.
  • “Stubby–spiky” and “animate–inanimate” are proposed interpretations of those directions. They need testing.

Test whether a direction predicts a response

Schematic simulated example: projecting points in model space onto a direction gives a scalar related to a simulated neural response

  • The gold ring follows one image from its position in model space to its measured response.
  • Choose a direction using training images, then ask whether each image's projection predicts the cell's response on held-out images, which were not used to choose it.
  • The preferred direction can combine several PCs; it need not coincide with PC1 or PC2.

Neural preferences occupy different regions

Grey stimulus points in the PC1-PC2 plane with four colored groups of points — the preferred images of four IT patch systems — landing one per quadrant

  • Grey circles are stimuli in the network-derived PC plane.
  • Colored points are images that strongly activate four IT systems. NML is the network Bao et al. found in IT's "no man's land", cortex that no earlier localizer had assigned a category; the numbers in parentheses count recorded neurons.
  • Different IT systems prefer different regions of the model space. A separate analysis tests whether a neuron's response varies with its projection along a direction.

Use the space to make a new prediction

Published coronal fMRI maps from two monkeys, showing the four networks at posterior, middle and anterior IT; magenta marks the predicted stubby-object network

  • The object space predicted a system responding to stubby, inanimate objects.
  • fMRI located it (magenta); targeted recordings confirmed the object preferences.
  • A description of a representation can guide a new experiment.

The measurement scale changes PCA

The same correlated points twice. Top, in raw units with feature 1 in millimetres and feature 2 in metres, PC1 equals 1.00 times feature 1 plus 0.00 times feature 2 and carries 100 percent of the variance. Bottom, after z-scoring, PC1 is 0.71 times each feature and carries 89 percent

  • PCA maximizes variance, so a feature measured in millimetres dominates one measured in metres: in raw units PC1 is 1.00 x1+0.00 x21.00\,x_1 + 0.00\,x_2 and the second feature is invisible.
  • After z-scoring each feature (subtract its mean, divide by its standard deviation), PC1 is 0.71 x1+0.71 x20.71\,x_1 + 0.71\,x_2 and recovers the diagonal structure the two features share.
  • Standardize when the units are not comparable. When they are comparable and the variances differ for real reasons, standardizing changes the relative weighting of features; it is a modelling decision, not a default.

Most variance need not mean most useful

Left: the 120 photographs laid out along the leading direction of their pixel coordinates, running from dark pictures to bright ones with animals of different categories interspersed. Right: position along that direction plotted against the mean pixel value of each photograph, an almost perfect straight line, r equals 0.99

  • The largest direction in the 120 photographs carries 29% of the variation and is strongly related to overall brightness: ordered along it they run dark to bright, different categories interspersed.
  • Would this direction help distinguish the animals? This plot alone cannot answer.
  • PCA ranks directions by variance; the task determines what is useful.

Which direction would you keep?

Toy classes with a high-variance horizontal direction and low-variance vertical separation; projection panels ask for a prediction

  • The groups vary mainly from left to right. Their labels differ from top to bottom.
  • Which direction would best reconstruct the points?
  • Which direction would best distinguish the groups?

Low variance can still separate the groups

Toy classes with a high-variance horizontal direction and low-variance vertical separation; PC1 scores overlap while PC2 scores separate the classes

  • PC1 retains 97.8% of variance, but both groups have the same PC1 coordinates.
  • PC2 retains 2.2%, yet a threshold at zero separates the groups.
  • Variance measures spread. Task relevance depends on what we want to distinguish.

What PCA provides

The two neighboring pixels again, with the direction that retains about 90 percent of their variance

  • PCA replaces many measurements with ordered scores and tells us how much variance each direction retains.
  • Those scores support reconstruction and can connect a model representation to neural responses. In the Bao study, that connection helped motivate a new experiment.
  • PCA describes variation. A separate test establishes what the variation is useful for.

What if the variation follows a curve?

Rotated, shifted and resized versions of an image projected into two PCA directions; rotation traces a closed curve

  • Changing one stimulus parameter can trace a curved path in a large response space.
  • Even one varying parameter can require two linear coordinates when the path bends back on itself.
  • PCA approximates data with a flat subspace.
  • Next: nonlinear maps, and how to check what their geometry preserves.

Further reading for the research trail

Handout: Statistics and the mathematics of PCA — variance, covariance, eigenvectors, projection and reconstruction, worked by hand in this lecture’s notation.

  • Bao et al. (2020): object-space directions and IT responses. Nature
  • Chang & Tsao (2017): face-space directions and individual neurons. Cell
  • Turk & Pentland (1991): eigenfaces. J. Cogn. Neurosci.
  • DiCarlo & Cox (2007): untangling object representations. Trends Cogn. Sci.
  • Stringer et al. (2019): a structured, high-dimensional neural spectrum. Nature

Trail entry: pick any paper cited in a lecture so far, read it, and post four to six sentences on Ed under Research trail — what it did, and what surprised you or what question it left you with. Reply once to a classmate. Next deadline Fri 9/25, best three of four entries count.

Minute paper

Submit on Canvas → Minute papers → Minute Paper 4.

Write three brief points in your own words:

  1. Why do we center the data before finding the principal directions?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Appendix: examples and implementation details

Additional examples, the PCA/SVD implementation, and reference material for the assignments.

One direction was not enough

Left: the same 120 photographs placed on a single direction, coloured by animal category, with the categories piled on top of one another. Right: the distortion of the ratings plotted against the number of directions kept, falling steeply from one to two and then flattening out

  • Fit the same ratings on one direction and the categories pile up: the distances in the fit no longer match the judged dissimilarities.
  • A second direction removes most of that distortion. A third and a fourth remove very little. Two directions greatly improve the fit, but some distortion remains.

How far does the dependence reach?

Three panels. Left: a pixel plotted against its immediate neighbour across the 120 photographs, the points hugging the diagonal at r equals 0.81. Middle: the same pixel against the one twenty places away, points with no visible trend at r equals 0.08. Right: the correlation plotted against separation, falling steeply over the first few pixels and flattening near zero

  • The same pixel-pair computation, at larger separations: one pixel against another, across the same 120 photographs.
  • Next door, the two pixel values strongly covary. Twenty pixels apart, they barely covary.
  • Neighbouring measurements are strongly dependent, and the dependence is local.

PCA in code: two routes to the same answer

centered = features - features.mean(axis=0)

# Route 1: eigendecomposition of sample covariance
C = np.cov(centered.T)
lam, V = np.linalg.eigh(C)                 # ascending order
lam, V = lam[::-1], V[:, ::-1]             # descending order

# Route 2: SVD, without forming C
_, s, Vt = np.linalg.svd(centered, full_matrices=False)
lam2 = s**2 / (len(centered) - 1)
V2 = Vt.T                                # directions as columns
Z = centered @ V2[:, :K]                   # (N, K) scores
  • Use eigh, not eig: C\mathbf{C} is symmetric, so the eigenvalues are real and eigh is faster and stabler. It returns them ascending, so the reversal is not optional.
  • Prefer route 2. Forming C\mathbf{C} costs D×DD\times D memory; for a 256×256256\times256 color image D≈196,608D\approx196{,}608 and C\mathbf{C} would be ~309 GB in float64. SVD never builds it. Sign is arbitrary: vk\mathbf{v}_k and −vk-\mathbf{v}_k are both valid; Assignment 2's helper fixes it so the largest entry is positive.

Eigenfaces

PCA on 400 faces, and a classic approach to face recognition.

PCA as compression: eigenfaces

Grid of the mean face and leading eigenfaces, with opposing signs shown in red and blue

Stack face images as vectors and run PCA → the leading eigenvectors are eigenfaces: ghostly, face-like basis images, one direction in pixel space each.

Turk & Pentland J. Cogn. Neurosci. 1991 · precursor: Sirovich & Kirby J. Opt. Soc. Am. A 1987 · Olivetti faces (AT&T Laboratories Cambridge)

Face space

Twelve face photographs placed on a plane at their coordinates along the first two eigenfaces, z1 on the horizontal axis and z2 on the vertical

  • 400 face images (64×6464\times64 pixels, so D=4096D = 4096), PCA on the centered pixels, and twelve of the faces drawn at their coordinates (z1,z2)(z_1, z_2) on the first two eigenvectors.
  • Two numbers per face instead of 4,096. They carry 24% and 14% of the variance: a coarse, lossy summary, but already one in which lighting and head pose vary smoothly across the plane.
  • This is a face space: every face is a point, and nearby points are faces that look alike to the two leading eigenfaces. The preceding slide shows those eigenfaces as images.

A face = mean + a few eigenfaces

A face reconstructed using 1, 5, 25, 100 and 400 components, beside the original image

  • These faces can be approximated by the mean face plus a weighted sum of eigenfaces. More components add finer detail.
  • Turk & Pentland used the coefficients for recognition: a new face is matched to the stored face with the nearest coefficients.

Mix your own face

  • Each slider sets one coefficient in mean + weighted eigenfaces.
  • Set a coefficient to zero to remove that direction’s contribution.

Two views, one answer

  • Maximize the variance of the projections: max⁡∥v∥=1v⊤C v\displaystyle\max_{\lVert\mathbf{v}\rVert=1}\mathbf{v}^\top\mathbf{C}\,\mathbf{v}
  • Minimize the squared reconstruction error: min⁡∥v∥=11N∑i∥x~(i)−(v⊤x~(i))v∥2\displaystyle\min_{\lVert\mathbf{v}\rVert=1}\tfrac{1}{N}\sum_i\lVert\tilde{\mathbf{x}}^{(i)} - (\mathbf{v}^\top\tilde{\mathbf{x}}^{(i)})\mathbf{v}\rVert^2
  • These are the same problem: total variance is fixed, so whatever the projection captures, the residual loses. Both are solved by the top eigenvector of C\mathbf{C}.
  • For KK retained PCs, the average squared fitting-data reconstruction error is ∑k>Kλk\sum_{k>K}\lambda_k (using covariance divided by NN).
  • Connection to MDS: classical MDS on a Euclidean dissimilarity matrix returns the PCA coordinates up to an orthogonal transformation of the centered data.

PCA as decorrelation & whitening

  • Project onto the eigenbasis and the new coordinates are uncorrelated: the covariance becomes diagonal. For positive eigenvalues, rescale each retained direction by 1/λi1/\sqrt{\lambda_i} and you get whitening: zero correlation, unit variance.

C=VΛV⊤    ⇒    z=Λ−1/2V⊤(x−μ),Cov(z)=I\mathbf{C} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^\top \;\;\Rightarrow\;\; \mathbf{z} = \boldsymbol{\Lambda}^{-1/2}\mathbf{V}^\top(\mathbf{x}-\boldsymbol{\mu}), \quad \mathrm{Cov}(\mathbf{z}) = \mathbf{I}

  • Why analysts reach for it:
    • remove redundancy between measured dimensions before modeling
    • decorrelate inputs, without assuming the resulting features are independent
    • compress by dropping directions; whether this denoises depends on what they contain

What do the components look like?

Principal components of synthetic image patterns, shown as red and blue signed weights: structured patterns dominate early components and noise appears in later ones

  • Each eigenvector is a vector in the original space, so each PC unrolls into a picture. Here the inputs are synthetic patterns: early components recover structured variation; later ones contain little variance.
  • The components can be read as weights on the original variables, then tested for what they capture.
Generated after Murphy Probabilistic Machine Learning 2022 ch. 20 (MIT License)

The rainbow cortex is three numbers per voxel

Panels from Huth et al. showing word-related features above cortical surfaces, where every voxel is colored by its first three PCA coordinates (red for PC1, green for PC2, blue for PC3), producing a smooth rainbow map across cortex

  • Over two hours of narrated stories, one regression per voxel: every voxel gets a vector of weights on word-related features, and cortex is one more N×DN\times D matrix.
  • PCA on it, keep three components, color each voxel by its coordinates: PC1 → red, PC2 → green, PC3 → blue. The palette is the projection z\mathbf{z}.

PCA finds flat subspaces

Top: two interleaved spiral arms, one blue and one green, with the PC1 line drawn through them. Bottom: the histogram of the coordinates along PC1, where the two arms overlap almost completely

  • If the data lie on a curved manifold, a smooth curved surface (a spiral, a Swiss roll, a torus), a low-dimensional linear projection can fold distinct regions onto each other.
  • Two points far apart along the manifold can land on top of each other after projection: the two arms are perfectly distinct in the plane and inseparable along PC1.
  • t-SNE and UMAP construct nonlinear maps emphasizing local relationships. Their displays can distort global structure and need their own checks.

- Same 120 photographs and ratings as Assignment 1 - Non-metric MDS fits the rank order of the dissimilarities, not their values - Orientation is arbitrary, and different initializations can reach different solutions: your notebook map may differ from this one by more than a rotation - The category blocks: dissimilarities are smaller within a category than between categories

- PCA works on the response matrix itself (stimuli × pixels, features or neurons), not on an RDM computed from it - Question: is the variation across many measured dimensions concentrated along a few directions? Every variable is still recorded; PCA changes the description, not the measurement - Because the response matrix names the measured variables, each PCA direction is a set of weights on pixels, features or units - Behavioral dissimilarities can come directly from judgments, with no response matrix behind them - Embedding dimension: the number of principal components needed to account for the variance at a chosen threshold, not a count of true latent causes. The nonlinear-dimensionality lecture names it formally, next to the ambient and intrinsic dimensions

- PCA chooses directions by variance; reconstruction error measures the cost of keeping fewer of them - Whether the retained directions help with a particular task is a separate question

- Intensities at one fixed pixel pair (row 33, columns 32 and 33, counting from 1) across the 120 photographs, resized to 64 × 64 - Raw values, 0 to 255; correlation r = 0.81

- Leading direction ≈ 0.71 x₁ + 0.70 x₂: a local brightness measurement, 90% of the variance - Perpendicular direction ≈ 0.70 x₁ − 0.71 x₂ (overall sign arbitrary): compares the two pixels, 10% - The weights come out close to average and difference because the two pixels have similar variances and strong positive covariance; "average" and "difference" describe these fitted directions and are not the definition of PC1 and PC2

- A distribution records how often each value occurs; here, the x₁ values of the same 120 photographs, mean 101.6 - Across photographs, a pixel intensity is a random variable: a quantity whose value changes from one photograph to the next - Each bar is the percentage of photographs in that intensity bin; the bars sum to 100% (a percentage per bin, not a probability density) - Center and spread are defined on one variable first, then carried back to the two pixels together

- The histogram was an empirical distribution: the values observed in 120 photographs. A Gaussian is a mathematical model of all the values one could observe - Drawn in standard units from −3σ to +3σ; the tails continue in both directions - PCA does not require Gaussian data

- Synthetic: 500 draws from a two-dimensional Gaussian - For a Gaussian, the mean vector and covariance matrix specify the whole distribution; for other data they summarize only center and spread - The ellipses mark constant standardized distance from the mean; in two dimensions they do not enclose 68% and 95%

- Each image i gives two raw intensities, x₁⁽ⁱ⁾ and x₂⁽ⁱ⁾; j indexes the pixel - Averaging over images gives one mean per pixel; together they form the mean vector, ≈ (102, 104) - Variance and standard deviation describe spread along each coordinate separately; covariance describes how the two vary together - Dividing by N summarizes this dataset; the sample estimator divides by N − 1

- Centering moves every point by the same vector; distances between points are unchanged - Without centering, the decomposition measures spread around the origin, and the mean can dominate the leading direction - Keep the mean: reconstruction adds it back

- `features` is N × D, the same table of responses as in the first lectures; the mean has length D, so the subtraction runs column by column - The helper in Assignment 2 starts with these two lines; scikit-learn's PCA centers the data the same way internally

- A tilde marks a centered response: x̃ⱼ⁽ⁱ⁾ = xⱼ⁽ⁱ⁾ − μⱼ, both coordinates measured across the same images - Positive covariance: same-signed deviations outweigh opposite-signed ones - Pearson correlation (previous lecture) divides covariance by the two standard deviations; for z-scored coordinates the two are equal. Covariance carries units; correlation has none - A coordinate paired with itself gives its variance (the diagonal of the covariance matrix); swapping the order gives the same covariance, so the matrix is symmetric - Neither implies causation, and zero covariance does not imply independence

- One variance per feature on the diagonal, one covariance per pair off it - Simulated Gaussian data with set covariances - Unequal variances alone give PCA a preferred direction, even with zero covariance; for the round set of points every direction is equally good - The ellipse captures second-order spread only, not all structure in a distribution

- The tilde marks a centered response vector - z is a signed coordinate: negative when the point lies on the opposite side of the origin from w - The arrow for w is enlarged for visibility; it is not on the pixel-intensity scale

- The dashed segment is the residual: the part of x̃ that one score cannot reconstruct - PCA picks the direction that minimizes the sum of squared residual lengths over the data

- For the marked point: z ≈ 200 (centered pixel units), residual length ≈ 48 - z is a number, the coordinate; z w is a vector, the reconstruction along w - D: number of measurements per stimulus; K: number of retained directions

- Same points, with each one's residual drawn in red - A direction across the points leaves long residuals: a poor summary

- Best direction ≈ 44.4°, keeping ≈ 90.4% of the total variance - Minimizing the summed squared residuals and maximizing the retained variance give the same direction

- The demo opens on the same two pixels across the 120 photographs - Turn w: retained variance and squared reconstruction error always sum to the total variance - The menu offers other feature pairs (half-image means; mean intensity against contrast); each is a different set of points with its own best direction - The demo divides by N and `np.cov` by N − 1; the constant factor changes neither the variance fractions nor the directions

- z is the projection coordinate from the earlier slides - For centered data, Var(z) is the mean of z²; substituting z = wᵀx̃ gives wᵀCw - The demo searched over angles; the eigendecomposition solves the same maximization directly

- Blue: an input direction, not a data point. Red: the same vector after multiplication by C - C is the pixel example's covariance, divided by a positive constant to fit on the slide; this keeps the directions and changes the eigenvalues, so no stretch factors are shown - Most input directions change both length and orientation

- The unit circle is a set of input directions, not data - The red ellipse shows what C does to those directions; it is not a covariance contour of the data - Both share the principal directions; the half-lengths of this ellipse scale with λ, those of a covariance contour with √λ

- On each dashed line, the input vector and its transformed version lie on the same line: C changes only the length, by the factor λ - Rescaling C to fit the drawing leaves the eigenvectors unchanged - In the pixel example, the first eigenvector lies at ≈ 44°, the direction found by turning the line in the demo

- For a unit eigenvector, vᵀv = 1, so its projected variance equals its eigenvalue - The largest eigenvalue is the maximum of wᵀCw over all unit directions (a standard linear-algebra result) - Turning the line, minimizing reconstruction error and the eigendecomposition select the same 44° direction; variance fractions ≈ 90% and 10%

- Recipe: center, compute the covariance, order its eigenvectors by eigenvalue - Eigenvectors give directions; eigenvalues give the variance along them - SVD of the centered data gives the same quantities without forming C (appendix)

- Schematic, centered: D = 3, K = 2 - Orthonormal basis: mutually perpendicular unit directions; a complete basis spans the whole space - The highlighted point lies off the plane; its projection is given by the two scores z₁ and z₂

- With the retained directions as the columns of V_K: z = V_Kᵀ(x − μ) - For the whole dataset (rows = stimuli): Z = X_c V_K, one row per stimulus, K columns - Weights (v_k, one entry per measurement) and scores (z_k, one per stimulus) are different quantities

- Both figures show the same 100 centered points. On the left they live in the three measured coordinates; on the right each is placed at its two scores, its coordinates along v1 and v2 - Going from left to right is the projection z = V_K transpose (x minus mu); the third coordinate, perpendicular to the plane, is dropped. Here it holds 2.1% of the squared length, so little is lost - This 2-D plot is what a PCA figure in a paper shows. The axes are no longer measurements such as pixels or neurons but the principal directions

- The axes are back in raw pixel units, not centered ones - Green arrow: from μ to x̂, that is z₁v₁. Dashed segment: the residual x − x̂, perpendicular to v₁ - Keeping all D directions reconstructs x exactly - The same equation is the decoder of a linear autoencoder (future lectures)

- Residual fraction for one face: ‖x − x̂‖² / ‖x − μ‖² - Eigenvalue tail: summed squared residuals over all 400 faces, divided by their total centered squared norm - The tail is not the average of the per-face fractions, and not any one face's fraction

- The elbow is a heuristic, not a statistical test - Same face matrix as the previous slide: the cumulative curve is the dataset-wide counterpart of the discarded-variance curve - In practice K is often set by the next analysis: keep as many components as it needs - Theme 2 asks the same question of neural population responses

- X is N × D, Z is N × K; choose K before running - PCA uses no labels: it learns directions from X alone - `svd_solver="full"` makes the result deterministic; Assignment 1 may use the default solver - The sign of each direction is arbitrary: flipping a direction and its scores leaves the reconstruction unchanged, so compare absolute correlations

- Each matrix shows identity and viewpoint together - Bright columns suggest view tolerance; cells can still prefer some views - The 1,224 centered image responses span at most 1,223 directions (their rank); the 51 object means, at most 50. Averaging over views changes what variation PCA sees

- PCA is fitted to AlexNet responses, not to the IT recordings - PCA gets no category labels, although AlexNet was trained with object labels - Assignment 1 tests the proposed interpretations of these directions

- Simulated to illustrate the test; not a fit from Bao et al. - A fitted direction links the network's responses (a population of units) to one cell's responses - Fit the direction on some images and evaluate it on others; otherwise a flexible model fits noise - The projection is the dot product from earlier in the lecture

- Colored dots are highly effective images for each cortical network; numbers in parentheses count recorded neurons - The result that each neuron responds along one direction ("axis coding" in the paper) comes from a separate analysis relating each neuron's response to projections along object-space directions - Assignment 1 covers the PCA plane and preferred directions, not every experiment in the paper - Bao et al. (2020), https://doi.org/10.1038/s41586-020-2350-5

- Rows: monkeys M3 and M4; columns: posterior, middle and anterior IT; colors as in the object-space plot - fMRI maps thresholded at P < .001, uncorrected - The face, body and NML networks motivated the prediction of a fourth network for stubby, inanimate objects; fMRI located it and targeted recordings confirmed its preferences - Bao et al. (2020), https://doi.org/10.1038/s41586-020-2350-5

- Rescaling a feature changes the covariance and can rotate the principal directions - Same observations in both panels - With known, nonzero scales, z-scoring is invertible: nothing is lost, but distances, PCA and regularized fits change

- The leading pixel-space score correlates with mean brightness at r = 0.99 - Brightness here mixes background, object and illumination - Whether the direction also carries category information needs a readout (a weighted sum of the retained components, then a threshold) fit on some images and tested on held-out ones

- Constructed toy data: the two groups share their horizontal coordinates; their vertical coordinates have opposite signs

- PCA ranks the horizontal direction first, yet the label depends only on the vertical coordinate - The small direction separates the groups perfectly because the label rule is noiseless; real readouts need held-out evaluation - In the score panels, identical PC1 markers overlap; the vertical jitter in the PC2 panel only makes repeated scores visible

- PCA gives ordered coordinates, a reconstruction and an account of the variance - In Bao et al., those coordinates supported a hypothesis about IT responses and led to a new experiment - What the coordinates are useful for needs a task or neural-response test; a low-dimensional description does not by itself measure task information

- Real computations on rotated, shifted and resized versions of one photograph, all projected onto two PCs fitted to the rotations - One varying parameter can trace a curve that needs more than one linear direction - Nonlinear maps (next lecture) give other summaries; an attractive map need not preserve global distances

- Huth et al. (2016), semantic maps: https://doi.org/10.1038/nature17637 - Gardner et al. (2022), grid-cell population geometry: https://doi.org/10.1038/s41586-021-04268-7 - Perich (2026), interpreting dimensionality: https://doi.org/10.53053/FTEU9810 - Marks et al. (2026), PCA of hippocampal measurements: https://doi.org/10.1002/brb3.71666 - Chang & Tsao use shape and appearance coordinates, not raw-pixel eigenfaces; Stringer et al.'s high-dimensional spectrum is not evidence that only a few dimensions matter

- Canvas → Minute papers → Minute Paper 4 - Credit is for a thoughtful attempt, not for being correct

- Stress falls from ≈ 0.50 with one dimension to ≈ 0.20 with two, then more slowly - Non-metric MDS stress compares map distances with fitted disparities that keep the rank order of the dissimilarities - A good two-dimensional approximation, not proof that the judgments have exactly two underlying dimensions - Vertical jitter in the one-dimensional panel only separates overlapping markers

- Same fixed pixel, same 120 photographs, same Pearson correlation, at increasing separation - Computed across images, not within one image: this across-image covariance is what the covariance matrix holds - Independent random pixels would give zero at every separation; photographs give 0.81 next door and 0.08 twenty pixels away - This local dependence is why a photograph is not a random point in pixel space

- Both routes give the same nonzero variances and principal subspaces, up to numerical error - Signs are arbitrary; with repeated eigenvalues, any orthonormal basis of the shared subspace is valid - With N < D, reduced SVD returns N columns; the covariance route also returns a large null space - A 196,608 × 196,608 float64 covariance takes ≈ 309 GB (288 GiB) before any workspace - N against N − 1 changes the variances by a constant factor; scikit-learn's solver depends on its settings and version

- Eigenfaces: a face as the mean image plus weighted directions of variation in pixel space - Turk & Pentland (1991), Eigenfaces for recognition, J. Cogn. Neurosci., https://doi.org/10.1162/jocn.1991.3.1.71

- Eigenfaces are signed weights on pixels, so each renders as grey with light and dark lobes rather than as a face - The first few capture overall brightness, lighting direction and coarse head shape; later ones carry finer detail - Olivetti set: 400 images of 40 people, 64 × 64 pixels

- Each of the 400 Olivetti images is a point in 4,096-dimensional pixel space; projecting onto the first two eigenvectors places it on this plane - The first two components carry 24% and 14% of the variance: a rough summary, but not a random arrangement - Brightness and lighting direction vary along the first direction; pose and face shape along the second - The eigenfaces are the direction weights; the scores place each face on the plane

- The reconstruction equation as images: x̂ = μ + Σₖ zₖ vₖ - 4,096 pixels down to about 100 coefficients, a forty-fold reduction, and the face is still recognisable

- Mix: the mean face plus or minus multiples of the first eight eigenfaces; each slider spans three standard deviations of that coefficient across the dataset - K: one of twelve stored faces rebuilt from its first K coefficients, with the residual fraction beside the eigenvalue tail - Plane: click a point on the PC1–PC2 plane to draw the face with exactly those two coordinates

- Classical MDS on Euclidean distances returns the PCA coordinates, up to an orthogonal transformation - The previous lecture taught the stress-minimising route to MDS, not this eigendecomposition route; the connection is for reference, not something to derive

- Whitening applies only to directions with positive eigenvalues; a zero-variance direction cannot be divided by its standard deviation - Small eigenvalues amplify noise; in practice, truncate or regularize - Uncorrelated features need not be independent - Whitening differs from z-scoring the original features

- A principal component of an image dataset is a vector in pixel space, so it can be displayed as an image - Synthetic patterns here, not the faces

- Huth et al. (2016), Natural speech reveals the semantic maps that tile human cerebral cortex, Nature, https://doi.org/10.1038/nature17637 - Each voxel's responses to the stories were modeled from word-related features; PCA ran on the fitted weights, and three component scores became the color channels - The colors encode a projection, not three named psychological properties; semantic labels come from examining the weights and predictions

- The two spiral arms overlap when projected onto PC1 - Nonlinear maps (next lecture) can show other relationships, but their displays also need checking; visual separation in a map is not a classifier