PCA & eigendecomposition

Reader-friendly version of the 60 lecture slides: the lecture content in normal flow, with figure descriptions and data tables where the figure carries data. Open the presentation.

Slide 1

Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 3 Β· Theme 1 β€” Representational spaces

PCA & eigendecomposition

Tuesday, September 22 Β· Fall 2026

Return to contents

Slide 2

Last time: comparing representational spaces

Left: the 120 animal photographs as a dissimilarity matrix built from people's ratings, blocked by animal category. Right: the two-dimensional map non-metric MDS recovers from that matrix, each point one photograph, coloured by category, with primates, birds and reptiles falling in separate regions

Left: the 120 animal photographs as a dissimilarity matrix built from people's ratings, blocked by animal category. Right: the two-dimensional map non-metric MDS recovers from that matrix, each point one photograph, coloured by category, with primates, birds and reptiles falling in separate regions

Open full-size figure

Return to contents

Slide 3

How many dimensions describe the variation?

A response matrix linked to the RDM and MDS map from the previous lecture. Below it, the new questions: what varies together, and how many dimensions describe the variation (the embedding dimension)?

A response matrix linked to the RDM and MDS map from the previous lecture. Below it, the new questions: what varies together, and how many dimensions describe the variation (the embedding dimension)?

Open full-size figure

Return to contents

Slide 4

Can fewer coordinates capture the variation?

Which directions capture most of the variation, and what do we lose when we retain only a lower-dimensional description?

Return to contents

Slide 6

Why look for a direction at all?

The same figure again, now with a long shared-brightness direction approximately corresponding to the two pixels' average and a short local-contrast direction approximately corresponding to their difference

The same figure again, now with a long shared-brightness direction approximately corresponding to the two pixels' average and a short local-contrast direction approximately corresponding to their difference

Open full-size figure
  • Along the long arrow, both pixels brighten or darken together. Its weights are close to (1,1)/2(1,1)/\sqrt{2}, so it behaves like their average and accounts for 90% of the variation.
  • Along the short arrow, one pixel brightens as the other darkens. Its weights are close to (1,βˆ’1)/2(1,-1)/\sqrt{2}, so it compares the two pixels and accounts for the remaining 10%.
  • For this pixel pair, the two fitted directions are close to average and difference.

Return to contents

Slide 7

From the same dots to one distribution

The same two-pixel scatter of dots from the previous slide appears on the left, with a vertical line at the mean x1 value of 101.6; an arrow labeled read only x1 leads to a histogram of those same 120 x1 coordinates with the identical mean marked

The same two-pixel scatter of dots from the previous slide appears on the left, with a vertical line at the mean x1 value of 101.6; an arrow labeled read only x1 leads to a histogram of those same 120 x1 coordinates with the identical mean marked

Open full-size figure

Return to contents

Slide 8

From observed values to a Gaussian model

A complete symmetric Gaussian bell curve from three standard deviations below its mean to three above, with the mean marked, the interval of one standard deviation shaded, and approximately 68 percent written inside it

A complete symmetric Gaussian bell curve from three standard deviations below its mean to three above, with the mean marked, the interval of one standard deviation shaded, and approximately 68 percent written inside it

Open full-size figure

Return to contents

Slide 9

The Gaussian model in two dimensions

Five hundred points sampled from a two-dimensional Gaussian, forming a tilted ellipse of points centered at the mean, with one- and two-standard-deviation elliptical contours

Five hundred points sampled from a two-dimensional Gaussian, forming a tilted ellipse of points centered at the mean, with one- and two-standard-deviation elliptical contours

Open full-size figure

Return to contents

Slide 11

Where are the points, and how do they spread?

The same points after subtracting their mean, now sitting at the origin with the axes crossing through the middle of the points

The same points after subtracting their mean, now sitting at the origin with the axes crossing through the middle of the points

Open full-size figure
  • So subtract the mean from every point: 𝐱~(i)=𝐱(i)βˆ’ΞΌ\tilde{\mathbf{x}}^{(i)} = \mathbf{x}^{(i)} - \boldsymbol{\mu}. This is centering, and it slides the points until their centre sits at the origin.
  • PCA describes variation around the mean.
  • A negative centered value means β€œdarker than this pixel’s average,” not a negative raw intensity.

Return to contents

Slide 12

Centering, in code

mean = features.mean(axis=0) # (D,) mean response across stimuli, for each coordinate centered = features - mean # (N, D) each stimulus response minus the mean response

Return to contents

Slide 13

Covariance: do the deviations go together?

The same centered pixel points and coordinate frame as before, with axes x1 minus mu1 and x2 minus mu2. Signs in each quadrant show whether multiplying the two deviations gives a positive or negative result.

The same centered pixel points and coordinate frame as before, with axes x1 minus mu1 and x2 minus mu2. Signs in each quadrant show whether multiplying the two deviations gives a positive or negative result.

Open full-size figure
  • These are the same centered responses: x~1=x1βˆ’ΞΌ1\tilde{x}_1=x_1-\mu_1 and x~2=x2βˆ’ΞΌ2\tilde{x}_2=x_2-\mu_2. Same signs give a positive product; opposite signs give a negative product.
  • Covariance is their average product:
    cov⁑(x~1,x~2)=1Nβˆ‘ix~1(i)x~2(i)\displaystyle \operatorname{cov}(\tilde{x}_1,\tilde{x}_2)=\frac{1}{N}\sum_i\tilde{x}_1^{(i)}\tilde{x}_2^{(i)}.
  • Pearson correlation is covariance measured in standard-deviation units:
    r=cov⁑(x~1,x~2)Οƒ1Οƒ2\displaystyle r=\frac{\operatorname{cov}(\tilde{x}_1,\tilde{x}_2)}{\sigma_1\sigma_2}.

Return to contents

Slide 14

The covariance matrix

Three sets of two-dimensional points side by side, each with its two-by-two covariance matrix printed underneath: a round set with equal spread in both directions, an ellipse stretched along the horizontal axis, and a tilted ellipse that runs diagonally

Three sets of two-dimensional points side by side, each with its two-by-two covariance matrix printed underneath: a round set with equal spread in both directions, an ellipse stretched along the horizontal axis, and a tilted ellipse that runs diagonally

Open full-size figure

Return to contents

Slide 17

Projection onto a unit direction

The two neighboring pixels, centered, with one point marked as the centered vector x tilde, a magenta line through the origin along the unit direction w, and the construction of the projection

The two neighboring pixels, centered, with one point marked as the centered vector x tilde, a magenta line through the origin along the unit direction w, and the construction of the projection

Open full-size figure
  • The foot of that perpendicular is 𝐩=(𝐰⊀𝐱~) 𝐰=z 𝐰\mathbf{p}=(\mathbf{w}^\top\tilde{\mathbf{x}})\,\mathbf{w} = z\,\mathbf{w}: the part of 𝐱~\tilde{\mathbf{x}} that lies along 𝐰\mathbf{w}.
  • Here K=1K=1: two coordinates become one score, zz.
  • More generally, KK retained directions replace DD coordinates with KK scores.

Return to contents

Slide 18

A direction keeps part of the spread

The two neighboring pixels again, a direction drawn at 110 degrees, and a short red segment from every point to the line showing what that direction does not capture

The two neighboring pixels again, a direction drawn at 110 degrees, and a short red segment from every point to the line showing what that direction does not capture

Open full-size figure
  • Project every point onto some direction. The red segments are what is left over: the part each point keeps to itself.
  • At this angle the direction keeps 23% of the total spread and leaves 77% over.

Return to contents

Slide 19

Turn it until the leftovers are smallest

The same points with the direction turned to 44 degrees, lengthwise through the points, and the red segments now short

The same points with the direction turned to 44 degrees, lengthwise through the points, and the red segments now short

Open full-size figure
  • Turn the direction until those segments are as short as possible: here at 44Β°, lengthwise through the points.
  • Now it keeps 90% and leaves 10%.

Return to contents

Slide 20

Spin the direction yourself

Return to contents

Slide 21

From turning the line to a calculation

The familiar centered pixel points with the best direction at 44 degrees, retaining 90 percent of the variance and leaving the smallest residual

The familiar centered pixel points with the best direction at 44 degrees, retaining 90 percent of the variance and leaving the smallest residual

Open full-size figure
  • For a unit direction 𝐰\mathbf{w}, each centered response has one score: z=𝐰⊀𝐱~z=\mathbf{w}^\top\tilde{\mathbf{x}}.
  • Across all stimuli, those scores have variance

Var⁑(z)=π°βŠ€π‚π°.\operatorname{Var}(z)=\mathbf{w}^\top\mathbf{C}\mathbf{w}.

  • The best direction is the 𝐰\mathbf{w} that makes this variance largest. The covariance matrix 𝐂\mathbf{C} lets us find it without searching every angle.

Return to contents

Slide 22

The covariance matrix changes a direction

A unit direction w and the vector Cw produced when the covariance matrix acts on it; Cw has a different length and orientation

A unit direction w and the vector Cw produced when the covariance matrix acts on it; Cw has a different length and orientation

Open full-size figure
  • The blue vector 𝐰\mathbf{w} is one possible direction through the response space.
  • Multiply it by the covariance matrix: 𝐂𝐰\mathbf{C}\mathbf{w} is the red vector.
  • Usually, the red vector has a different length and points in a different direction.

Return to contents

Slide 23

Most directions turn

The unit circle with several directions and their transformed vectors; their tips trace an ellipse and almost every vector changes orientation

The unit circle with several directions and their transformed vectors; their tips trace an ellipse and almost every vector changes orientation

Open full-size figure
  • Apply 𝐂\mathbf{C} to every direction on the unit circle. The transformed tips trace an ellipse.
  • Almost every direction comes out turned as well as rescaled.

Return to contents

Slide 24

Eigenvectors are directions that do not turn

Two special directions shown before and after multiplication by the covariance matrix; each transformed vector remains on the same dashed line

Two special directions shown before and after multiplication by the covariance matrix; each transformed vector remains on the same dashed line

Open full-size figure
  • For most vectors, 𝐂𝐰\mathbf{Cw} points in a different direction from 𝐰\mathbf{w}.
  • An eigenvector 𝐯k\mathbf{v}_k is special: 𝐂\mathbf{C} changes its length but leaves it on the same line.
  • The amount of scaling is its eigenvalue Ξ»k\lambda_k:

𝐂𝐯k=Ξ»k𝐯k.\mathbf{C}\mathbf{v}_k=\lambda_k\mathbf{v}_k.

Return to contents

Slide 25

An eigenvalue gives variance along its eigenvector

The same centered pixel points and the same 44-degree line as the best-direction slide, now with the eigenvector v1 drawn along it as an arrow and the variance along it given as 90 percent of the total

The same centered pixel points and the same 44-degree line as the best-direction slide, now with the eigenvector v1 drawn along it as an arrow and the variance along it given as 90 percent of the total

Open full-size figure
  • Put 𝐰=𝐯k\mathbf{w}=\mathbf{v}_k into the variance formula, then replace 𝐂𝐯k\mathbf{C}\mathbf{v}_k by Ξ»k𝐯k\lambda_k\mathbf{v}_k:

Var⁑(zk)=𝐯kβŠ€π‚π―k=𝐯k⊀(Ξ»k𝐯k)=Ξ»k 𝐯k⊀𝐯k=Ξ»k,\operatorname{Var}(z_k)=\mathbf{v}_k^\top\mathbf{C}\mathbf{v}_k=\mathbf{v}_k^\top(\lambda_k\mathbf{v}_k)=\lambda_k\,\mathbf{v}_k^\top\mathbf{v}_k=\lambda_k,

since 𝐯k\mathbf{v}_k has length 1.

  • The largest eigenvalue identifies the direction with the most variance. Here it is the same 44Β° direction found by turning the line.
  • Ranked by eigenvalue, these directions are the principal components. The first is PC1; computing them is principal component analysis (PCA).

Return to contents

Slide 26

PCA = eigendecomposition of 𝐂\mathbf{C}

Covariance (centered, NN samples): β€…β€Šπ‚=1Nβˆ‘i𝐱~(i)𝐱~(i)⊀\;\mathbf{C} = \dfrac{1}{N}\sum_{i} \tilde{\mathbf{x}}^{(i)}\tilde{\mathbf{x}}^{(i)\top}   (code divides by Nβˆ’1N-1; same eigenvectors)

Eigendecomposition (orthonormal 𝐕\textcolor{#B5396B}{\mathbf{V}}):

𝐂=π•β€‰Ξ›β€‰π•βŠ€,𝐂 𝐯k=Ξ»k 𝐯k,Ξ»1β‰₯Ξ»2β‰₯β‹―β‰₯Ξ»D\mathbf{C} = \textcolor{#B5396B}{\mathbf{V}}\,\textcolor{#3F7A6B}{\boldsymbol{\Lambda}}\,\textcolor{#B5396B}{\mathbf{V}}^\top,\qquad \mathbf{C}\,\textcolor{#B5396B}{\mathbf{v}_k} = \textcolor{#3F7A6B}{\lambda_k}\,\textcolor{#B5396B}{\mathbf{v}_k},\qquad \textcolor{#3F7A6B}{\lambda_1}\ge\textcolor{#3F7A6B}{\lambda_2}\ge\cdots\ge\textcolor{#3F7A6B}{\lambda_D}

Return to contents

Slide 27

From a line to a plane

A three-dimensional set of points near a plane spanned by two principal directions; one point and its projection on the plane

A three-dimensional set of points near a plane spanned by two principal directions; one point and its projection on the plane

Open full-size figure
  • In this picture, D=3D=3 measured coordinates and we retain K=2K=2 perpendicular directions.
  • The two directions define a plane; each centered point gets two scores, (z1,z2)(z_1,z_2).
  • Keeping all DD directions changes coordinates without losing information.
  • Keeping K<DK<D projects each point onto a lower-dimensional flat (linear) subspace: a line, a plane, or its analogue in more dimensions.

Return to contents

Slide 28

The new coordinates are the PC scores

The same three-dimensional centered points and retained plane; the highlighted point projects to a location with coordinates z1 and z2 in the principal-component basis

The same three-dimensional centered points and retained plane; the highlighted point projects to a location with coordinates z1 and z2 in the principal-component basis

Open full-size figure

For one stimulus, take its dot product with each retained direction:

zk=𝐯k⊀(π±βˆ’ΞΌ).z_k=\mathbf{v}_k^\top(\mathbf{x}-\boldsymbol{\mu}).

  • 𝐯k\mathbf{v}_k gives weights on the original measurements.
  • The score zkz_k is this stimulus's coordinate along that direction.
  • Keeping KK directions gives the point KK scores.

Return to contents

Slide 29

The same points, in the new coordinates

A three-dimensional set of points lying close to a plane spanned by v1 and v2, with axes x-tilde 1, 2 and 3; one point off the plane and its projection

A three-dimensional set of points lying close to a plane spanned by v1 and v2, with axes x-tilde 1, 2 and 3; one point off the plane and its projection

Open full-size figure
The same 100 points drawn in two dimensions at their scores z1 and z2 along v1 and v2; the projected point is marked in green

The same 100 points drawn in two dimensions at their scores z1 and z2 along v1 and v2; the projected point is marked in green

Open full-size figure

Return to contents

Slide 30

Reconstruction: the way back

The raw two-pixel points with their mean, one original point, and its reconstruction formed by adding the score-weighted first principal direction to the mean; the residual joins the reconstruction to the original

The raw two-pixel points with their mean, one original point, and its reconstruction formed by adding the score-weighted first principal direction to the mean; the residual joins the reconstruction to the original

Open full-size figure
  • Return to the original measurement coordinates and start from the mean.
  • Add each retained direction, weighted by the image's score:

𝐱^=ΞΌ+βˆ‘k=1Kzk𝐯k.\hat{\mathbf{x}}=\boldsymbol{\mu}+\sum_{k=1}^{K}z_k\mathbf{v}_k.

  • The residual is the difference between the original and reconstructed measurements.

Return to contents

Slide 31

What do we lose when we keep fewer components?

Top: one face reconstructed with K equal to 1, 5, 20, 50, 150 and 400 components, from a blurry near-average to the exact original. Bottom: the residual fraction for this face at each K as dots, compared with the eigenvalue tail computed from the whole dataset

Top: one face reconstructed with K equal to 1, 5, 20, 50, 150 and 400 components, from a blurry near-average to the exact original. Bottom: the residual fraction for this face at each K as dots, compared with the eigenvalue tail computed from the whole dataset

Open full-size figure

Return to contents

Slide 32

How many components? The scree plot

Scree plot for the 400 Olivetti faces: the variance fraction carried by each principal component, and the cumulative fraction

Scree plot for the 400 Olivetti faces: the variance fraction carried by each principal component, and the cumulative fraction

Open full-size figure

Return to contents

Slide 33

PCA in scikit-learn

from sklearn.decomposition import PCA pca = PCA(n_components=K, svd_solver="full") Z = pca.fit_transform(X) # X: (N, D) -> Z: (N, K) fraction = pca.explained_variance_ratio_ directions = pca.components_ # one principal direction per row

Return to contents

Slide 34

Can these directions help explain a neuron?

Three example IT cells from Bao and colleagues, each shown as a 24-view by 51-object response matrix with four views of the cell's preferred object picked out and shown underneath: an aeroplane, a swan and a basket

Three example IT cells from Bao and colleagues, each shown as a 24-view by 51-object response matrix with four views of the cell's preferred object picked out and shown underneath: an aeroplane, a swan and a basket

Open full-size figure
Bao et al. Nature 2020 Β· Fig. 3b, three of the example cells

Return to contents

Slide 35

Build an object space from network responses

Bao et al.'s schematic of the AlexNet object space: stimulus images arranged in the plane of the first two PCs, in four colored groups

Bao et al.'s schematic of the AlexNet object space: stimulus images arranged in the plane of the first two PCs, in four colored groups

Open full-size figure
  • Present the 1,224 images to AlexNet and collect responses from fc6, a late network layer.
  • Run PCA on those response vectors. In the paper, 50 components account for 85% of the variance.
  • The picture is a schematic of the first two PCs, arranged for legibility rather than a raw scatter plot.
  • β€œStubby–spiky” and β€œanimate–inanimate” are proposed interpretations of those directions. They need testing.

Return to contents

Slide 36

Test whether a direction predicts a response

Schematic simulated example: projecting points in model space onto a direction gives a scalar related to a simulated neural response

Schematic simulated example: projecting points in model space onto a direction gives a scalar related to a simulated neural response

Open full-size figure

Return to contents

Slide 37

Neural preferences occupy different regions

Grey stimulus points in the PC1-PC2 plane with four colored groups of points β€” the preferred images of four IT patch systems β€” landing one per quadrant

Grey stimulus points in the PC1-PC2 plane with four colored groups of points β€” the preferred images of four IT patch systems β€” landing one per quadrant

Open full-size figure
  • Grey circles are stimuli in the network-derived PC plane.
  • Colored points are images that strongly activate four IT systems. NML is the network Bao et al. found in IT's "no man's land", cortex that no earlier localizer had assigned a category; the numbers in parentheses count recorded neurons.
  • Different IT systems prefer different regions of the model space. A separate analysis tests whether a neuron's response varies with its projection along a direction.

Return to contents

Slide 38

Use the space to make a new prediction

Published coronal fMRI maps from two monkeys, showing the four networks at posterior, middle and anterior IT; magenta marks the predicted stubby-object network

Published coronal fMRI maps from two monkeys, showing the four networks at posterior, middle and anterior IT; magenta marks the predicted stubby-object network

Open full-size figure

Return to contents

Slide 39

The measurement scale changes PCA

The same correlated points twice. Top, in raw units with feature 1 in millimetres and feature 2 in metres, PC1 equals 1.00 times feature 1 plus 0.00 times feature 2 and carries 100 percent of the variance. Bottom, after z-scoring, PC1 is 0.71 times each feature and carries 89 percent

The same correlated points twice. Top, in raw units with feature 1 in millimetres and feature 2 in metres, PC1 equals 1.00 times feature 1 plus 0.00 times feature 2 and carries 100 percent of the variance. Bottom, after z-scoring, PC1 is 0.71 times each feature and carries 89 percent

Open full-size figure
  • PCA maximizes variance, so a feature measured in millimetres dominates one measured in metres: in raw units PC1 is 1.00 x1+0.00 x21.00\,x_1 + 0.00\,x_2 and the second feature is invisible.
  • After z-scoring each feature (subtract its mean, divide by its standard deviation), PC1 is 0.71 x1+0.71 x20.71\,x_1 + 0.71\,x_2 and recovers the diagonal structure the two features share.
  • Standardize when the units are not comparable. When they are comparable and the variances differ for real reasons, standardizing changes the relative weighting of features; it is a modelling decision, not a default.

Return to contents

Slide 40

Most variance need not mean most useful

Left: the 120 photographs laid out along the leading direction of their pixel coordinates, running from dark pictures to bright ones with animals of different categories interspersed. Right: position along that direction plotted against the mean pixel value of each photograph, an almost perfect straight line, r equals 0.99

Left: the 120 photographs laid out along the leading direction of their pixel coordinates, running from dark pictures to bright ones with animals of different categories interspersed. Right: position along that direction plotted against the mean pixel value of each photograph, an almost perfect straight line, r equals 0.99

Open full-size figure

Return to contents

Slide 41

Which direction would you keep?

Toy classes with a high-variance horizontal direction and low-variance vertical separation; projection panels ask for a prediction

Toy classes with a high-variance horizontal direction and low-variance vertical separation; projection panels ask for a prediction

Open full-size figure

Return to contents

Slide 42

Low variance can still separate the groups

Toy classes with a high-variance horizontal direction and low-variance vertical separation; PC1 scores overlap while PC2 scores separate the classes

Toy classes with a high-variance horizontal direction and low-variance vertical separation; PC1 scores overlap while PC2 scores separate the classes

Open full-size figure

Return to contents

Slide 43

What PCA provides

The two neighboring pixels again, with the direction that retains about 90 percent of their variance

The two neighboring pixels again, with the direction that retains about 90 percent of their variance

Open full-size figure
  • PCA replaces many measurements with ordered scores and tells us how much variance each direction retains.
  • Those scores support reconstruction and can connect a model representation to neural responses. In the Bao study, that connection helped motivate a new experiment.
  • PCA describes variation. A separate test establishes what the variation is useful for.

Return to contents

Slide 44

What if the variation follows a curve?

Rotated, shifted and resized versions of an image projected into two PCA directions; rotation traces a closed curve

Rotated, shifted and resized versions of an image projected into two PCA directions; rotation traces a closed curve

Open full-size figure

Return to contents

Slide 45

Further reading for the research trail

Handout: Statistics and the mathematics of PCA β€” variance, covariance, eigenvectors, projection and reconstruction, worked by hand in this lecture’s notation.

Trail entry: pick any paper cited in a lecture so far, read it, and post four to six sentences on Ed under Research trail β€” what it did, and what surprised you or what question it left you with. Reply once to a classmate. Next deadline Fri 9/25, best three of four entries count.

Return to contents

Slide 46

Minute paper

Submit on Canvas β†’ Minute papers β†’ Minute Paper 4.

Write three brief points in your own words:

  1. Why do we center the data before finding the principal directions?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

Return to contents

Slide 47

Appendix: examples and implementation details

Additional examples, the PCA/SVD implementation, and reference material for the assignments.

Return to contents

Slide 48

One direction was not enough

Left: the same 120 photographs placed on a single direction, coloured by animal category, with the categories piled on top of one another. Right: the distortion of the ratings plotted against the number of directions kept, falling steeply from one to two and then flattening out

Left: the same 120 photographs placed on a single direction, coloured by animal category, with the categories piled on top of one another. Right: the distortion of the ratings plotted against the number of directions kept, falling steeply from one to two and then flattening out

Open full-size figure

Return to contents

Slide 49

How far does the dependence reach?

Three panels. Left: a pixel plotted against its immediate neighbour across the 120 photographs, the points hugging the diagonal at r equals 0.81. Middle: the same pixel against the one twenty places away, points with no visible trend at r equals 0.08. Right: the correlation plotted against separation, falling steeply over the first few pixels and flattening near zero

Three panels. Left: a pixel plotted against its immediate neighbour across the 120 photographs, the points hugging the diagonal at r equals 0.81. Middle: the same pixel against the one twenty places away, points with no visible trend at r equals 0.08. Right: the correlation plotted against separation, falling steeply over the first few pixels and flattening near zero

Open full-size figure

Return to contents

Slide 50

PCA in code: two routes to the same answer

centered = features - features.mean(axis=0) # Route 1: eigendecomposition of sample covariance C = np.cov(centered.T) lam, V = np.linalg.eigh(C) # ascending order lam, V = lam[::-1], V[:, ::-1] # descending order # Route 2: SVD, without forming C _, s, Vt = np.linalg.svd(centered, full_matrices=False) lam2 = s**2 / (len(centered) - 1) V2 = Vt.T # directions as columns Z = centered @ V2[:, :K] # (N, K) scores

Return to contents

Slide 51

Eigenfaces

PCA on 400 faces, and a classic approach to face recognition.

Return to contents

Slide 52

PCA as compression: eigenfaces

Grid of the mean face and leading eigenfaces, with opposing signs shown in red and blue

Grid of the mean face and leading eigenfaces, with opposing signs shown in red and blue

Open full-size figure

Stack face images as vectors and run PCA β†’ the leading eigenvectors are eigenfaces: ghostly, face-like basis images, one direction in pixel space each.

Turk & Pentland J. Cogn. Neurosci. 1991 Β· precursor: Sirovich & Kirby J. Opt. Soc. Am. A 1987 Β· Olivetti faces (AT&T Laboratories Cambridge)

Return to contents

Slide 53

Face space

Twelve face photographs placed on a plane at their coordinates along the first two eigenfaces, z1 on the horizontal axis and z2 on the vertical

Twelve face photographs placed on a plane at their coordinates along the first two eigenfaces, z1 on the horizontal axis and z2 on the vertical

Open full-size figure
  • 400 face images (64Γ—6464\times64 pixels, so D=4096D = 4096), PCA on the centered pixels, and twelve of the faces drawn at their coordinates (z1,z2)(z_1, z_2) on the first two eigenvectors.
  • Two numbers per face instead of 4,096. They carry 24% and 14% of the variance: a coarse, lossy summary, but already one in which lighting and head pose vary smoothly across the plane.
  • This is a face space: every face is a point, and nearby points are faces that look alike to the two leading eigenfaces. The preceding slide shows those eigenfaces as images.

Return to contents

Slide 54

A face = mean + a few eigenfaces

A face reconstructed using 1, 5, 25, 100 and 400 components, beside the original image

A face reconstructed using 1, 5, 25, 100 and 400 components, beside the original image

Open full-size figure

Return to contents

Slide 55

Mix your own face

Return to contents

Slide 56

Two views, one answer

Return to contents

Slide 57

PCA as decorrelation & whitening

𝐂=π•Ξ›π•βŠ€β€…β€Šβ€…β€Šβ‡’β€…β€Šβ€…β€Šπ³=Ξ›βˆ’1/2π•βŠ€(π±βˆ’ΞΌ),Cov(𝐳)=𝐈\mathbf{C} = \mathbf{V}\boldsymbol{\Lambda}\mathbf{V}^\top \;\;\Rightarrow\;\; \mathbf{z} = \boldsymbol{\Lambda}^{-1/2}\mathbf{V}^\top(\mathbf{x}-\boldsymbol{\mu}), \quad \mathrm{Cov}(\mathbf{z}) = \mathbf{I}

Return to contents

Slide 58

What do the components look like?

Principal components of synthetic image patterns, shown as red and blue signed weights: structured patterns dominate early components and noise appears in later ones

Principal components of synthetic image patterns, shown as red and blue signed weights: structured patterns dominate early components and noise appears in later ones

Open full-size figure
Generated after Murphy Probabilistic Machine Learning 2022 ch. 20 (MIT License)

Return to contents

Slide 59

The rainbow cortex is three numbers per voxel

Panels from Huth et al. showing word-related features above cortical surfaces, where every voxel is colored by its first three PCA coordinates (red for PC1, green for PC2, blue for PC3), producing a smooth rainbow map across cortex

Panels from Huth et al. showing word-related features above cortical surfaces, where every voxel is colored by its first three PCA coordinates (red for PC1, green for PC2, blue for PC3), producing a smooth rainbow map across cortex

Open full-size figure
  • Over two hours of narrated stories, one regression per voxel: every voxel gets a vector of weights on word-related features, and cortex is one more NΓ—DN\times D matrix.
  • PCA on it, keep three components, color each voxel by its coordinates: PC1 β†’ red, PC2 β†’ green, PC3 β†’ blue. The palette is the projection 𝐳\mathbf{z}.

Return to contents

Slide 60

PCA finds flat subspaces

Top: two interleaved spiral arms, one blue and one green, with the PC1 line drawn through them. Bottom: the histogram of the coordinates along PC1, where the two arms overlap almost completely

Top: two interleaved spiral arms, one blue and one green, with the PC1 line drawn through them. Bottom: the histogram of the coordinates along PC1, where the two arms overlap almost completely

Open full-size figure
  • If the data lie on a curved manifold, a smooth curved surface (a spiral, a Swiss roll, a torus), a low-dimensional linear projection can fold distinct regions onto each other.
  • Two points far apart along the manifold can land on top of each other after projection: the two arms are perfectly distinct in the plane and inseparable along PC1.
  • t-SNE and UMAP construct nonlinear maps emphasizing local relationships. Their displays can distort global structure and need their own checks.

Return to contents