These are the same 120 dots. For the moment, read each dot's horizontal x1 coordinate and set x2 aside.
The histogram is the distribution of those 120 values. Their mean is 101.6; their spread tells us how much this measurement varies across photographs.
From observed values to a Gaussian model
The previous histogram showed observed pixel values. This is an idealized reference model, not an assumption today's method requires.
A Gaussian is smooth and bell-shaped. Its meanμ gives the center and its standard deviationσ gives the spread; about 68% lies between μ−σ and μ+σ.
The Gaussian model in two dimensions
Each simulated point contains two measurements, (x1,x2). These are draws from the model, not the 120 photographs.
The mean becomes a mean vectorμ, the center of the points; the covariance matrix describes their width, elongation and tilt.
PCA will find the directions along which the ellipse is longest and shortest.
Where are the points, and how do they spread?
For each pixel j, average over the N=120 images:
Mean:μj=N1i∑xj(i) The center of the points is μ=(μ1,μ2).
Variance:σj2=N1i∑(xj(i)−μj)2 Spread around the mean along that coordinate.
Standard deviation:σj=σj2 Spread in pixel-intensity units.
Where are the points, and how do they spread?
So subtract the mean from every point: x~(i)=x(i)−μ. This is centering, and it slides the points until their centre sits at the origin.
PCA describes variation around the mean.
A negative centered value means “darker than this pixel’s average,” not a negative raw intensity.
Centering, in code
mean = features.mean(axis=0) # (D,) mean response across stimuli, for each coordinate
centered = features - mean # (N, D) each stimulus response minus the mean response
features is the response matrix from the first lectures: one row per stimulus (images here), one column per coordinate.
The mean is a vector with one entry per coordinate. Subtracting it from every row leaves each column summing to zero.
Keep the mean. PCA works on differences from it, and reconstruction adds it back.
Covariance: do the deviations go together?
These are the same centered responses: x~1=x1−μ1 and x~2=x2−μ2. Same signs give a positive product; opposite signs give a negative product.
Covariance is their average product: cov(x~1,x~2)=N1i∑x~1(i)x~2(i).
Pearson correlation is covariance measured in standard-deviation units: r=σ1σ2cov(x~1,x~2).
The covariance matrix
In C, the diagonal contains the variances. The same covariance appears on both sides of the diagonal because cov(x1,x2)=cov(x2,x1).
Geometrically C summarizes the points’ spread and orientation: round, an ellipse, or an ellipse tilted because the two vary together.
The ellipse's long direction is its direction of greatest variance.
Projection onto a unit direction
Which direction gives the best one-number summary?
A unit directionw is any direction, written as a vector of length 1.
The projection length of a centered point is z=w⊤x~=∥x~∥cosϕ: one number, its coordinate along w.
Projection onto a unit direction
Drop a perpendicular from x~ onto the line of w.
The dashed segment is what the direction does not capture: the part of the point left over once its position along w is known.
Projection onto a unit direction
The foot of that perpendicular is p=(w⊤x~)w=zw: the part of x~ that lies along w.
Here K=1: two coordinates become one score, z.
More generally, K retained directions replace D coordinates with K scores.
A direction keeps part of the spread
Project every point onto some direction. The red segments are what is left over: the part each point keeps to itself.
At this angle the direction keeps 23% of the total spread and leaves 77% over.
Turn it until the leftovers are smallest
Turn the direction until those segments are as short as possible: here at 44°, lengthwise through the points.
Now it keeps 90% and leaves 10%.
Spin the direction yourself
Before turning the direction, predict which direction leaves the smallest error.
Projected variance and squared reconstruction error add to a fixed total. Maximizing one minimizes the other.
From turning the line to a calculation
For a unit direction w, each centered response has one score: z=w⊤x~.
Across all stimuli, those scores have variance
Var(z)=w⊤Cw.
The best direction is the w that makes this variance largest. The covariance matrix C lets us find it without searching every angle.
The covariance matrix changes a direction
The blue vector w is one possible direction through the response space.
Multiply it by the covariance matrix: Cw is the red vector.
Usually, the red vector has a different length and points in a different direction.
Most directions turn
Apply C to every direction on the unit circle. The transformed tips trace an ellipse.
Almost every direction comes out turned as well as rescaled.
Eigenvectors are directions that do not turn
For most vectors, Cw points in a different direction from w.
An eigenvectorvk is special: C changes its length but leaves it on the same line.
The amount of scaling is its eigenvalueλk:
Cvk=λkvk.
An eigenvalue gives variance along its eigenvector
Put w=vk into the variance formula, then replace Cvk by λkvk:
Var(zk)=vk⊤Cvk=vk⊤(λkvk)=λkvk⊤vk=λk,
since vk has length 1.
The largest eigenvalue identifies the direction with the most variance. Here it is the same 44° direction found by turning the line.
Ranked by eigenvalue, these directions are the principal components. The first is PC1; computing them is principal component analysis (PCA).
PCA = eigendecomposition of C
Covariance (centered, N samples): C=N1∑ix~(i)x~(i)⊤(code divides by N−1; same eigenvectors)
Eigendecomposition (orthonormal V):
C=VΛV⊤,Cvk=λkvk,λ1≥λ2≥⋯≥λD
Principal directions = eigenvectorsvk of C.
Variance along vk = λk.
PC1=v1, the eigenvector with the largest eigenvalue.
From a line to a plane
In this picture, D=3 measured coordinates and we retain K=2perpendicular directions.
The two directions define a plane; each centered point gets two scores, (z1,z2).
Keeping all D directions changes coordinates without losing information.
Keeping K<D projects each point onto a lower-dimensional flat (linear) subspace: a line, a plane, or its analogue in more dimensions.
The new coordinates are the PC scores
For one stimulus, take its dot product with each retained direction:
zk=vk⊤(x−μ).
vk gives weights on the original measurements.
The scorezk is this stimulus's coordinate along that direction.
Keeping K directions gives the point K scores.
The same points, in the new coordinates
Left: points in the original measured coordinates x~1,x~2,x~3. Right: the same points at their scores(z1,z2): the PCA embedding.
The green point is the same stimulus in both. Its distance from the plane, the dashed red residual, is what the 2-D map leaves out.
Reconstruction: the way back
Return to the original measurement coordinates and start from the mean.
Add each retained direction, weighted by the image's score:
x^=μ+k=1∑Kzkvk.
The residual is the difference between the original and reconstructed measurements.
What do we lose when we keep fewer components?
PCA on the pixels of 400 faces: keeping more components adds detail and reduces reconstruction error.
A residual fraction of 0 means exact reconstruction; larger values mean more of that face was left out.
Dots: this face. Curve: the discarded variance across all 400 faces, which need not equal one face's error.
The gold ring follows one image from its position in model space to its measured response.
Choose a direction using training images, then ask whether each image's projection predicts the cell's response on held-out images, which were not used to choose it.
The preferred direction can combine several PCs; it need not coincide with PC1 or PC2.
Neural preferences occupy different regions
Grey circles are stimuli in the network-derived PC plane.
Colored points are images that strongly activate four IT systems. NML is the network Bao et al. found in IT's "no man's land", cortex that no earlier localizer had assigned a category; the numbers in parentheses count recorded neurons.
Different IT systems prefer different regions of the model space. A separate analysis tests whether a neuron's response varies with its projection along a direction.
PCA maximizes variance, so a feature measured in millimetres dominates one measured in metres: in raw units PC1 is 1.00x1+0.00x2 and the second feature is invisible.
After z-scoring each feature (subtract its mean, divide by its standard deviation), PC1 is 0.71x1+0.71x2 and recovers the diagonal structure the two features share.
Standardize when the units are not comparable. When they are comparable and the variances differ for real reasons, standardizing changes the relative weighting of features; it is a modelling decision, not a default.
Most variance need not mean most useful
The largest direction in the 120 photographs carries 29% of the variation and is strongly related to overall brightness: ordered along it they run dark to bright, different categories interspersed.
Would this direction help distinguish the animals? This plot alone cannot answer.
PCA ranks directions by variance; the task determines what is useful.
Which direction would you keep?
The groups vary mainly from left to right. Their labels differ from top to bottom.
Which direction would best reconstruct the points?
Which direction would best distinguish the groups?
Low variance can still separate the groups
PC1 retains 97.8% of variance, but both groups have the same PC1 coordinates.
PC2 retains 2.2%, yet a threshold at zero separates the groups.
Variance measures spread. Task relevance depends on what we want to distinguish.
What PCA provides
PCA replaces many measurements with ordered scores and tells us how much variance each direction retains.
Those scores support reconstruction and can connect a model representation to neural responses. In the Bao study, that connection helped motivate a new experiment.
PCA describes variation. A separate test establishes what the variation is useful for.
What if the variation follows a curve?
Changing one stimulus parameter can trace a curved path in a large response space.
Even one varying parameter can require two linear coordinates when the path bends back on itself.
PCA approximates data with a flat subspace.
Next: nonlinear maps, and how to check what their geometry preserves.
Further reading for the research trail
Handout:Statistics and the mathematics of PCA — variance, covariance, eigenvectors, projection and reconstruction, worked by hand in this lecture’s notation.
Bao et al. (2020): object-space directions and IT responses. Nature
Chang & Tsao (2017): face-space directions and individual neurons. Cell
Stringer et al. (2019): a structured, high-dimensional neural spectrum. Nature
Trail entry: pick any paper cited in a lecture so far, read it, and post four to six sentences on Ed under Research trail — what it did, and what surprised you or what question it left you with. Reply once to a classmate. Next deadline Fri 9/25, best three of four entries count.
Minute paper
Submit on Canvas → Minute papers → Minute Paper 4.
Write three brief points in your own words:
Why do we center the data before finding the principal directions?
Something you do not yet understand, or a question still open
Another idea you found interesting, and why it matters for brains, behavior or AI
Credit for a thoughtful attempt, not for being correct.
Appendix: examples and implementation details
Additional examples, the PCA/SVD implementation, and reference material for the assignments.
One direction was not enough
Fit the same ratings on one direction and the categories pile up: the distances in the fit no longer match the judged dissimilarities.
A second direction removes most of that distortion. A third and a fourth remove very little. Two directions greatly improve the fit, but some distortion remains.
How far does the dependence reach?
The same pixel-pair computation, at larger separations: one pixel against another, across the same 120 photographs.
Next door, the two pixel values strongly covary. Twenty pixels apart, they barely covary.
Neighbouring measurements are strongly dependent, and the dependence is local.
PCA in code: two routes to the same answer
centered = features - features.mean(axis=0)
# Route 1: eigendecomposition of sample covariance
C = np.cov(centered.T)
lam, V = np.linalg.eigh(C) # ascending order
lam, V = lam[::-1], V[:, ::-1] # descending order# Route 2: SVD, without forming C
_, s, Vt = np.linalg.svd(centered, full_matrices=False)
lam2 = s**2 / (len(centered) - 1)
V2 = Vt.T # directions as columns
Z = centered @ V2[:, :K] # (N, K) scores
Use eigh, not eig: C is symmetric, so the eigenvalues are real and eigh is faster and stabler. It returns them ascending, so the reversal is not optional.
Prefer route 2. Forming C costs D×D memory; for a 256×256 color image D≈196,608 and C would be ~309 GB in float64. SVD never builds it. Sign is arbitrary: vk and −vk are both valid; Assignment 2's helper fixes it so the largest entry is positive.
Eigenfaces
PCA on 400 faces, and a classic approach to face recognition.
PCA as compression: eigenfaces
Stack face images as vectors and run PCA → the leading eigenvectors are eigenfaces: ghostly, face-like basis images, one direction in pixel space each.
400 face images (64×64 pixels, so D=4096), PCA on the centered pixels, and twelve of the faces drawn at their coordinates (z1,z2) on the first two eigenvectors.
Two numbers per face instead of 4,096. They carry 24% and 14% of the variance: a coarse, lossy summary, but already one in which lighting and head pose vary smoothly across the plane.
This is a face space: every face is a point, and nearby points are faces that look alike to the two leading eigenfaces. The preceding slide shows those eigenfaces as images.
A face = mean + a few eigenfaces
These faces can be approximated by the mean face plus a weighted sum of eigenfaces. More components add finer detail.
Turk & Pentland used the coefficients for recognition: a new face is matched to the stored face with the nearest coefficients.
Each slider sets one coefficient in mean + weighted eigenfaces.
Set a coefficient to zero to remove that direction’s contribution.
Two views, one answer
Maximize the variance of the projections: ∥v∥=1maxv⊤Cv
Minimize the squared reconstruction error: ∥v∥=1minN1i∑∥x~(i)−(v⊤x~(i))v∥2
These are the same problem: total variance is fixed, so whatever the projection captures, the residual loses. Both are solved by the top eigenvector of C.
For K retained PCs, the average squared fitting-data reconstruction error is ∑k>Kλk (using covariance divided by N).
Connection to MDS: classical MDS on a Euclidean dissimilarity matrix returns the PCA coordinates up to an orthogonal transformation of the centered data.
PCA as decorrelation & whitening
Project onto the eigenbasis and the new coordinates are uncorrelated: the covariance becomes diagonal. For positive eigenvalues, rescale each retained direction by 1/λi and you get whitening: zero correlation, unit variance.
C=VΛV⊤⇒z=Λ−1/2V⊤(x−μ),Cov(z)=I
Why analysts reach for it:
remove redundancy between measured dimensions before modeling
decorrelate inputs, without assuming the resulting features are independent
compress by dropping directions; whether this denoises depends on what they contain
What do the components look like?
Each eigenvector is a vector in the original space, so each PC unrolls into a picture. Here the inputs are synthetic patterns: early components recover structured variation; later ones contain little variance.
The components can be read as weights on the original variables, then tested for what they capture.
Over two hours of narrated stories, one regression per voxel: every voxel gets a vector of weights on word-related features, and cortex is one more N×D matrix.
PCA on it, keep three components, color each voxel by its coordinates: PC1 → red, PC2 → green, PC3 → blue. The palette is the projection z.
If the data lie on a curved manifold, a smooth curved surface (a spiral, a Swiss roll, a torus), a low-dimensional linear projection can fold distinct regions onto each other.
Two points far apart along the manifold can land on top of each other after projection: the two arms are perfectly distinct in the plane and inseparable along PC1.
t-SNE and UMAP construct nonlinear maps emphasizing local relationships. Their displays can distort global structure and need their own checks.
- Same 120 photographs and ratings as Assignment 1
- Non-metric MDS fits the rank order of the dissimilarities, not their values
- Orientation is arbitrary, and different initializations can reach different solutions: your notebook map may differ from this one by more than a rotation
- The category blocks: dissimilarities are smaller within a category than between categories
- PCA works on the response matrix itself (stimuli × pixels, features or neurons), not on an RDM computed from it
- Question: is the variation across many measured dimensions concentrated along a few directions? Every variable is still recorded; PCA changes the description, not the measurement
- Because the response matrix names the measured variables, each PCA direction is a set of weights on pixels, features or units
- Behavioral dissimilarities can come directly from judgments, with no response matrix behind them
- Embedding dimension: the number of principal components needed to account for the variance at a chosen threshold, not a count of true latent causes. The nonlinear-dimensionality lecture names it formally, next to the ambient and intrinsic dimensions
- PCA chooses directions by variance; reconstruction error measures the cost of keeping fewer of them
- Whether the retained directions help with a particular task is a separate question
- Intensities at one fixed pixel pair (row 33, columns 32 and 33, counting from 1) across the 120 photographs, resized to 64 × 64
- Raw values, 0 to 255; correlation r = 0.81
- Leading direction ≈ 0.71 x₁ + 0.70 x₂: a local brightness measurement, 90% of the variance
- Perpendicular direction ≈ 0.70 x₁ − 0.71 x₂ (overall sign arbitrary): compares the two pixels, 10%
- The weights come out close to average and difference because the two pixels have similar variances and strong positive covariance; "average" and "difference" describe these fitted directions and are not the definition of PC1 and PC2
- A distribution records how often each value occurs; here, the x₁ values of the same 120 photographs, mean 101.6
- Across photographs, a pixel intensity is a random variable: a quantity whose value changes from one photograph to the next
- Each bar is the percentage of photographs in that intensity bin; the bars sum to 100% (a percentage per bin, not a probability density)
- Center and spread are defined on one variable first, then carried back to the two pixels together
- The histogram was an empirical distribution: the values observed in 120 photographs. A Gaussian is a mathematical model of all the values one could observe
- Drawn in standard units from −3σ to +3σ; the tails continue in both directions
- PCA does not require Gaussian data
- Synthetic: 500 draws from a two-dimensional Gaussian
- For a Gaussian, the mean vector and covariance matrix specify the whole distribution; for other data they summarize only center and spread
- The ellipses mark constant standardized distance from the mean; in two dimensions they do not enclose 68% and 95%
- Each image i gives two raw intensities, x₁⁽ⁱ⁾ and x₂⁽ⁱ⁾; j indexes the pixel
- Averaging over images gives one mean per pixel; together they form the mean vector, ≈ (102, 104)
- Variance and standard deviation describe spread along each coordinate separately; covariance describes how the two vary together
- Dividing by N summarizes this dataset; the sample estimator divides by N − 1
- Centering moves every point by the same vector; distances between points are unchanged
- Without centering, the decomposition measures spread around the origin, and the mean can dominate the leading direction
- Keep the mean: reconstruction adds it back
- `features` is N × D, the same table of responses as in the first lectures; the mean has length D, so the subtraction runs column by column
- The helper in Assignment 2 starts with these two lines; scikit-learn's PCA centers the data the same way internally
- A tilde marks a centered response: x̃ⱼ⁽ⁱ⁾ = xⱼ⁽ⁱ⁾ − μⱼ, both coordinates measured across the same images
- Positive covariance: same-signed deviations outweigh opposite-signed ones
- Pearson correlation (previous lecture) divides covariance by the two standard deviations; for z-scored coordinates the two are equal. Covariance carries units; correlation has none
- A coordinate paired with itself gives its variance (the diagonal of the covariance matrix); swapping the order gives the same covariance, so the matrix is symmetric
- Neither implies causation, and zero covariance does not imply independence
- One variance per feature on the diagonal, one covariance per pair off it
- Simulated Gaussian data with set covariances
- Unequal variances alone give PCA a preferred direction, even with zero covariance; for the round set of points every direction is equally good
- The ellipse captures second-order spread only, not all structure in a distribution
- The tilde marks a centered response vector
- z is a signed coordinate: negative when the point lies on the opposite side of the origin from w
- The arrow for w is enlarged for visibility; it is not on the pixel-intensity scale
- The dashed segment is the residual: the part of x̃ that one score cannot reconstruct
- PCA picks the direction that minimizes the sum of squared residual lengths over the data
- For the marked point: z ≈ 200 (centered pixel units), residual length ≈ 48
- z is a number, the coordinate; z w is a vector, the reconstruction along w
- D: number of measurements per stimulus; K: number of retained directions
- Same points, with each one's residual drawn in red
- A direction across the points leaves long residuals: a poor summary
- Best direction ≈ 44.4°, keeping ≈ 90.4% of the total variance
- Minimizing the summed squared residuals and maximizing the retained variance give the same direction
- The demo opens on the same two pixels across the 120 photographs
- Turn w: retained variance and squared reconstruction error always sum to the total variance
- The menu offers other feature pairs (half-image means; mean intensity against contrast); each is a different set of points with its own best direction
- The demo divides by N and `np.cov` by N − 1; the constant factor changes neither the variance fractions nor the directions
- z is the projection coordinate from the earlier slides
- For centered data, Var(z) is the mean of z²; substituting z = wᵀx̃ gives wᵀCw
- The demo searched over angles; the eigendecomposition solves the same maximization directly
- Blue: an input direction, not a data point. Red: the same vector after multiplication by C
- C is the pixel example's covariance, divided by a positive constant to fit on the slide; this keeps the directions and changes the eigenvalues, so no stretch factors are shown
- Most input directions change both length and orientation
- The unit circle is a set of input directions, not data
- The red ellipse shows what C does to those directions; it is not a covariance contour of the data
- Both share the principal directions; the half-lengths of this ellipse scale with λ, those of a covariance contour with √λ
- On each dashed line, the input vector and its transformed version lie on the same line: C changes only the length, by the factor λ
- Rescaling C to fit the drawing leaves the eigenvectors unchanged
- In the pixel example, the first eigenvector lies at ≈ 44°, the direction found by turning the line in the demo
- For a unit eigenvector, vᵀv = 1, so its projected variance equals its eigenvalue
- The largest eigenvalue is the maximum of wᵀCw over all unit directions (a standard linear-algebra result)
- Turning the line, minimizing reconstruction error and the eigendecomposition select the same 44° direction; variance fractions ≈ 90% and 10%
- Recipe: center, compute the covariance, order its eigenvectors by eigenvalue
- Eigenvectors give directions; eigenvalues give the variance along them
- SVD of the centered data gives the same quantities without forming C (appendix)
- Schematic, centered: D = 3, K = 2
- Orthonormal basis: mutually perpendicular unit directions; a complete basis spans the whole space
- The highlighted point lies off the plane; its projection is given by the two scores z₁ and z₂
- With the retained directions as the columns of V_K: z = V_Kᵀ(x − μ)
- For the whole dataset (rows = stimuli): Z = X_c V_K, one row per stimulus, K columns
- Weights (v_k, one entry per measurement) and scores (z_k, one per stimulus) are different quantities
- Both figures show the same 100 centered points. On the left they live in the three measured coordinates; on the right each is placed at its two scores, its coordinates along v1 and v2
- Going from left to right is the projection z = V_K transpose (x minus mu); the third coordinate, perpendicular to the plane, is dropped. Here it holds 2.1% of the squared length, so little is lost
- This 2-D plot is what a PCA figure in a paper shows. The axes are no longer measurements such as pixels or neurons but the principal directions
- The axes are back in raw pixel units, not centered ones
- Green arrow: from μ to x̂, that is z₁v₁. Dashed segment: the residual x − x̂, perpendicular to v₁
- Keeping all D directions reconstructs x exactly
- The same equation is the decoder of a linear autoencoder (future lectures)
- Residual fraction for one face: ‖x − x̂‖² / ‖x − μ‖²
- Eigenvalue tail: summed squared residuals over all 400 faces, divided by their total centered squared norm
- The tail is not the average of the per-face fractions, and not any one face's fraction
- The elbow is a heuristic, not a statistical test
- Same face matrix as the previous slide: the cumulative curve is the dataset-wide counterpart of the discarded-variance curve
- In practice K is often set by the next analysis: keep as many components as it needs
- Theme 2 asks the same question of neural population responses
- X is N × D, Z is N × K; choose K before running
- PCA uses no labels: it learns directions from X alone
- `svd_solver="full"` makes the result deterministic; Assignment 1 may use the default solver
- The sign of each direction is arbitrary: flipping a direction and its scores leaves the reconstruction unchanged, so compare absolute correlations
- Each matrix shows identity and viewpoint together
- Bright columns suggest view tolerance; cells can still prefer some views
- The 1,224 centered image responses span at most 1,223 directions (their rank); the 51 object means, at most 50. Averaging over views changes what variation PCA sees
- PCA is fitted to AlexNet responses, not to the IT recordings
- PCA gets no category labels, although AlexNet was trained with object labels
- Assignment 1 tests the proposed interpretations of these directions
- Simulated to illustrate the test; not a fit from Bao et al.
- A fitted direction links the network's responses (a population of units) to one cell's responses
- Fit the direction on some images and evaluate it on others; otherwise a flexible model fits noise
- The projection is the dot product from earlier in the lecture
- Colored dots are highly effective images for each cortical network; numbers in parentheses count recorded neurons
- The result that each neuron responds along one direction ("axis coding" in the paper) comes from a separate analysis relating each neuron's response to projections along object-space directions
- Assignment 1 covers the PCA plane and preferred directions, not every experiment in the paper
- Bao et al. (2020), https://doi.org/10.1038/s41586-020-2350-5
- Rows: monkeys M3 and M4; columns: posterior, middle and anterior IT; colors as in the object-space plot
- fMRI maps thresholded at P < .001, uncorrected
- The face, body and NML networks motivated the prediction of a fourth network for stubby, inanimate objects; fMRI located it and targeted recordings confirmed its preferences
- Bao et al. (2020), https://doi.org/10.1038/s41586-020-2350-5
- Rescaling a feature changes the covariance and can rotate the principal directions
- Same observations in both panels
- With known, nonzero scales, z-scoring is invertible: nothing is lost, but distances, PCA and regularized fits change
- The leading pixel-space score correlates with mean brightness at r = 0.99
- Brightness here mixes background, object and illumination
- Whether the direction also carries category information needs a readout (a weighted sum of the retained components, then a threshold) fit on some images and tested on held-out ones
- Constructed toy data: the two groups share their horizontal coordinates; their vertical coordinates have opposite signs
- PCA ranks the horizontal direction first, yet the label depends only on the vertical coordinate
- The small direction separates the groups perfectly because the label rule is noiseless; real readouts need held-out evaluation
- In the score panels, identical PC1 markers overlap; the vertical jitter in the PC2 panel only makes repeated scores visible
- PCA gives ordered coordinates, a reconstruction and an account of the variance
- In Bao et al., those coordinates supported a hypothesis about IT responses and led to a new experiment
- What the coordinates are useful for needs a task or neural-response test; a low-dimensional description does not by itself measure task information
- Real computations on rotated, shifted and resized versions of one photograph, all projected onto two PCs fitted to the rotations
- One varying parameter can trace a curve that needs more than one linear direction
- Nonlinear maps (next lecture) give other summaries; an attractive map need not preserve global distances
- Huth et al. (2016), semantic maps: https://doi.org/10.1038/nature17637
- Gardner et al. (2022), grid-cell population geometry: https://doi.org/10.1038/s41586-021-04268-7
- Perich (2026), interpreting dimensionality: https://doi.org/10.53053/FTEU9810
- Marks et al. (2026), PCA of hippocampal measurements: https://doi.org/10.1002/brb3.71666
- Chang & Tsao use shape and appearance coordinates, not raw-pixel eigenfaces; Stringer et al.'s high-dimensional spectrum is not evidence that only a few dimensions matter
- Canvas → Minute papers → Minute Paper 4
- Credit is for a thoughtful attempt, not for being correct
- Stress falls from ≈ 0.50 with one dimension to ≈ 0.20 with two, then more slowly
- Non-metric MDS stress compares map distances with fitted disparities that keep the rank order of the dissimilarities
- A good two-dimensional approximation, not proof that the judgments have exactly two underlying dimensions
- Vertical jitter in the one-dimensional panel only separates overlapping markers
- Same fixed pixel, same 120 photographs, same Pearson correlation, at increasing separation
- Computed across images, not within one image: this across-image covariance is what the covariance matrix holds
- Independent random pixels would give zero at every separation; photographs give 0.81 next door and 0.08 twenty pixels away
- This local dependence is why a photograph is not a random point in pixel space
- Both routes give the same nonzero variances and principal subspaces, up to numerical error
- Signs are arbitrary; with repeated eigenvalues, any orthonormal basis of the shared subspace is valid
- With N < D, reduced SVD returns N columns; the covariance route also returns a large null space
- A 196,608 × 196,608 float64 covariance takes ≈ 309 GB (288 GiB) before any workspace
- N against N − 1 changes the variances by a constant factor; scikit-learn's solver depends on its settings and version
- Eigenfaces: a face as the mean image plus weighted directions of variation in pixel space
- Turk & Pentland (1991), Eigenfaces for recognition, J. Cogn. Neurosci., https://doi.org/10.1162/jocn.1991.3.1.71
- Eigenfaces are signed weights on pixels, so each renders as grey with light and dark lobes rather than as a face
- The first few capture overall brightness, lighting direction and coarse head shape; later ones carry finer detail
- Olivetti set: 400 images of 40 people, 64 × 64 pixels
- Each of the 400 Olivetti images is a point in 4,096-dimensional pixel space; projecting onto the first two eigenvectors places it on this plane
- The first two components carry 24% and 14% of the variance: a rough summary, but not a random arrangement
- Brightness and lighting direction vary along the first direction; pose and face shape along the second
- The eigenfaces are the direction weights; the scores place each face on the plane
- The reconstruction equation as images: x̂ = μ + Σₖ zₖ vₖ
- 4,096 pixels down to about 100 coefficients, a forty-fold reduction, and the face is still recognisable
- Mix: the mean face plus or minus multiples of the first eight eigenfaces; each slider spans three standard deviations of that coefficient across the dataset
- K: one of twelve stored faces rebuilt from its first K coefficients, with the residual fraction beside the eigenvalue tail
- Plane: click a point on the PC1–PC2 plane to draw the face with exactly those two coordinates
- Classical MDS on Euclidean distances returns the PCA coordinates, up to an orthogonal transformation
- The previous lecture taught the stress-minimising route to MDS, not this eigendecomposition route; the connection is for reference, not something to derive
- Whitening applies only to directions with positive eigenvalues; a zero-variance direction cannot be divided by its standard deviation
- Small eigenvalues amplify noise; in practice, truncate or regularize
- Uncorrelated features need not be independent
- Whitening differs from z-scoring the original features
- A principal component of an image dataset is a vector in pixel space, so it can be displayed as an image
- Synthetic patterns here, not the faces
- Huth et al. (2016), Natural speech reveals the semantic maps that tile human cerebral cortex, Nature, https://doi.org/10.1038/nature17637
- Each voxel's responses to the stories were modeled from word-related features; PCA ran on the fitted weights, and three component scores became the color channels
- The colors encode a projection, not three named psychological properties; semantic labels come from examining the weights and predictions
- The two spiral arms overlap when projected onto PC1
- Nonlinear maps (next lecture) can show other relationships, but their displays also need checking; visual separation in a map is not a classifier