=========================================================================== CPSY 1291 — RECITATION 3: Matrices, curve fitting, and methods with knobs TA-led, 80 minutes. Optional. Week 3 — hold this AFTER Lecture 4 and BEFORE Assignment 1 is due (Tue 9/29). HOW TO RUN THIS DECK: every slide's presenter note gives a timing, what to say, and (where relevant) the answer to the exercise on the slide. Slides marked EXERCISE are meant to be done live — put the prompt up, let them try for 2-3 minutes, then take answers from the room before revealing. SCOPE: the four things Assignment 1 needs that Recitations 1-2 did not give them — matrix products, curve fitting and R^2, Spearman on the upper triangle, and knobs (random_state, perplexity). Do NOT re-teach PCA; R02 did that, and repeating it costs the time this session does not have. MATERIAL: handout-03-matrices-fitting-knobs.pdf covers the same ground in prose. Say so at the start so nobody takes notes. DO NOT: work any part of Assignment 1 on the board. Every example here uses different data on purpose. ============================================================================
2 min. Say the frame: Assignment 1 is due Tuesday. This session is the last scheduled help before then, and it covers only what R01 and R02 did not. Set the expectation out loud: nothing from the assignment will be worked on the board. Every example uses different data, deliberately. Point at office hours for assignment-specific questions.
2 min. Read the last line aloud; it is the through-line of the session and of Part 4 of the assignment. If the room is thin on R02 attendance, do NOT back up — point at linear-algebra-notes.pdf / a1-programming-notes.pdf and recitation-02, and keep going. Backing up costs the whole session.
4 min. Do the shape arithmetic on the board once, slowly, with the inner pair circled and struck out. It is worth the chalk. The payoff line is the last one: naming the summed index tells you instantly whether the operation means anything. "Summed over units" is sensible; "summed over images" usually is not, and that is how they will catch a transpose error before it costs an hour.
4 min. This is the single most common silent error in the first assignment. Make it concrete: X * X.T on a square matrix runs and returns garbage. There is no error message and no nan. The only defense is the habit on the next slide.
4 min. Land this table — it is most of the assignment in three lines. Row 1 is where every dissimilarity matrix comes from. Row 2, after centering and dividing by n-1, IS the covariance matrix whose eigenvectors PCA returns — connect it back to last session explicitly. Row 3 is what pca.transform does, one component at a time. Then give the habit: every time you write .T, write the resulting shape in a comment. Three seconds, highest-yield habit in the course.
4 min. Answers: (120, 64); (64, 120); (120, 120). The third — square, and indexed by stimuli on both sides. Push one step after revealing: what does (1) mean? "Each of 120 images described by 64 numbers." What does (2) mean? The same thing, transposed — which is why the transpose error is invisible.
4 min. R02 gave them what an eigenvector is; this is the part they need for Part 3, and it is only about how to READ the output. Draw both curve shapes on the board. The cliff and the slow decay look completely different and lead to completely different claims, and students tend to report only the single number at 90%.
4 min. Plant this hard. Part 3 of the assignment is built on it and it is the most common way a student produces a confident wrong claim. The moral to state: dimensionality is a property of the data AND the stimulus set, never of the network alone. Do NOT give away the numbers they will get — that is theirs to find.
4 min. Emphasize p0. When a fit comes back absurd, the first move is a different starting guess, not a conclusion about the data. They will hit this. Note that they write the model function themselves — which means they have to have decided what shape they believe the relationship has. That decision is the science; curve_fit is the arithmetic.
4 min. R02 covered point 1; do it fast and spend the time on point 2, which is new and which they need for the three-model comparison in Part 2. Point 2 in one sentence: if you compare a 2-parameter law against a 3-parameter law and the 3-parameter one wins, you have learned nothing. Compare models with the same number of free parameters, or say plainly that you did not.
4 min. Answers: (i) they binned/averaged before fitting, shrinking the denominator; (ii) they fitted a model with more free parameters. Also fine: they averaged over subjects first, or their stimulus set had more spread in y. The point to land: a published number is a number computed under choices, and you cannot compare against it until you know the choices. This comes up directly in the assignment.
4 min. Give the reason properly: a network's distances and a human's similarity judgments are on unrelated scales, and there is no reason to expect a linear relation. The question you actually want is "do the two systems ORDER the pairs the same way?" — that is Spearman's question. Pearson would answer a different question and report a lower number for a reason that has nothing to do with the models. Students who use it will conclude the network is a worse match than it is.
4 min. Draw the matrix and shade the upper triangle. k=1 vs k=0 is worth one sentence: k=0 includes the diagonal, which is 120 guaranteed zeros in both matrices and will inflate any correlation.
4 min. This is a research-methods point, not a course technicality, and it is one they can carry into a lab. Say that. If asked "so is the correlation meaningless?" — no. The correlation is a fine descriptive statistic. It is the significance claim that is unsupported, and the distinction between those two is the thing to learn.
4 min. This is the conceptual centre of the session, so do not rush it. The analogy that lands: quoting one run of a random algorithm is like reporting one participant and calling it an experiment. Reproducible, and still not a result.
3 min. Short slide on purpose; it is the takeaway, and Part 4 of the assignment is exactly this exercise. Ask the room: how many seeds is enough? There is no correct number — the point is that one is definitely not enough, and that you say in the write-up how many you ran.
5 min. Spend the time. Every one of these three appears in published figures. perplexity in one sentence: roughly how many neighbors each point tries to stay near — small preserves local detail, large preserves coarse structure. Sweeping it changes the picture substantially, which is the point of the sweep in Part 4.
4 min. Answers: real = the two groups genuinely occupy separate regions of the high-dimensional space. Not real = t-SNE manufactured the separation at this perplexity, or this seed. The check: rerun across several seeds AND several perplexities and see whether the split survives. Accept "compare against distances in the original space" as an even better answer, and use it to introduce the next slide.
4 min. Neighborhood preservation. The `1:k+1` slice is worth pointing at: every point is its own nearest neighbor at position 0, and forgetting that shifts every answer by one. The last line is the good idea: a single average is a summary, but a per-point map tells you WHERE the embedding lied. That is a better figure than any average.
4 min. Read these out; do not elaborate. It is a checklist to photograph, and they will use it on Sunday night. Note for step 4: correlation distance produces nan for any constant row — a dead unit — which is exactly why the assignment makes them look for those.
2 min. Close by pointing at handout-03.pdf for the prose version, at office hours for anything assignment-specific, and at the fact that the next session (Week 4) starts the calculus you will need for Assignment 2 — derivatives, gradients, and the chain rule. Nothing in today's session carries forward to it.