=========================================================================== CPSY 1291 — RECITATION 1: Python and NumPy TA-led, 80 minutes. Optional session; students who already program comfortably in Python should skip it. HOW TO RUN THIS DECK: every slide's presenter note gives a timing, what to say, and (where relevant) the answer to the exercise on the slide. Slides marked EXERCISE are meant to be done live — put the prompt up, let them try for 2-3 minutes, then take answers from the room before revealing. SCOPE: reconciled against the FINAL Assignment 1 (2026-09-01). Everything here is required by A1 parts 1-2, which students are starting this week: axis/keepdims and the broadcasting error (1b, 2c), boolean arrays (1f, 2i, 3e), and fancy indexing (1e, 1f, 2g). Nothing else is here. Do NOT add gradients, the chain rule, or how networks are trained — those arrive with Assignment 2 and its own recitation. Concepts (Shepard, psychological space) are carried by L1/L2, not here (Recitation 2 keeps only pointers). MATERIAL: the companion notes (linear-algebra-notes.pdf / a1-programming-notes.pdf) cover the same ground in prose. Tell students at the start that they do not need to take notes. NOTATION (recorded decision, 2026-09-01): the recitation decks follow the notebooks — n items x p features, lowercase, matching the code ((n,), (n, n), n x p). Lecture's D (input dimensionality) and N (sample count) are the same quantities; the mapping is said once, in the slide-4 presenter note. D on a slide means one thing only: a distance matrix. ============================================================================
2 min. Introduce yourself, say office hours, and set expectations: this session and the next are OPTIONAL, and they exist so that nobody loses a weekend to a shape error. Anyone comfortable with numpy broadcasting can leave now with no penalty — say that out loud, it keeps the room self-selected and the pace honest. Say what this is NOT: not a Python course, not comprehensive. It is the specific set of habits Assignment 1 rewards. A1 parts 1 and 2 are out; everything in this session appears in them, under its real section label.
2 min. The third bullet catches more people than they expect, including strong CS students. Encourage staying for the broadcasting and indexing sections even if the early part is familiar — those are the parts that actually cost people time in A1.
3 min. This is the most important slide in the deck; do not rush it. Make it concrete with a story: a student spends an hour debugging a distance function, and the actual problem is that they never decided whether the distance should ignore brightness. No error message will ever tell you that. Connect it to how they will really work: outside this course they will use AI tools to write code, and those tools are only useful if you can state what you want and recognize a wrong answer. That judgment is what we are building here. Thomas is explicit about this — assignments are their own work, but the skill being trained is deciding WHAT to compute, not typing it.
2 min. Lecture writes D for the dimensionality; the notebooks and these decks write p (n items x p features) — same thing, say so once here. This is the bridge from "programming session" to "why we are here", and it is the one conceptual slide in an otherwise practical deck. The psychology behind it — Shepard, psychological space — is lecture material and Recitation 2's opening; the notebooks also carry glossary cells for the vocabulary. Draw it: two axes on the board, a few dots. Ask what "close" means, and take two or three answers — you want someone to say "depends what you mean by distance". That is exactly the assignment's first question, so tell them they have just written 1b themselves.
3 min. The `xs * 2` row is the demo to actually run live — it surprises people from every background, and it is a real source of silent bugs. Keep the asides below to one sentence each; the table carries the slide. MATLAB refugees: reassure them arrays behave much like MATLAB matrices, with the crucial difference that indexing starts at 0 and the last index is n-1. Java refugees: no type declarations, no compile step, and indentation is syntax.
5 min. "Unit test" needs a plain version for the psych-only half of the room: your automatic correctness check — the thing you look at to know a step worked. Say both. Have them type along. Then break something on purpose: transpose a matrix and show the error message, so the first time they meet "could not be broadcast together with shapes ..." it is in a room with help rather than at 2am. The stimuli-by-features convention is assumed by every assignment and by the lectures. Say it twice. A1's own arrays follow it: 120 images x 4,096 late-layer units, and it is the shape every function in 1b takes.
3 min. Slices are VIEWS, not copies — modifying a slice modifies the original. Mention it once; do not dwell. IF AHEAD: np.shares_memory(X, X[::24]) -> True — proof that slices are views, and why mutating a "copy" can corrupt the original; X[::24].copy() is the escape hatch. The X[::24] example lands better if you ask first: "if the data goes object 1 from 24 angles, then object 2 from 24 angles, how would you get one picture of each object?" Let them derive the stride. In A1 3c they will use F[::24] to get 51 images at a single viewpoint.
6 min. THE central slide of the session. Do not move on until the rule has been said back to you by someone in the room. Draw the (120, 4096) grid on the board and physically cross out the axis being collapsed. The mnemonic is reliable and students keep it. "axis=0 collapses images, leaving one number per unit" — that sentence is 2c's dead-unit check (variance across the 120 images, per unit) in a nutshell. keepdims will look pointless right now. Say only: "it exists so subtraction lines up — we will see why in three slides." Then deliver on that.
4 min. Answers: (1) (4096,) (2) (120,) (3) X.mean(axis=0) — averaging down the rows gives one number per feature, which is the average stimulus. Question 3 is the one that separates rule-following from understanding. If the room struggles, go back to the crossed-out grid.
6 min. Here is the promised payoff for keepdims, and the single most common error in A1 part 1 — show the traceback live, on purpose, so the first time they meet it is in this room. Make the difference between the two working lines concrete: centering each column removes "this pixel is bright on average"; centering each row removes "this image is bright overall". They lead to different answers to the same question, and in 1b the row version is what turns cosine distance into correlation distance. Walk the error: (120,4096) vs (120,) aligned from the RIGHT pairs 4096 with 120. Neither is 1, so NumPy refuses. keepdims=True makes it (120,1), which stretches. That is all keepdims is for.
4 min. Answer: B. A removes each FEATURE's mean across stimuli — a different, also useful operation (PCA does exactly A, as part 3 will point out). Then ask what `X - X.mean(axis=1)` (without keepdims) does: the error from the previous slide. Having just seen it, someone will say so — good, that means it stuck. B is one line of correlation_dist in 1b.
3 min. Three vocabulary items, nothing else — the next slide is unreadable without them, so land each one on a shape print. @ between arrays is matrix multiplication, NOT the decorator @ that sits above function definitions in Python tutorials — say that out loud, the Java/CS students have seen the decorator and will misread it. [:, None] is two-axis indexing (row spec, column spec) with None standing in for "new axis of size 1". Tie it straight back to keepdims: X.mean(axis=1, keepdims=True) and X.mean(axis=1)[:, None] produce the same (120, 1) array.
4 min. This is 1b's HINT block, read slowly, line by line. The notebook is explicit that the Euclidean idiom is given because the definitions are the point, not the idiom — so do not re-derive it, just decode it with the three symbols from the previous slide: sq[:, None] is (n,1), sq[None, :] is (1,n), and broadcasting turns their sum into every pairwise combination; X @ X.T fills the cross term. The algebra bullet is the same expansion they saw in high school, (a-b)^2 = a^2 + b^2 - 2ab, done to whole rows — say it that way. IF AHEAD: np.einsum('ij,ij->i', X, X) is the show-off spelling of (X**2).sum(axis=1), and scipy.spatial.distance.cdist(X, X) is what you import in real life — 1b makes you write it yourself because the definitions are the point.
3 min. The cosine skeleton is the two lines on the slide: normalize rows (keepdims again, doing real work), then 1 - n @ n.T — the same @-and-.T move as the Euclidean cross term. Correlation is the same two lines after row-centering — the notebook says "one line different" and this is what it means. The maximum(D2, 0) clip is not paranoia; they WILL hit sqrt of -1e-13 -> nan. Say it now so they recognize it later.
4 min. argsort returns INDICES, not values — say it, then show it, because the confusion is universal. mask.mean() is the quiet star: 1f asks for the PERCENTAGE of 200,000 triples that violate the triangle inequality, and that is one .mean() on a boolean array. Flag it here, deliver it three slides on. Motivate argsort: "the five images people rated most similar to this one" is an argsort on a row of the similarity matrix — 2a makes them do exactly that, and it is how they check the data was loaded correctly.
5 min. Every one of these lines is lifted from A1. np.isin(cat, [0, 2]) is 3e's animate indicator — Animal is category 0 and Face is category 2, so that one call builds the binary variable the whole animacy claim rests on. (The boolean-argmax "first True" trick for 3b now lives in Recitation 2's variance-explained slide, next to cumsum — do not teach it here.) Show the and/or error live: (taxon == 2) and (taxon == 5). Read the message aloud — Python is asking "do you mean ANY of these, or ALL?", because a 120-element boolean is not one truth value. any()/all() answer that question when you really do want one truth value; &/| when you want 120 of them. 2i's contingency table is a nested loop with a boolean & — mention it as the place parenthesized & shows up next.
3 min. Answer: np.argsort(S[7])[-2] The trap is [-1], which returns stimulus 7 itself — nothing is more similar to something than itself. Let the room fall into it; do not warn first. It is a memorable lesson about checking that an answer is sensible, and 2a says it in so many words: "an image is always maximally similar to itself, so exclude the query image from its own neighbor list." Optional follow-up, only if the room is quick: how would you get the five most similar, excluding itself? Answer: np.argsort(S[7])[::-1][1:6] IF AHEAD: np.argpartition(v, 5)[:5] gets a top-k without a full sort — O(n) vs O(n log n); one %timeit on a 10^6-element array makes the point in ten seconds.
6 min. Work the 4x4 on the board until M[a, b] is obvious: element-wise pairs, NOT a submatrix. Without this idiom students write a Python loop over 200,000 triples, watch nothing happen for minutes, and conclude Colab hung — that is the failure mode this slide exists to prevent. The bottom block is 1f's own skeleton (the notebook gives the sampling lines). Two things to say about it: compute each full distance matrix ONCE and index into it — recomputing distances inside a loop takes minutes rather than seconds, and the notebook warns exactly that; and the + 1e-9 tolerance is not decoration — 1d showed your functions agree with SciPy to ~10 decimal places, and a bare > counts every rounding disagreement as a violation. And there is the promised boolean .mean(): the violation rate in one call. IF AHEAD: %timeit the vectorized triple check against the naive Python loop — the "minutes rather than seconds" claim, measured. And np.add.at is the write-side sibling of fancy read-indexing.
4 min. D[iu] is the same move as the last slide — iu is literally a pair of index arrays — and 7,140 is no magic number: 120·119/2, every pair once, diagonal excluded. It matters because the assignment leans on it repeatedly: 1e's provided check correlates upper triangles, 2g pairs s_ij = S_human[iu] with MDS distances, and 2j/2k reuse the same pair order. Say the phrase "same pair order" now; it is why two D[iu] vectors can be correlated at all. np.ix_ is 1e verbatim: order = np.argsort(taxon), then np.ix_(order, order). The contrast on the slide is the lesson — D[order, order] silently applies the PAIRED rule from the previous slide and hands back 120 numbers, no error. If a student's "sorted" RDM comes out as a vector, this is why. np.ix_ builds the open grid: every row in order against every column in order.
2 min. These lines are provided in the notebook — read, don't write. Three names to gloss in passing: plt is matplotlib (imported in the setup cell), DATA is the setup cell's data-path variable, and f'{...}' is Python's string formatting — new syntax for the Java folks, and not something they must write. An .npz is a zip of named arrays — index by name, not position. The loading cells in A1 are PROVIDED, so nobody has to write this; the point is reading them and knowing what came back. The second block is the habit to instill: look at the data before trusting a number computed from it. 1a makes them do this explicitly (and its output — most activations are exactly zero, because ReLU — answers a question shapes alone cannot), and part 3's recordings have real NaNs waiting for exactly the np.isnan check.
2 min. Walk down the list in order — the discipline is the point, and the first two catch most problems. Item 4 is 2c verbatim: the untrained network has hundreds of zero-variance units, and 2c asks what they do to a correlation. Tell them to bring a failing cell to office hours having already done steps 1-3, and that doing so usually means they no longer need office hours.
2 min. Closing the loop matters — students sit through a methods session much more willingly once they can see the assignment on the other side of it. Read the table out; every row carries a real section label they can go find. The concepts — psychological space, Shepard's law, why any of this is science — are what L1/L2 carry this week — Recitation 2 keeps them to pointers; the notebooks also start with glossary cells for the vocabulary. Point them at linear-algebra-notes.pdf / a1-programming-notes.pdf for the prose version, and take questions with whatever time is left.