=========================================================================== CPSY 1291 — RECITATION 7: Midterm review (Lectures 1-8) TA-led, 80 minutes. Optional. Week 7 — MUST be held before Thu 10/22. FORMAT: this is a PROBLEM SESSION, not a lecture. The slides are prompts. Put a problem up, give the room 3-4 minutes, take answers from students before revealing. Resist the urge to talk through the summary slides — they exist so the room can check itself, and the handout carries the full version. BUDGET: about 15 min on the compressed summary (slides 2-5), about 55 min on problems, and 8 min of open questions at the end. If you fall behind, cut problems 5 and 6 — not the open-question block, which is what the students who came actually came for. DO NOT: predict what is on the exam, or rank topics by likelihood. Say you do not know. Point at the learning objectives in the syllabus instead. MATERIAL: handout-07-midterm-review.pdf has the full compressed summary of L01-L08 and fourteen problems with worked answers. Tell them at the start that it exists so nobody transcribes the board. ============================================================================
2 min. Say the format immediately: this is a problem session. You will put questions up and the room answers them. Fifteen minutes of summary, then an hour of problems. Say the thing about the handout: it has the full compressed summary and fourteen worked problems, so nobody needs to transcribe anything. If asked what is on the exam: you do not know. Point at the syllabus learning objectives. Do not speculate — it is unfair to whoever is not in the room.
3 min. This is the most useful framing you can give them, so give it first. Illustrate with one example they know: PCA. The object is the eigendecomposition of the covariance. What it is for is finding directions of greatest variance. When it misleads you: it is linear, it is scale-sensitive, and without centering PC1 is the mean. All three are examinable, and only the first is what students revise.
5 min. Read the bullets, do not expand them. This slide is a checklist for the room to audit itself against — anyone who cannot expand a bullet knows what to revise tonight. Ask for a show of hands on which bullet is least familiar, and spend two minutes on whatever wins. Do not spend six.
5 min. Same treatment. The two most misunderstood items on this slide are "universal approximation says possible, not findable" and "backprop is not an optimizer" — flag both explicitly, because both are natural exam questions and both are usually answered wrong.
5 min. Answers: both are 0. Cosine because the vectors point in exactly the same direction. Correlation because after centering each row both become proportional to (-1, 0, 1). Part 3 is the real question: these units carry the same information about the stimuli and differ only in gain. Neither distance can see that; Euclidean can. Which one you want is a modelling decision. Follow-up if the room is fast: what if the second unit were (3, 2, 1)?
4 min. Answers: the triangle inequality, since 0.9 > 0.4. Cosine distance and correlation distance both permit it. Push: is that a defect? No — and this is worth landing. Human similarity judgments violate the triangle inequality too, so a distance that permits it may be the better model of the behaviour. A "violation" is only a defect if you needed the axiom.
4 min. Answer: it points approximately at the MEAN image, because the uncentered second-moment matrix is dominated by it. You have spent your first component describing the average face, and the structure moves to PC2 onward. Connect it forward: this is the same reason Hebb's rule needs centered data — Problem 6. Both are "the second moment is not the covariance unless you center." That one sentence covers two lectures.
5 min. Answers: (1) MDS axes carry no inherent meaning — the solution can be rotated or reflected freely, so "the horizontal axis" is not a property of the data at all. (2) Independent evidence: correlate the coordinate with an external animacy rating, or show a classifier trained on that coordinate alone predicts animacy on held-out stimuli. The general lesson is worth naming: naming an axis is a claim that requires evidence outside the plot. Assignment 1 Part 3j is exactly this.
5 min. Answers: toward the mean DIRECTION of the data, not the direction of greatest variance. Hebbian dynamics follow the second-moment matrix, which equals the covariance only after centering. Fix: subtract the mean. Ask them to connect it to Problem 3. Same fact, two lectures apart.
5 min. Answer: dw = eta * a * (x - a*w). The -eta a^2 w term is a decay proportional to the unit's own activity: it grows exactly when the weight grows, driving ||w|| toward 1. What it does NOT change: the DIRECTION of convergence. Hebb already finds PC1's direction; Oja fixes the magnitude. Students often say Oja "makes it find PC1", which is half wrong and is a good exam trap.
6 min. Answers: (1) the perceptron updates only on mistakes and halts as soon as there are none, so it stops at whichever separating boundary it reaches first — and which one that is depends on the order the mistakes arrived in. (2) No: the delta rule minimizes a squared error with a single minimum, and converges to the same answer regardless of order. Land the general point: "it converged" and "it converged to a unique answer" are different claims. Assignment 2 Part 3c measures exactly this spread.
5 min. Answers: V_B stays near zero. The error term is lambda - (V_A + V_B), and V_A already accounts for the reward, so there is no error left to drive learning about B. This is BLOCKING. Why it mattered historically: it shows animals learn from prediction ERROR, not from mere co-occurrence. B co-occurs with reward on every one of 50 trials and is still not learned. That result is why R-W is a landmark rather than a footnote. Follow-up: what if phase 1 is removed? Both cues learn, sharing the association.
6 min. Answers: (1) a hidden unit whose pre-activation is negative for all four inputs — its ReLU gradient is zero, it never updates, and the network is effectively 2-1-1, which cannot solve XOR. (2) Run the four inputs through the trained network and count, per hidden unit, how many produce a non-zero output. Zero for all four means dead. Also accept: both hidden units converging to the same function, or a saddle. This is Assignment 2 Part 4e. If anyone has done it, let them describe what they found rather than describing it yourself.
6 min. Answers: 1 is true as stated. 2 is FALSE — existence is not findability, and the theorem says nothing about training. 3 is FALSE — backprop computes the gradient; gradient descent (or Adam) does the optimizing. Backprop supplies the derivative, nothing more. Both false statements are things students write on exams because both sound like reasonable paraphrases of lecture. Say that out loud.
4 min. Read them out. This is the "night before" list and it is in the handout. Each of these has appeared in a lecture, an assignment, AND a recitation — which is the reason they are on the list, and worth saying, because it tells them the list is not arbitrary.
8 min, and protect it. This is what the people who came actually came for. If nobody speaks, prime the room with a question of your own: "who can tell me why we center the data before PCA?" — then let the discussion go where it wants. Close by saying that Assignment 3 is due Tue 10/27 and that recitation 8 covers comparing representations — RSA, decoding, and permutation tests.