=========================================================================== CPSY 1291 — RECITATION 8: Comparing representations TA-led, 80 minutes. Optional. Week 8 — hold this after Lecture 12. Assignment 4 is due Tue 11/10. SCOPE: the machinery behind RSA and decoding, and the three statistical traps that make published versions of both analyses wrong. Lecture 12 supplies the motivation and the science; this supplies the code and the caveats. THE DEBT: Recitation 3 promised that the meaningless RSA p-value would be fixed by a permutation test "in Recitation 8". This is that session. Say so explicitly — students remember being told something was owed. TIMING: the permutation-test block (slides 8-11) is the core. If you run short, cut CKA and shorten the decoding block; do not cut permutation or the noise ceiling. MATERIAL: handout-08-rsa-decoding-permutation.pdf, with six exercises. ============================================================================
2 min. Open by paying the debt: five weeks ago we said the p-value from correlating two RDMs was meaningless and that a permutation test would fix it. Today is the fix. Frame the session: Lecture 12 gave you the science. This is the code, plus the three ways these analyses go wrong in print.
4 min. This is the frame for the whole session, and it is the single most portable idea in it. Give the concrete version now so it lands: a model scoring 0.35 against a ceiling of 0.40 is doing very well; the same 0.35 against a ceiling of 0.85 is doing poorly. The raw number cannot tell those apart.
6 min. Name the four: which distance; whether to standardize the units first; Spearman or Pearson; and the stimulus set. Say why correlation distance is the field's default for neural data: it removes each stimulus's overall response level, which is often driven by attention or arousal rather than by stimulus identity. Euclidean keeps it. Neither is neutral and you have to say which you used.
6 min. This is the most important methodological slide of the session, and it is not in the lecture. Make it vivid: two models that disagree about everything can produce nearly identical RSA scores on a stimulus set with one dominant division. The literature has this problem, and it is why stimulus-set design gets as much scrutiny as model design in careful papers.
6 min. Walk through the logic before the code. A p-value from a formula is an answer to this question computed under assumptions; a permutation test answers it by direct simulation, and works when the assumptions do not. np.ix_(p, p) is the line to point at: it reindexes rows AND columns with the same permutation, which keeps a valid dissimilarity matrix.
6 min. The first point is the one that produces wrong papers. If a student's permutation test says every model including a random one is significant, this is why, every time. The +1 has a reason worth giving: your observed value is itself one possible arrangement, so it belongs in the null count. Reporting p = 0 claims you have proved something impossible with 10,000 samples.
5 min. This is the diagnostic that catches the previous slide's error, and it takes two lines. General habit worth naming: whenever a procedure produces a single summary number, look at the distribution it came from at least once. That applies to p-values, to cross-validation scores, and to loss curves.
6 min. Say why this changes conclusions and not just presentation: model rankings in this literature sometimes REVERSE when the ceiling is added, because different datasets have different ceilings. Counterintuitive point worth making: a high noise ceiling makes your models look WORSE. Clean data is harder to explain, and that is honest rather than unfortunate.
6 min. Answers: (1) the noise ceiling — is 0.02 large relative to the achievable range? (2) A confidence interval on the DIFFERENCE, obtained by bootstrapping stimuli with replacement. The second is the one they will not say. Emphasize: bootstrap the difference, not the two scores separately — overlapping intervals on two quantities do not imply their difference is uncertain.
5 min. Two different questions, two different resampling schemes — permutation shuffles WITHOUT replacement and destroys the relationship; bootstrap samples WITH replacement and preserves it. Students conflate these constantly. Say the distinction in those terms and it sticks.
5 min. Same leak as Recitation 6, in its most common disguise — and it is invisible, because the code looks tidy and the numbers look better. Tell them the size of it: a few percentage points, which is exactly the size of effect people publish.
6 min. Spend the time; all three appear in published work. Point 3 is the interpretive limit of the entire method and belongs in any write-up that uses it. The clean way to say it: decoding tells you what information is PRESENT, not what the brain DOES with it. Point 1's fix: report balanced accuracy or the confusion matrix, ideally both.
6 min. Model answer: category information is linearly readable from V1's population response under these conditions. That does not establish that V1 represents category — low-level features correlated with category in this stimulus set would produce the same result, and nothing here shows that the rest of the brain uses this information. This is the single most useful thing in the session for anyone going into a lab. Let the room argue about it.
6 min. Both centerings are per-unit (column). Forgetting them inflates the score toward 1 for almost any pair of matrices — a good bug to warn about, because the result looks like a strong finding. Callback to Lecture 5: the metric zoo disagrees with itself, and that is a live problem in the field rather than a technicality. Running two is cheap.
3 min. Point at handout-08.pdf, especially the cheat sheet and the six exercises. Preview next week honestly: today was about comparing a network's representations from the outside; next week is about opening it up, and the methods for doing that are much less trustworthy than the ones today.