=========================================================================== CPSY 1291 — RECITATION 5: PyTorch, tensors, autograd, and the training loop TA-led, 80 minutes. Optional. Week 5 — hold this AFTER Lecture 8 (MLPs and backpropagation). Assignment 2 is due Tue 10/13. HOW TO RUN THIS DECK: presenter notes carry timing, what to say, and the answers to the exercises. SCOPE: reading and debugging a training loop, not writing one from scratch. The assignment PROVIDES its training loops; the skill being taught is knowing which lines matter and how to diagnose one that fails. RUN THE CODE LIVE if the room has a projector and a Colab tab. The single most valuable five minutes of this session is deleting opt.zero_grad() in front of them and watching the loss explode. MATERIAL: handout-05-pytorch-training-loop.pdf covers the same ground in prose, with a seven-item debugging checklist and six exercises. ============================================================================
2 min. Frame the session honestly: you will not have to WRITE a training loop in this course. You will have to read one, change one, and fix one. That is what today is. If you can run code live, say so now and promise the zero_grad demo — it gives the room a reason to stay awake through the tensor mechanics.
4 min. Land this properly; everything else follows from it. Connect back to last week explicitly: they computed dL/dw by hand and checked it with a two-sided difference. backward() is that same quantity, obtained by bookkeeping instead of by algebra, for every parameter at once.
5 min. Put the error message on the board in full — recognizing it by sight saves each of them twenty minutes. Second gotcha, worth stating here: classification TARGETS must be long (integers), shape (N,), as class indices — not one-hot, not floats. That error message is much less readable, so warn them in advance.
4 min. The permute point produces a confusing error and comes up the moment they try to look at a filter, which is Part 1a of the assignment. Say it now. The (N, C, H, W) convention has a name worth mentioning — NCHW — because they will see it in documentation and in error messages.
4 min. Run this live if you can. Seeing 4.0 appear where they know the answer is 2w is what makes autograd stop being magic. Then perturb it: change to w**3 and ask the room what .grad will print before you run it. 3w^2 = 12.
7 min. Fact 1 is the demo. Delete opt.zero_grad() from a working loop in front of them: the loss falls for a few epochs and then explodes. Then put it back. If asked WHY accumulation is the default rather than a bug: it lets you sum gradients over several batches before stepping, which is how people train models too big to fit a batch in memory. It is a feature that costs beginners one line of attention.
4 min. The symptom worth naming: a script that trains fine and then runs out of memory during evaluation. Almost always a missing no_grad, because the graph is being kept for every evaluation batch. Feature extraction from a pretrained network is the case they will hit in the assignment.
5 min. The (out, in) convention confuses everyone once. Say it now so that when they inspect AlexNet's first layer and get (64, 3, 11, 11) they read it as "64 filters" rather than "64 inputs". Ask the room what happens if you delete the ReLU. Answer: the stack collapses to a single linear map — more parameters, no more power. That is Lecture 8's point and it is worth hearing twice.
7 min. Spend the time — this is the single highest-value slide in the deck, because it is invisible and because they will meet it in copied code. Make the ReLU case vivid: a classifier works by accumulating evidence for AND against each class. ReLU destroys all the evidence against. The model still trains, still improves, and still ends up much worse, with no error anywhere. Give the rule to write down: if the loss has "CrossEntropy" in its name, the last thing in your model is a Linear layer. Full stop.
6 min. Read it as five jobs, not eleven lines. Have someone in the room name each job before you reveal the comments. The highlighted line matters for debugging: if weights are not changing, the question is whether step() is being called, not whether backward() is working. Mention .item(): it pulls a Python float out of a one-element tensor. Printing the tensor works but drags the whole graph along.
4 min. Answers: (i) a missing opt.zero_grad(), so gradients accumulate; (ii) a learning rate far too large. Telling them apart: a missing zero_grad usually gives you a few GOOD epochs first, because early gradients are small. Too large a learning rate usually misbehaves from the very first steps. Fix zero_grad first; if it still diverges, it is the learning rate.
4 min. When someone says "the network did not learn", the first question is which optimizer and the second is what learning rate. Make that a habit. Do not explain Adam's internals. It is enough that it keeps a per-parameter step size and that this makes it forgiving.
5 min. Read the shape out loud: 64 filters, 3 color channels, 11 x 11 pixels. Then note that VGG's first conv is 3 x 3 — too small to show much structure at all, which is a useful fact and a reason model choice is not arbitrary. To display one filter: transpose(1,2,0) then rescale to [0,1], because imshow of arbitrary floats is not meaningful.
6 min. Ordered by how often each is the answer. Item 3 fixes more cases than everything else combined; say that. Item 7 is the one to remember and deserves its own moment — next slide.
4 min. This is what an experienced person does first, and it separates "broken pipeline" from "hard problem" faster than any other check. Note the connection forward: next week is about the opposite failure — a model that memorizes and does NOT generalize. Both are about the same quantity, seen from two sides.
4 min. Answers: (64*16 + 16) + (16*10 + 10) = 1040 + 170 = 1210. Function class: exactly the linear maps from R^64 to R^10 — the same class as a single nn.Linear(64, 10), with 1,210 parameters instead of 650 and no more expressive power. Push: so is the extra layer useless? Not quite — it constrains the map to rank at most 16, which is a real (and sometimes useful) restriction. Nice place to mention that this is exactly what a linear autoencoder does, which is Lecture 10.
3 min. Point at handout-05.pdf, especially the seven-item checklist, which is the part worth photographing. Close on the preview: today was "can it learn at all", next week is "did it learn the right thing", and those turn out to be almost unrelated questions.