- The readouts fitted in the geometry lectures and in the assignment, and every unit of the networks whose activations the first assignment analysed, compute the same thing: a weighted sum of their inputs, then a nonlinearity - Until now a library fitted the weights for us. This lecture opens the box: first the single neuron and the real neuron it abstracts, then how its weights are learned from examples. How a whole network of them is trained by backpropagation
- A spike is a brief electrical pulse, about a millisecond long, that travels along an axon. Other neurons' spikes arrive at this neuron's synapses, mostly on its dendrites - At each synapse, an arriving spike lets in a small current (next slide). A cortical neuron receives about 10,000 synapses; one input spike changes its voltage by roughly 0.1 to 1 mV, and firing needs about 15 mV above rest, so many inputs have to arrive close together - The currents spread to the soma and add up, only roughly: dendrites can compute on their own (London & Häusser 2005), and inhibition near the soma partly divides the input rather than subtracting from it - When the voltage at the axon hillock, the start of the axon, crosses threshold, voltage-gated sodium channels open and the neuron fires a spike, which travels down the axon to the next neurons' synapses (Hodgkin & Huxley 1952 described how) - The output the model keeps is the firing rate, the number of spikes per second; it discards when each spike happens - The neuroscience background notes on the course site (Recitations) cover these parts, spikes, synapses and ion channels in four pages, and map each onto the model neuron. Other short, non-technical introductions: OpenStax Psychology 2e, section 3.2 "Cells of the Nervous System"; OpenStax Biology 2e, sections 35.1 "Neurons and Glial Cells" and 35.2 "How Neurons Communicate"; Neuroscience Online (UTHealth), chapters 1 and 6; Gerstner et al., Neuronal Dynamics, section 1.1, the closest bridge to the model neuron
- The cell body and the axon hillock: where the inputs are summed and the spike starts
- The axon and its terminals: the spike travels to the next neurons. The notes of the first frame of this slide have the details
- The synapse produces a current; the small voltage change that current causes is called a postsynaptic potential (PSP). It is graded, not a spike - The neurotransmitter transporters in the drawing do reuptake: they pump transmitter out of the cleft, back into the axon terminal (or into nearby glial cells), which ends the signal and recycles the transmitter. Many antidepressants (SSRIs) block the serotonin transporter - Each neuron's synapses onto others are all excitatory or all inhibitory (Dale's principle): the sign belongs to the presynaptic cell - Learning changes synaptic strengths, and also how excitable a neuron is; the model keeps the first as its weights and the second, roughly, as its bias
- Each part of the model stands for one part of the neuron: the inputs and weights for the synapses, the sum for the soma, g for the spike threshold at the axon hillock, a for the firing rate. The model keeps one number per input and one number out, the firing rates, and discards when the spikes happen - Weights can change sign during learning; a real synapse keeps its sign. The bias has no single anatomical counterpart: it stands for how excitable the neuron is, how far its resting voltage sits below threshold, and any steady background input - Real rates are never negative; model inputs can be, once data are centred or z-scored - The abstraction is that of McCulloch & Pitts (1943). The next two slides build it in that order: weighted inputs and their sum, then the nonlinearity
- The same correspondences, one by one. Real rates are never negative; model inputs can be, once data are centred or z-scored - Minute paper 9 asks about this slide
- The indexing convention: x_1 to x_D, with x_i a generic input and D the number of inputs, the input dimension (the same D as the number of units in earlier lectures: a neuron's inputs are the units of the population it reads) - z is a weighted sum, not an average: the weights need not be positive or add up to one - The bias b shifts the sum up or down, and so sets how much input it takes to drive the neuron
- The sum and the dot product are the same number; the vector form is only shorter to write. Vectors are bold lowercase columns, and w is one unit's incoming weights - You met the dot product as a measure of alignment in the similarity lecture, and the projection onto a direction in the PCA lecture. A neuron computes that projection, then adds a bias and applies a nonlinearity - In the figure, the dark band is the projection of x onto the direction of w; its length is w^T x / |w|, and w^T x is that length times |w|. φ is the angle between x and w - The cosine form makes alignment literal: for inputs of a fixed length, the response depends only on the angle φ to w - The linear readout of the geometry lecture computed exactly this weighted sum and answered "yes" when it was positive. Here we give it a nonlinearity and read it as a neuron - w^T x is the one transpose in the course; a whole layer is written z^[1] = W^[1] x + b^[1] (next lecture) with no transpose, as in PyTorch's nn.Linear
- The readouts of the first part (decoding, shattering, CCGP) all fit a w and a b. The gold planes in the figures of hypothetical three-neuron populations are their decision boundaries, in three units instead of two - Why w is perpendicular: moving along the line keeps w^T x constant, so every step along the line is perpendicular to w. Moving along w changes z fastest - The bias sets where the line sits: the line crosses the direction of w at a distance −b/|w| from the origin - The next slide reads the same weights a second way, as a template
- One model neuron: 2,304 inputs (the pixels), one weight per pixel, a bias, and a sigmoid (the S-shaped nonlinearity of a later slide). Training it is logistic regression, which this lecture derives shortly - The photographs come from Caltech-101, a public dataset: its "Faces_easy" category was cropped and roughly centred by the dataset's authors, so no alignment was done here. Against the 435 faces, 435 photographs drawn at random from the other categories; all in grayscale, shrunk to 48 × 48 pixels - Both readings describe this same single neuron. The boundary reading asks which side of the line an input falls on; the template reading asks which input drives the neuron most. The next slide reads a unit inside a trained network the second way - The readout is right on 95.4% of held-out photographs. Its weights resemble the difference between the average face and the average non-face (correlation 0.73 across pixels), but they are not the same: the readout also learns to give little weight to pixels that vary a lot within each class, such as the background - A caution that matters for decoding brain data (Haufe et al. 2014): a decoder's weights are a filter, not a picture of the signal. A voxel or neuron can get a large weight because it helps cancel noise, not because it responds to the class
- With w and x both of length 1, wᵀx is the cosine of the angle between them; the ReLU nonlinearity, max(0, z), next slides, clips the negative ones to 0 - The tuning curve: the response as the input pattern is rotated from 0° to 180°. A later lecture measures such curves in real neurons - A two-pixel image and a two-weight neuron: set w to the bright pattern and the neuron responds most to that pattern - The comparison is fair only for inputs of the same length, since w^T x = |w| |x| cos θ; a brighter input raises the response too - The same weights have two readings, as the face readout showed: a decision boundary (a classifier) and the input pattern that drives the neuron most (a template). For a unit inside a network the template reading is the useful one; it comes back when we measure tuning in neurons and in networks
- ReLU passes positive inputs unchanged and sets negative ones to zero: a threshold at zero, so with ReLU the neuron is active only when z > 0, that is when the weighted sum exceeds −b - The activation a is one number. With ReLU it is never negative, like a firing rate - The whole neuron in one line: a = g(Σ_i w_i x_i + b), a weighted sum, a bias, a nonlinearity. Most networks in this course are built from many of these units
- g is a design choice. Inject a steady current into a neuron and it fires nothing below a threshold, then fires faster the stronger the current, and levels off because after each spike it cannot fire again for a millisecond or two (the refractory period). Each function keeps part of that curve - Step: McCulloch & Pitts (1943) and Rosenblatt's perceptron (1958) used an all-or-none output, 1 above threshold and 0 below. It captures the threshold, but its slope is zero everywhere, so gradient descent cannot use it - Sigmoid, from 0 to 1, and tanh, from −1 to 1: smooth S-shaped curves that made backpropagation possible (Rumelhart, Hinton & Williams 1986); tanh, centred on 0, was the usual choice in the 1990s (LeCun et al. 1998). The sigmoid is the most rate-like of the five: zero-ish below threshold, then rising and saturating. Its weakness is the flat ends: where the curve is flat, the gradient that learning follows is nearly zero, and deep stacks of sigmoids learn very slowly. This is the vanishing-gradient problem, which comes back in a later lecture - ReLU, max(0, z): the threshold-linear unit. Hahnloser et al. (2000) used it to model cortical circuits; Nair & Hinton (2010) and Glorot, Bordes & Bengio (2011) showed that it makes deep networks much easier to train, and AlexNet (2012) used it. Biologically it keeps the threshold and the rise, not the ceiling - Leaky ReLU (Maas et al. 2013) lets a small negative output through so that a unit never stops learning entirely. GELU (Hendrycks & Gimpel 2016) is a smooth ReLU with a small dip below zero; BERT and the GPT models use it - Negative outputs (tanh, leaky ReLU, GELU's dip) have no firing-rate counterpart unless the output is read as a change from a baseline rate - Without g, the neuron is linear: its output is the weighted sum itself. A weighted sum followed by a threshold-like g is the linear–nonlinear (LN) model, a common first model of a cell
- The rest of the lecture is about where the weights come from, for one neuron, a readout. The hidden layers of a network come next lecture - The four kinds of output differ only in the last step: what g is, and what the output is compared with. A yes/no output is compared with a label, a number with a measured value, a probability with the label, 1 or 0, and a set of class scores with the correct class. Each comparison gives its own loss, and the same downhill procedure, gradient descent, lowers any of them - Every weight learned today is judged the way the last lecture judged a score: on held-out images, against shuffled labels. With many units and few images, a readout fits even shuffled labels, so keep the generalization gap in mind whenever a training loss looks good
- The model neuron of the first part is the smallest such mapping: D numbers in, one number out. A network chains many of them - The next slides show four common choices of inputs and targets. The first needs people to label every example; the other three take their targets from the data themselves
- In every frame the training loop is the same: compute the output, compare it with a target, change the weights to shrink the difference. Only the source of the target changes. Three of the four need no hand-assigned labels - Labels: training lowers the cross-entropy (defined later in this lecture), which raises the probability given to the correct class. Someone had to label every image: ImageNet's 1.2 million photographs were labeled by hand, while text and unlabeled photographs are nearly unlimited - Next word: a language model predicts the next word (strictly, the next token, a word or piece of a word) from the words before it; the objective behind today's chatbots (Brown et al. 2020, GPT-3). BERT (Devlin et al. 2019) hides about 15% of the words and predicts them from both sides - Two views: DINO (Caron et al. 2021) trains a student network to match a teacher network (a slowly updated average of the student) on two crops of one image with changed colours. Tricks in the loss stop both from giving one output for every image. Masked autoencoders (He et al. 2022) hide about three quarters of an image's patches and reconstruct them - Image and caption: CLIP (Radford et al. 2021) trains an image network and a text network so that each image matches its own caption better than the other captions in the batch; about 400 million pairs from the web. People wrote the captions, but not to train a network
- A language model predicts the next word (strictly, the next token) from the words before it; the objective behind today's chatbots (Brown et al. 2020). No one labels the text
- Self-supervised learning (SSL): the data supply their own targets, with no labels from people. Next-word prediction (the previous slide) is self-supervised too - DINO (Caron et al. 2021) trains a network so that two crops of one photograph, with changed colours, give similar responses. No labels are used
- CLIP (Radford et al. 2021) trains an image network and a text network so that each photograph matches its own caption better than the other captions; about 400 million pairs from the web - Networks trained these four ways differ in how closely they match people's odd-one-out choices on the THINGS images (Muttenthaler et al. 2023). You test this yourself in the assignment
- The readout of the geometry lecture: a weighted sum of the responses and a threshold. Until now we fit it with a library call; here is what the call does - One neuron, a label for every input, and a rule that changes the weights. Three rules follow, and all have one form: change each weight by $\eta$ × error × input
- The perceptron is the model neuron from the start of the lecture, a = g(wᵀx + b), with a threshold for g and a rule that sets w from labeled examples: on a mistake, add ηyx to w (labels coded +1 and −1; η is a small step size); otherwise do nothing - Novikoff (1962): if a line separates the classes with some gap, the number of mistakes is bounded. The rule stops at the first line that works; which line depends on the starting weights and the order of the examples - If no line separates the classes, as for XOR, the rule never stops - Minsky and Papert (1969) proved that a single layer cannot compute XOR, and other limits like it. The perceptron is still the basis for understanding supervised learning; since then, the learning rule comes from an explicit loss, lowered by gradient descent, in many stacked layers. Hidden layers trained by backpropagation, the second half of this lecture, came in the 1980s - Rosenblatt's book-length account: Principles of Neurodynamics (1962) - The Mark I Perceptron (designed from 1958, publicly demonstrated 1960) had a 20 × 20 grid of photocells as input and motor-driven potentiometers as weights - After Minsky and Papert's book, interest and funding for neural networks fell for more than a decade - The figure: the same simulated points and order of examples, three different starting weights, no bias (the classes are separable by a line through the origin); each faint line is the boundary after one update, and the mistakes per pass are printed under each run. Compare the XOR figure later in the lecture, where the mistakes never reach zero
- The top curve is a made-up loss, ℓ(w) = (w − 2)² + 0.5, chosen so the minimum is easy to see; a real loss is computed from the training examples, but it is still a function of the weights - The gradient generalizes the derivative from one variable to many. In calculus you found the minimum of a function by setting its derivative to 0; with millions of weights there is no formula to solve, so we walk downhill instead - The partial derivative ∂ℓ/∂w_i is the ordinary derivative with every other weight held fixed - The gradient points in the direction in which the loss grows fastest, and it is perpendicular to the contour lines of the loss. Its opposite is the steepest way down - η, the learning rate, sets the size of each step. Too large and the steps overshoot; too small and training is slow. A later part of this lecture returns to it - If derivatives are new to you: the Recitation 4 handout (gradients and the chain rule, course site, Recitations) starts from the slope of a line
- The perceptron's rule was a heuristic: Rosenblatt did not derive it from a loss, and its error is either 0 or not, with no slope to follow. Widrow and Hoff (1960) compared the target with the output before the threshold, ŷ = wᵀx, which makes the error smooth, and derived the rule as gradient descent on that squared error. The result has the perceptron's form, η × error × input, which is why the delta rule is often presented as its generalization - A loss is the number a learning rule tries to make small: here the squared error ℓ of one example. Averaged over the training examples it is the loss ℒ of the whole training set; the course writes ℓ for one example and ℒ for the average - Stepping against the gradient, one example at a time, is stochastic gradient descent - Sign convention: the gradient is (ŷ − y)x; the update is minus the gradient, (y − ŷ)x. Books use both forms - Names: the delta rule, the Widrow–Hoff rule, the least-mean-squares (LMS) rule - Unlike the perceptron, it updates on every example, even correct ones: it aims at the value y = ±1, not only the correct side. With the threshold put back, (y − ŷ) is 0 on correct examples and ±2 on mistakes, which is the perceptron's update - With a continuous target, such as a firing rate, the same rule is linear regression: a model of a neuron's response, in the next lecture
- The demo is on the course site with the other demos. Start from w = 0: the first example is always a mistake - Watch the number of mistakes per pass fall to zero when a line separates the classes. Add flipped labels and it never does: the perceptron keeps moving - The menu switches to the delta rule and logistic regression, which update on every example by how wrong it is, the rules of the next two slides
- The delta rule aims at y = ±1 even for points far on the correct side. Logistic regression outputs a probability instead, and its error for a clear case is near zero - Despite the name, a classifier: one neuron whose nonlinearity is the sigmoid σ of the transfer-function slide - The cross-entropy is minus the log of the probability given to the correct class: −log p̂ when y = 1, −log(1 − p̂) when y = 0 (bottom panel). At p̂ = 0.5 it is 0.69; a confident mistake, p̂ = 0.01 for a true class of 1, costs 4.61 - Three updates: perceptron ηyx on mistakes only (y = ±1); delta rule η(y − ŷ)x on every example; logistic regression η(y − p̂)x on every example (y = 1 or 0) - The gradient is this simple because the slope of the sigmoid cancels against the slope of the log. The update is the delta rule's with p̂ in place of ŷ - Averaged over the training set, the cross-entropy is the loss ℒ. Confident and wrong means, for example, p̂ = 0.1 when the true class is 1
- Rosenblatt (1958), The perceptron: a probabilistic model for information storage and organization in the brain. The original proposal, written for psychologists - Zipser & Andersen (1988), A back-propagation programmed network that simulates response properties of a subset of posterior parietal neurons. The hidden units of a small trained network, compared with recorded neurons: the template for many later model–brain comparisons - Lehky & Sejnowski (1988), Network model of shape-from-shading: neural function arises from both receptive and projective fields. A neuron's receptive field alone may not reveal what it computes - Lillicrap, Santoro, Marris, Akerman & Hinton (2020), Backpropagation and the brain. A readable review of how error signals might be carried in cortex; the companion Lillicrap et al. (2016, Nat Commun) shows that random feedback weights can replace the transposed weights of step 3 - Muttenthaler, Dippel, Linhardt, Vandermeulen & Kornblith (2023), Human alignment of neural network representations - Francioni, Tang, Toloza, Ding, Brown & Harnett (2026), Vectorized instructive signals in cortical dendrites - For the mechanics, the 3Blue1Brown videos on neural networks (chapters 3 and 4) animate the four steps of backprop - Before posting, skim the existing entries on your paper; your post must add something not already said
- Submitted on Canvas (Minute papers → Minute Paper 9); credit for a thoughtful attempt