Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 8 · Theme 2: Learning representations

Learning the weights: from neurons to gradient descent

Thursday, October 8 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

You have been using model neurons all along

Two clouds of points in the space of three units x1, x2, x3, blue upper right and red lower left, separated by a gold hyperplane, the decision boundary, labelled w transpose x plus b equals 0; an arrow labelled w points from the plane toward the blue points
  • The readouts of the last lectures (animate or not, context, shattering, CCGP): a weighted sum of the responses, then a threshold: a decision boundary, a hyperplane in the space of units
  • The activations of the networks in the assignments: each number is one unit's output
  • Both are model neurons. Today: what one computes, how it relates to a real neuron, and how it learns its weights

What a real neuron does

A drawing of one neuron: branching dendrites on the left, the cell body with its nucleus, the axon hillock, a long axon and branching axon terminals on the right. The dendrites and the incoming synapses from other neurons' axons are highlighted; the rest is faded
  • The dendrites collect inputs from other neurons: at each synapse, an incoming spike becomes a small electrical current

What a real neuron does

The same neuron drawing with the cell body (soma) and the axon hillock highlighted and a small spike leaving the hillock; the rest is faded
  • The dendrites collect inputs from other neurons: at each synapse, an incoming spike becomes a small electrical current
  • The cell body (soma) integrates them; past a threshold, the axon hillock fires a spike, a brief electrical pulse

What a real neuron does

The same neuron drawing with the axon and the axon terminals highlighted and small spikes travelling along the axon to the right; the rest is faded
  • The dendrites collect inputs from other neurons: at each synapse, an incoming spike becomes a small electrical current
  • The cell body (soma) integrates them; past a threshold, the axon hillock fires a spike, a brief electrical pulse
  • The spike travels down the axon to the synaptic terminals: the inputs of the next neurons
  • Output: spikes; their number per second is the firing rate

At a synapse: a spike becomes a current

Schematic of a chemical synapse: the axon terminal of the presynaptic neuron holds vesicles filled with neurotransmitter; an arriving action potential opens voltage-gated calcium channels, vesicles fuse and release neurotransmitter into the synaptic cleft; it binds to receptors on the postsynaptic dendrite, which open and let ions flow in
  • A spike reaching the axon terminal releases neurotransmitter into the gap
  • It binds receptors on the next neuron, which open channels: ions flow in, a small current
  • Excitatory synapses push the neuron toward firing, inhibitory ones away; the currents of many synapses roughly add up at the soma
  • How much current one spike causes, the synaptic strength, is the main thing learning changes

From a neuron to a model neuron

The model neuron: inputs x1 to xD, each with its weight w1 to wD, summed with a bias b into z, passed through the nonlinearity g, giving the output a = g(z)
The neuron drawing of the earlier slide, all parts shown and labelled: dendrites, synapse, cell body (soma), axon hillock, axon, axon terminals

From a neuron to a model neuron

The model neuron: inputs x1 to xD, each with its weight w1 to wD, summed with a bias b into z, passed through the nonlinearity g, giving the output a = g(z)
  • Inputs xi\textcolor{#3E6D8E}{x_i}: the firing rates of the neurons that connect to it
  • Weights wi\textcolor{#B5396B}{w_i}: the strengths of those synapses, positive (excitatory) or negative (inhibitory) (minute paper)
  • Sum zz, with a bias b\textcolor{#3F7A6B}{b}: the currents adding up at the soma
  • Nonlinearity g\textcolor{#7A4FA3}{g} and output aa: the threshold at the axon hillock, and the firing rate
  • We drop spike timing, dendrites and chemistry: a point neuron

Building the neuron: inputs, weights and sum

The complete model neuron with the inputs x1 to xD, their weights w1 to wD, the bias b and the summation node in full colour, and the nonlinearity g and the output a faded. The equation reads z equals the sum over i of w i x i plus b
  • Each of the DD inputs xi\textcolor{#3E6D8E}{x_i} is a number: another neuron's output, or a stimulus feature such as a pixel
  • Each has a weight wi\textcolor{#B5396B}{w_i}, positive (excitatory) or negative (inhibitory). Their sum, z=∑iwi xi+bz = \sum_i \textcolor{#B5396B}{w_i}\, \textcolor{#3E6D8E}{x_i} + \textcolor{#3F7A6B}{b}, is the pre-activation
  • Learning changes the weights and the bias. First we set them by hand

The weighted sum is a dot product

Axes x1 and x2. The input vector x in blue and the weight vector w in magenta, both from the origin, at an angle phi; a dark band along w marks the projection of x onto w, of length w transpose x over the length of w
  • Stack the inputs and weights into vectors x, w∈RD\textcolor{#3E6D8E}{\mathbf{x}},\,\textcolor{#B5396B}{\mathbf{w}}\in\mathbb{R}^{D}: the weighted sum is their dot product

z=∑i=1Dwi xi+b=w⊤x+bz = \sum_{i=1}^{D} \textcolor{#B5396B}{w_i}\,\textcolor{#3E6D8E}{x_i} + \textcolor{#3F7A6B}{b} = \textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}} + \textcolor{#3F7A6B}{b}

  • Geometry (left): w⊤x=∥w∥ ∥x∥cos⁡ϕ\textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}} = \lVert\mathbf{w}\rVert\,\lVert\mathbf{x}\rVert\cos\phi, the length of the projection of x\textcolor{#3E6D8E}{\mathbf{x}} onto w\textcolor{#B5396B}{\mathbf{w}}, times ∥w∥\lVert\mathbf{w}\rVert
  • Aligned with w\textcolor{#B5396B}{\mathbf{w}} → large; perpendicular → zero; opposite → negative. The neuron asks how much of w\textcolor{#B5396B}{\mathbf{w}} is in the input

A neuron's weights set a decision boundary

The same axes, w and x as on the previous slide, plus the decision boundary, the gold line where w transpose x plus b equals zero, perpendicular to w. The side where z is positive is shaded and holds filled dots; hollow dots lie on the other side. The projection of x onto w passes the line, so x falls on the positive side
  • The neuron computes z=w⊤x+bz = \textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}} + \textcolor{#3F7A6B}{b}; its decision boundary is the line where z=0z = 0
  • w\textcolor{#B5396B}{\mathbf{w}} is perpendicular to the boundary and points to the side where z>0z > 0; b\textcolor{#3F7A6B}{b} slides the line without turning it
  • Add a threshold, and the neuron fires only on that side: one neuron is a linear classifier, the same readout that read context from the hypothetical three-neuron populations of the last two lectures (there, a plane in 3 units)
  • Every readout in this lecture is one model neuron

One weight vector, two readings

Four 48 by 48 pixel images in a row: the average face, the average non-face, the weights w of a readout trained to tell faces from other objects in Caltech-101 photographs, drawn with positive weights in magenta and negative in blue, and the average face minus the average non-face in the same colours. The weight image has a face-like layout of eyes, nose and mouth. Underneath, the readout's accuracy on held-out photographs
  • One neuron, one weight per pixel (2,304), trained to answer "face or not?"
  • As a boundary: it says "face" when w⊤x+b>0\textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}} + \textcolor{#3F7A6B}{b} > 0. As a template: drawn as an image, w\textcolor{#B5396B}{\mathbf{w}} has the layout of a face, and inputs that look like it give the largest w⊤x\textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}}
  • Magenta weights push toward "face", blue toward "non-face". The weights best separate the classes; they are not the average face
Haufe et al. NeuroImage 2014 · Photographs: Caltech-101, Fei-Fei, Fergus & Perona 2004

A neuron is a feature detector: template matching

One unit of AlexNet's first layer: its weights shown as a small colour image of an oriented edge. Three inputs of the same length: the same pattern, the pattern rotated by 90 degrees, and a different pattern; under each, the unit's response relative to the first: 1.00 for the same pattern, 0.00 rotated, 0.02 for the different pattern. Right, the tuning curve: the response as the pattern is rotated from 0 to 180 degrees, 1 at 0 degrees, falling to 0 by about 45 degrees, 0 at 90, with a small bump near 160
  • A unit of the network you analyse in the assignments: its weights w\textcolor{#B5396B}{\mathbf{w}}, drawn as an image, look like the pattern it detects; rotating the input traces its tuning curve (right)
  • w\textcolor{#B5396B}{\mathbf{w}} and every input x\textcolor{#3E6D8E}{\mathbf{x}} have length 1, so w⊤x\textcolor{#B5396B}{\mathbf{w}}^\top\textcolor{#3E6D8E}{\mathbf{x}} is the cosine of their angle: 1 when x\textcolor{#3E6D8E}{\mathbf{x}} looks like w\textcolor{#B5396B}{\mathbf{w}}. The weights act as a template

Building the neuron: activation and output

The same model neuron with the inputs, weights, bias and sum faded, and z, the nonlinearity g and the output a in full colour. The equation reads a equals g of z
  • a=g(z)a = \textcolor{#7A4FA3}{g}(z): the nonlinearity, most often ReLU, g(z)=max⁡(0,z)\textcolor{#7A4FA3}{g}(z)=\max(0,z)
  • The output aa is the neuron's activation, its stand-in for a firing rate
  • Why nonlinear? A real neuron fires nothing below threshold, then faster as its input grows

Transfer functions: the nonlinearity gg

Five transfer-function curves on one set of axes: ReLU (thick), leaky ReLU, sigmoid, tanh and GELU
  • A real neuron's f–I curve (firing rate against input current): 0 below threshold, then rising, then saturating
  • History: a step (McCulloch & Pitts 1943, the perceptron), then sigmoid and tanh, smooth enough for backpropagation (1986), then ReLU (2010s), then GELU in transformers
  • None is exact: ReLU and GELU never saturate, the sigmoid never reaches 0, tanh, leaky ReLU and GELU give negative "rates"
  • ReLU (thick line) dominates: cheap, and deep networks train well with it

From weights set by hand to weights learned

The complete single neuron: weighted inputs summed with a bias into z, z passed through the nonlinearity g, giving the output a, labelled a = g(z)
  • So far we chose w\textcolor{#B5396B}{\mathbf{w}} and b\textcolor{#3F7A6B}{b} by hand. A cortical neuron has thousands of synapses, a network millions, even billions, of weights: the weights must be learned from examples
  • Learning needs a loss, one number that measures how wrong the outputs are, and gradient descent, which changes every weight to lower it
  • The nonlinearity g\textcolor{#7A4FA3}{g} sets the kind of output, each with its own loss: a threshold gives yes or no (perceptron), no nonlinearity (g(z)=z\textcolor{#7A4FA3}{g}(z)=z) a number (linear regression), a sigmoid a probability (logistic regression), a softmax a probability per class

A network maps inputs to outputs

A diagram from left to right: an input x goes into a network drawn as layers of units with weights w, which produces an output y hat. The output is compared with a target y; the comparison gives the loss, and a dashed arrow labelled adjust w runs back into the network. Caption: inputs and targets decide which mapping the network learns
  • Training adjusts the weights so that, on many examples, the output matches the target. To train a network, choose inputs and targets

Where does the target come from?

The same diagram for supervised learning: the input is an image of a kitten, the output is class scores for cat, dog and car, and the target is the label cat, supplied by a person
  • Image → class label. The target, "cat", comes from a person: supervised learning

Where does the target come from?

The same diagram for a language model: the input is the words the kitten sat on the, the output is a predicted word, rug, and the target is the next word in the text, mat
  • The words so far → the next word. The text supplies the target: no one labels it

Where does the target come from?

The same diagram for learning from two views: the input is one crop of a kitten photograph, the output is a response pattern, and the target is the network's response to a second crop of the same photograph
  • One crop of a photograph → the response to another crop. The second crop supplies the target: self-supervised learning (SSL), like next-word prediction

Where does the target come from?

The same diagram for images and captions: the input is an image of a kitten, the output is an image embedding, and the target is the embedding of its caption, a kitten, computed by a text network; captions written by people for the web
  • Image → its caption, written by people for the web. An image network and a text network turn each into an embedding (a vector of numbers), and the two vectors are compared. The inputs and targets we choose decide which mapping the network learns

How can a readout learn its weights from labeled examples?

The first learning neuron: the perceptron

Three runs of the perceptron rule without a bias, on the same 2-D points separable by a line through the origin, each from a different starting weight vector. In each panel the decision line after every update is drawn from faint to dark, ending in a solid gold line that separates the classes, with the starting and final weight vectors as magenta arrows; under each panel, the mistakes per pass fall to zero
  • Rosenblatt (1958): a model neuron with a threshold that learns its weights from labeled examples, built in hardware
  • The rule is a heuristic: after a mistake, move w\textcolor{#B5396B}{\mathbf{w}} toward the missed example, or away from a false alarm. No loss lies behind it, yet if a line separates the classes, it stops (all three runs)
  • It learns from its mistakes: w\textcolor{#B5396B}{\mathbf{w}} changes only when the neuron is wrong

Which way is downhill? From derivative to gradient

Two panels. Top, one weight: a U-shaped loss curve over w with a dot at w = 0.5, its tangent of slope minus 3, an arrow pointing toward the minimum, labelled step against the slope, and the minimum marked slope 0. Bottom, two weights: elliptical contours of the loss over w1 and w2 with the minimum marked; at a starting point, a gray arrow perpendicular to the contour labelled gradient, uphill, and a magenta arrow in the opposite direction labelled minus the gradient
  • The loss ℓ(w)\ell(\textcolor{#B5396B}{w}) is a function of the weight: one number saying how wrong the outputs are
  • One weight: the derivative dℓ/dwd\ell/d\textcolor{#B5396B}{w} is its slope. Slope negative: increase w\textcolor{#B5396B}{w}; positive: decrease it. At the minimum the slope is 0
  • Many weights: the gradient ∇ℓ\nabla\ell lists one slope per weight, ∂ℓ/∂wi\partial\ell/\partial \textcolor{#B5396B}{w_i}: how fast ℓ\ell grows when wi\textcolor{#B5396B}{w_i} alone grows a little
  • The gradient points uphill. Gradient descent: a small step against it, w←w−η ∇ℓ\textcolor{#B5396B}{\mathbf{w}} \leftarrow \textcolor{#B5396B}{\mathbf{w}} - \eta\,\nabla\ell, again and again; the learning rate η\eta sets the step size

The delta rule (Widrow–Hoff): step down a smooth error

Contours of the loss for a two-weight linear regression, with the best weights marked by a star. From the start at the upper left, the path of the delta rule with a constant learning rate, one example per step, winds downhill toward the star; a short arrow at the start shows the direction minus the gradient

Widrow & Hoff (1960): write a loss and step down its slope. Drop the threshold so the loss is smooth: the output is y^=w⊤x\hat y = \mathbf{w}^\top\mathbf{x}.

ℓ=12 (y−y^)2,∂ℓ∂w=−(y−y^) x\ell = \tfrac12\,(y - \hat y)^2, \qquad \frac{\partial \ell}{\partial \mathbf{w}} = -(y - \hat y)\,\mathbf{x}

  Δw=−η ∂ℓ∂w=η (y−y^) x  \boxed{\;\Delta\mathbf{w} = -\eta\,\frac{\partial \ell}{\partial \mathbf{w}} = \eta\,(y - \hat y)\,\mathbf{x}\;}

  • The perceptron is a special case: the same rule with the thresholded output as y^\hat y, so its error is 0 unless the neuron is wrong
  • One recipe: change each weight by η\eta × error × input

Try it: perceptron and delta rule, one example at a time

  • Perceptron: correct, nothing changes; a mistake, w\textcolor{#B5396B}{\mathbf{w}} moves. Delta rule: every example moves w\textcolor{#B5396B}{\mathbf{w}} a little. Watch the loss after each update. Play with it after class

Logistic regression: a probability, then its error

p^=σ(z)=11+e−z\textcolor{#3F7A6B}{\hat p} = \textcolor{#7A4FA3}{\sigma}(z) = \frac{1}{1+e^{-z}}

ℓ=−[ ylog⁡p^+(1−y)log⁡(1−p^) ]\ell = -\big[\,y\log \textcolor{#3F7A6B}{\hat p} + (1-y)\log(1-\textcolor{#3F7A6B}{\hat p})\,\big]

  • Labels coded y=1y=1 or 00. The sigmoid converts the logit z=w⊤xz = \mathbf{w}^\top\mathbf{x} into p^\textcolor{#3F7A6B}{\hat p}, the probability of class 1
  • Minimize the cross-entropy ℓ\ell, minus the log of the probability given to the correct class: 0 if certain and right, 0.69 for a 50/50 guess, 2.30 if confident and wrong
  • Gradient ∂ℓ/∂w=(p^−y) x\partial \ell/\partial \mathbf{w} = (\textcolor{#3F7A6B}{\hat p} - y)\,\mathbf{x}, so Δw=η (y−p^) x\Delta\mathbf{w} = \eta\,(y - \textcolor{#3F7A6B}{\hat p})\,\mathbf{x}
  • Same recipe, another loss: again η\eta × error × input

Two stacked panels. Top: the sigmoid curve, probability p-hat rising from 0 to 1 as the logit z goes from negative to positive, crossing 0.5 at z equals 0. Bottom: cross-entropy against p-hat; for a true class of 1 the loss falls from above 5 to 0 as p-hat goes to 1, and for a true class of 0 it rises the other way

Further reading for the research trail

Optional: a few research-trail entries over the semester. Any paper mentioned in any lecture qualifies, but papers from this block are preferred.

Minute paper

Submit on Canvas → Minute papers → Minute Paper 9 (access code read out in class).

Write three brief points in your own words:

  1. In the model neuron, what does a weight wiw_i stand for in a real neuron, and what does a negative weight mean?
  2. Something you do not yet understand, or a question still open
  3. Another idea you found interesting, and why it matters for brains, behavior or AI

Credit for a thoughtful attempt, not for being correct.

- The readouts fitted in the geometry lectures and in the assignment, and every unit of the networks whose activations the first assignment analysed, compute the same thing: a weighted sum of their inputs, then a nonlinearity - Until now a library fitted the weights for us. This lecture opens the box: first the single neuron and the real neuron it abstracts, then how its weights are learned from examples. How a whole network of them is trained by backpropagation

- A spike is a brief electrical pulse, about a millisecond long, that travels along an axon. Other neurons' spikes arrive at this neuron's synapses, mostly on its dendrites - At each synapse, an arriving spike lets in a small current (next slide). A cortical neuron receives about 10,000 synapses; one input spike changes its voltage by roughly 0.1 to 1 mV, and firing needs about 15 mV above rest, so many inputs have to arrive close together - The currents spread to the soma and add up, only roughly: dendrites can compute on their own (London & Häusser 2005), and inhibition near the soma partly divides the input rather than subtracting from it - When the voltage at the axon hillock, the start of the axon, crosses threshold, voltage-gated sodium channels open and the neuron fires a spike, which travels down the axon to the next neurons' synapses (Hodgkin & Huxley 1952 described how) - The output the model keeps is the firing rate, the number of spikes per second; it discards when each spike happens - The neuroscience background notes on the course site (Recitations) cover these parts, spikes, synapses and ion channels in four pages, and map each onto the model neuron. Other short, non-technical introductions: OpenStax Psychology 2e, section 3.2 "Cells of the Nervous System"; OpenStax Biology 2e, sections 35.1 "Neurons and Glial Cells" and 35.2 "How Neurons Communicate"; Neuroscience Online (UTHealth), chapters 1 and 6; Gerstner et al., Neuronal Dynamics, section 1.1, the closest bridge to the model neuron

- The cell body and the axon hillock: where the inputs are summed and the spike starts

- The axon and its terminals: the spike travels to the next neurons. The notes of the first frame of this slide have the details

- The synapse produces a current; the small voltage change that current causes is called a postsynaptic potential (PSP). It is graded, not a spike - The neurotransmitter transporters in the drawing do reuptake: they pump transmitter out of the cleft, back into the axon terminal (or into nearby glial cells), which ends the signal and recycles the transmitter. Many antidepressants (SSRIs) block the serotonin transporter - Each neuron's synapses onto others are all excitatory or all inhibitory (Dale's principle): the sign belongs to the presynaptic cell - Learning changes synaptic strengths, and also how excitable a neuron is; the model keeps the first as its weights and the second, roughly, as its bias

- Each part of the model stands for one part of the neuron: the inputs and weights for the synapses, the sum for the soma, g for the spike threshold at the axon hillock, a for the firing rate. The model keeps one number per input and one number out, the firing rates, and discards when the spikes happen - Weights can change sign during learning; a real synapse keeps its sign. The bias has no single anatomical counterpart: it stands for how excitable the neuron is, how far its resting voltage sits below threshold, and any steady background input - Real rates are never negative; model inputs can be, once data are centred or z-scored - The abstraction is that of McCulloch & Pitts (1943). The next two slides build it in that order: weighted inputs and their sum, then the nonlinearity

- The same correspondences, one by one. Real rates are never negative; model inputs can be, once data are centred or z-scored - Minute paper 9 asks about this slide

- The indexing convention: x_1 to x_D, with x_i a generic input and D the number of inputs, the input dimension (the same D as the number of units in earlier lectures: a neuron's inputs are the units of the population it reads) - z is a weighted sum, not an average: the weights need not be positive or add up to one - The bias b shifts the sum up or down, and so sets how much input it takes to drive the neuron

- The sum and the dot product are the same number; the vector form is only shorter to write. Vectors are bold lowercase columns, and w is one unit's incoming weights - You met the dot product as a measure of alignment in the similarity lecture, and the projection onto a direction in the PCA lecture. A neuron computes that projection, then adds a bias and applies a nonlinearity - In the figure, the dark band is the projection of x onto the direction of w; its length is w^T x / |w|, and w^T x is that length times |w|. φ is the angle between x and w - The cosine form makes alignment literal: for inputs of a fixed length, the response depends only on the angle φ to w - The linear readout of the geometry lecture computed exactly this weighted sum and answered "yes" when it was positive. Here we give it a nonlinearity and read it as a neuron - w^T x is the one transpose in the course; a whole layer is written z^[1] = W^[1] x + b^[1] (next lecture) with no transpose, as in PyTorch's nn.Linear

- The readouts of the first part (decoding, shattering, CCGP) all fit a w and a b. The gold planes in the figures of hypothetical three-neuron populations are their decision boundaries, in three units instead of two - Why w is perpendicular: moving along the line keeps w^T x constant, so every step along the line is perpendicular to w. Moving along w changes z fastest - The bias sets where the line sits: the line crosses the direction of w at a distance −b/|w| from the origin - The next slide reads the same weights a second way, as a template

- One model neuron: 2,304 inputs (the pixels), one weight per pixel, a bias, and a sigmoid (the S-shaped nonlinearity of a later slide). Training it is logistic regression, which this lecture derives shortly - The photographs come from Caltech-101, a public dataset: its "Faces_easy" category was cropped and roughly centred by the dataset's authors, so no alignment was done here. Against the 435 faces, 435 photographs drawn at random from the other categories; all in grayscale, shrunk to 48 × 48 pixels - Both readings describe this same single neuron. The boundary reading asks which side of the line an input falls on; the template reading asks which input drives the neuron most. The next slide reads a unit inside a trained network the second way - The readout is right on 95.4% of held-out photographs. Its weights resemble the difference between the average face and the average non-face (correlation 0.73 across pixels), but they are not the same: the readout also learns to give little weight to pixels that vary a lot within each class, such as the background - A caution that matters for decoding brain data (Haufe et al. 2014): a decoder's weights are a filter, not a picture of the signal. A voxel or neuron can get a large weight because it helps cancel noise, not because it responds to the class

- With w and x both of length 1, wᵀx is the cosine of the angle between them; the ReLU nonlinearity, max(0, z), next slides, clips the negative ones to 0 - The tuning curve: the response as the input pattern is rotated from 0° to 180°. A later lecture measures such curves in real neurons - A two-pixel image and a two-weight neuron: set w to the bright pattern and the neuron responds most to that pattern - The comparison is fair only for inputs of the same length, since w^T x = |w| |x| cos θ; a brighter input raises the response too - The same weights have two readings, as the face readout showed: a decision boundary (a classifier) and the input pattern that drives the neuron most (a template). For a unit inside a network the template reading is the useful one; it comes back when we measure tuning in neurons and in networks

- ReLU passes positive inputs unchanged and sets negative ones to zero: a threshold at zero, so with ReLU the neuron is active only when z > 0, that is when the weighted sum exceeds −b - The activation a is one number. With ReLU it is never negative, like a firing rate - The whole neuron in one line: a = g(Σ_i w_i x_i + b), a weighted sum, a bias, a nonlinearity. Most networks in this course are built from many of these units

- g is a design choice. Inject a steady current into a neuron and it fires nothing below a threshold, then fires faster the stronger the current, and levels off because after each spike it cannot fire again for a millisecond or two (the refractory period). Each function keeps part of that curve - Step: McCulloch & Pitts (1943) and Rosenblatt's perceptron (1958) used an all-or-none output, 1 above threshold and 0 below. It captures the threshold, but its slope is zero everywhere, so gradient descent cannot use it - Sigmoid, from 0 to 1, and tanh, from −1 to 1: smooth S-shaped curves that made backpropagation possible (Rumelhart, Hinton & Williams 1986); tanh, centred on 0, was the usual choice in the 1990s (LeCun et al. 1998). The sigmoid is the most rate-like of the five: zero-ish below threshold, then rising and saturating. Its weakness is the flat ends: where the curve is flat, the gradient that learning follows is nearly zero, and deep stacks of sigmoids learn very slowly. This is the vanishing-gradient problem, which comes back in a later lecture - ReLU, max(0, z): the threshold-linear unit. Hahnloser et al. (2000) used it to model cortical circuits; Nair & Hinton (2010) and Glorot, Bordes & Bengio (2011) showed that it makes deep networks much easier to train, and AlexNet (2012) used it. Biologically it keeps the threshold and the rise, not the ceiling - Leaky ReLU (Maas et al. 2013) lets a small negative output through so that a unit never stops learning entirely. GELU (Hendrycks & Gimpel 2016) is a smooth ReLU with a small dip below zero; BERT and the GPT models use it - Negative outputs (tanh, leaky ReLU, GELU's dip) have no firing-rate counterpart unless the output is read as a change from a baseline rate - Without g, the neuron is linear: its output is the weighted sum itself. A weighted sum followed by a threshold-like g is the linear–nonlinear (LN) model, a common first model of a cell

- The rest of the lecture is about where the weights come from, for one neuron, a readout. The hidden layers of a network come next lecture - The four kinds of output differ only in the last step: what g is, and what the output is compared with. A yes/no output is compared with a label, a number with a measured value, a probability with the label, 1 or 0, and a set of class scores with the correct class. Each comparison gives its own loss, and the same downhill procedure, gradient descent, lowers any of them - Every weight learned today is judged the way the last lecture judged a score: on held-out images, against shuffled labels. With many units and few images, a readout fits even shuffled labels, so keep the generalization gap in mind whenever a training loss looks good

- The model neuron of the first part is the smallest such mapping: D numbers in, one number out. A network chains many of them - The next slides show four common choices of inputs and targets. The first needs people to label every example; the other three take their targets from the data themselves

- In every frame the training loop is the same: compute the output, compare it with a target, change the weights to shrink the difference. Only the source of the target changes. Three of the four need no hand-assigned labels - Labels: training lowers the cross-entropy (defined later in this lecture), which raises the probability given to the correct class. Someone had to label every image: ImageNet's 1.2 million photographs were labeled by hand, while text and unlabeled photographs are nearly unlimited - Next word: a language model predicts the next word (strictly, the next token, a word or piece of a word) from the words before it; the objective behind today's chatbots (Brown et al. 2020, GPT-3). BERT (Devlin et al. 2019) hides about 15% of the words and predicts them from both sides - Two views: DINO (Caron et al. 2021) trains a student network to match a teacher network (a slowly updated average of the student) on two crops of one image with changed colours. Tricks in the loss stop both from giving one output for every image. Masked autoencoders (He et al. 2022) hide about three quarters of an image's patches and reconstruct them - Image and caption: CLIP (Radford et al. 2021) trains an image network and a text network so that each image matches its own caption better than the other captions in the batch; about 400 million pairs from the web. People wrote the captions, but not to train a network

- A language model predicts the next word (strictly, the next token) from the words before it; the objective behind today's chatbots (Brown et al. 2020). No one labels the text

- Self-supervised learning (SSL): the data supply their own targets, with no labels from people. Next-word prediction (the previous slide) is self-supervised too - DINO (Caron et al. 2021) trains a network so that two crops of one photograph, with changed colours, give similar responses. No labels are used

- CLIP (Radford et al. 2021) trains an image network and a text network so that each photograph matches its own caption better than the other captions; about 400 million pairs from the web - Networks trained these four ways differ in how closely they match people's odd-one-out choices on the THINGS images (Muttenthaler et al. 2023). You test this yourself in the assignment

- The readout of the geometry lecture: a weighted sum of the responses and a threshold. Until now we fit it with a library call; here is what the call does - One neuron, a label for every input, and a rule that changes the weights. Three rules follow, and all have one form: change each weight by $\eta$ × error × input

- The perceptron is the model neuron from the start of the lecture, a = g(wᵀx + b), with a threshold for g and a rule that sets w from labeled examples: on a mistake, add ηyx to w (labels coded +1 and −1; η is a small step size); otherwise do nothing - Novikoff (1962): if a line separates the classes with some gap, the number of mistakes is bounded. The rule stops at the first line that works; which line depends on the starting weights and the order of the examples - If no line separates the classes, as for XOR, the rule never stops - Minsky and Papert (1969) proved that a single layer cannot compute XOR, and other limits like it. The perceptron is still the basis for understanding supervised learning; since then, the learning rule comes from an explicit loss, lowered by gradient descent, in many stacked layers. Hidden layers trained by backpropagation, the second half of this lecture, came in the 1980s - Rosenblatt's book-length account: Principles of Neurodynamics (1962) - The Mark I Perceptron (designed from 1958, publicly demonstrated 1960) had a 20 × 20 grid of photocells as input and motor-driven potentiometers as weights - After Minsky and Papert's book, interest and funding for neural networks fell for more than a decade - The figure: the same simulated points and order of examples, three different starting weights, no bias (the classes are separable by a line through the origin); each faint line is the boundary after one update, and the mistakes per pass are printed under each run. Compare the XOR figure later in the lecture, where the mistakes never reach zero

- The top curve is a made-up loss, ℓ(w) = (w − 2)² + 0.5, chosen so the minimum is easy to see; a real loss is computed from the training examples, but it is still a function of the weights - The gradient generalizes the derivative from one variable to many. In calculus you found the minimum of a function by setting its derivative to 0; with millions of weights there is no formula to solve, so we walk downhill instead - The partial derivative ∂ℓ/∂w_i is the ordinary derivative with every other weight held fixed - The gradient points in the direction in which the loss grows fastest, and it is perpendicular to the contour lines of the loss. Its opposite is the steepest way down - η, the learning rate, sets the size of each step. Too large and the steps overshoot; too small and training is slow. A later part of this lecture returns to it - If derivatives are new to you: the Recitation 4 handout (gradients and the chain rule, course site, Recitations) starts from the slope of a line

- The perceptron's rule was a heuristic: Rosenblatt did not derive it from a loss, and its error is either 0 or not, with no slope to follow. Widrow and Hoff (1960) compared the target with the output before the threshold, ŷ = wᵀx, which makes the error smooth, and derived the rule as gradient descent on that squared error. The result has the perceptron's form, η × error × input, which is why the delta rule is often presented as its generalization - A loss is the number a learning rule tries to make small: here the squared error ℓ of one example. Averaged over the training examples it is the loss ℒ of the whole training set; the course writes ℓ for one example and ℒ for the average - Stepping against the gradient, one example at a time, is stochastic gradient descent - Sign convention: the gradient is (ŷ − y)x; the update is minus the gradient, (y − ŷ)x. Books use both forms - Names: the delta rule, the Widrow–Hoff rule, the least-mean-squares (LMS) rule - Unlike the perceptron, it updates on every example, even correct ones: it aims at the value y = ±1, not only the correct side. With the threshold put back, (y − ŷ) is 0 on correct examples and ±2 on mistakes, which is the perceptron's update - With a continuous target, such as a firing rate, the same rule is linear regression: a model of a neuron's response, in the next lecture

- The demo is on the course site with the other demos. Start from w = 0: the first example is always a mistake - Watch the number of mistakes per pass fall to zero when a line separates the classes. Add flipped labels and it never does: the perceptron keeps moving - The menu switches to the delta rule and logistic regression, which update on every example by how wrong it is, the rules of the next two slides

- The delta rule aims at y = ±1 even for points far on the correct side. Logistic regression outputs a probability instead, and its error for a clear case is near zero - Despite the name, a classifier: one neuron whose nonlinearity is the sigmoid σ of the transfer-function slide - The cross-entropy is minus the log of the probability given to the correct class: −log p̂ when y = 1, −log(1 − p̂) when y = 0 (bottom panel). At p̂ = 0.5 it is 0.69; a confident mistake, p̂ = 0.01 for a true class of 1, costs 4.61 - Three updates: perceptron ηyx on mistakes only (y = ±1); delta rule η(y − ŷ)x on every example; logistic regression η(y − p̂)x on every example (y = 1 or 0) - The gradient is this simple because the slope of the sigmoid cancels against the slope of the log. The update is the delta rule's with p̂ in place of ŷ - Averaged over the training set, the cross-entropy is the loss ℒ. Confident and wrong means, for example, p̂ = 0.1 when the true class is 1

- Rosenblatt (1958), The perceptron: a probabilistic model for information storage and organization in the brain. The original proposal, written for psychologists - Zipser & Andersen (1988), A back-propagation programmed network that simulates response properties of a subset of posterior parietal neurons. The hidden units of a small trained network, compared with recorded neurons: the template for many later model–brain comparisons - Lehky & Sejnowski (1988), Network model of shape-from-shading: neural function arises from both receptive and projective fields. A neuron's receptive field alone may not reveal what it computes - Lillicrap, Santoro, Marris, Akerman & Hinton (2020), Backpropagation and the brain. A readable review of how error signals might be carried in cortex; the companion Lillicrap et al. (2016, Nat Commun) shows that random feedback weights can replace the transposed weights of step 3 - Muttenthaler, Dippel, Linhardt, Vandermeulen & Kornblith (2023), Human alignment of neural network representations - Francioni, Tang, Toloza, Ding, Brown & Harnett (2026), Vectorized instructive signals in cortical dendrites - For the mechanics, the 3Blue1Brown videos on neural networks (chapters 3 and 4) animate the four steps of backprop - Before posting, skim the existing entries on your paper; your post must add something not already said

- Submitted on Canvas (Minute papers → Minute Paper 9); credit for a thoughtful attempt