Brown University crest
CPSY 1291 Computational Methods for Mind, Brain & Behavior
Lecture 0 · Welcome

What this course is about

Thomas Serre · Professor of Cognitive & Psychological Sciences
Carney Institute for Brain Science, Brown University

Thursday, September 10 · Fall 2026

Reader-friendly version — lecture content with figure descriptions

QR code that opens the course Canvas page

Scan for Canvas
canvas.brown.edu/courses/1103531

Animal, or no animal?

What I spent my PhD building

Performance in d-prime for the model and for human observers across four viewing distances, head, close-body, medium-body and far-body: the two curves lie on top of each other within error bars The stimulus set: animals photographed at four viewing distances, head, close-body, medium-body and far-body, each with a matched distractor scene
HMAX: the ventral visual hierarchy mapped onto model layers, from oriented filters through alternating simple and complex cell stages to an animal versus non-animal decision

Where we are today: vision

30%

of 350+ vision models now land inside the human accuracy range — and about 1% land above it. Not the binary task you just did: 1,000 categories, 1.4 million photographs, no time limit.

Five phone photographs of a room are now enough to rebuild it as a navigable 3D world.

QR code linking to the World Labs Marble post, which shows single photographs turned into navigable 3D scenes
Two ordinary phone photographs of the same office room with a ping-pong table, taken from different angles, stacked one above the other

Where we are today: language

73%

In a controlled three-party Turing test, a persona-prompted GPT-4.5 was picked as the human more often than the real person was — five-minute conversations, the same judge talking to both. Without the persona prompt: about 36%.

A judge holds two conversations at once, one with a person and one with a machine, and has to say which is which. That is the Turing test, proposed by Turing in 1950.

The three-party Turing test: a judge on the left holds two simultaneous conversations, one with a person and one with a machine, and must say which is which

Where we are today: mathematics

125

pages of proof, generated from the problem statement alone, disproving a conjecture Erdős posed in 1946. OpenAI, May 2026: the unit-distance problem.

  • Dec 2025: DeepMind's Aletheia solves four Erdős problems on its own
  • Sept 2026: OpenAI reports cracking Navier–Stokes, a Millennium Prize problem: 10,000 agents, 88 hours. Contested, and not yet through peer review
  • Most of the new results are coming from hobbyists and undergraduates using models anyone can access

Where we are today: code

17,000

automated steps over one weekend — an AI being tested for safety broke out of the isolated space it was being tested in, and walked into Hugging Face, the site where the world keeps its AI models.

  • In July 2026 it found a flaw nobody knew about, used it to escape, spread across the company's internal machines, and took the answers to the test it was sitting
  • The first break-in of its kind carried out by AI agents rather than people
  • The models and datasets the public downloads were checked and found untouched

So do these models have anything to do with the brain?

Same accuracy ≠ same strategy

We can measure where a person looks to recognise an object, and where a model looks, on the same images.

They do not agree.

QR code linking to the ClickMe game, where players reveal the parts of a photograph that let someone name the object

Play ClickMe: the human maps come from people playing it.

Five rows — snake, bear, hare, leopard and ball. The left pair is the photograph and the human ClickMe importance map, concentrated on the head or face. The right six columns are the importance maps of ViT, MLP-Mixer, ResNet50, SimCLR, Robust ResNet50 and ConvNext, which scatter across body and background instead

Texture, not shape

Three panels: a photograph of a ginger cat labelled shape cat; a close-up of grey wrinkled elephant skin labelled texture elephant; and the two combined, a cat silhouette carrying elephant skin, labelled cue conflict

People say cat. A standard ImageNet CNN says elephant — it weights texture far above shape, the opposite of the human bias.

Geirhos et al. ICLR 2019 · stimuli CC BY 4.0

Neural alignment

The most accurate model is not the best model of monkey IT.

Scatter of hundreds of models: ImageNet multi-label accuracy on the horizontal axis against neural alignment with monkey inferotemporal cortex on the vertical axis. The trend rises and then turns down, so the most accurate models predict IT worse
Work from this lab · Linsley, Feng & Serre TICS 2026

Behavioural alignment

Accurate models rely on different features than we do.

The same scatter of hundreds of models, now against alignment with human feature-importance maps. Past a point, more accurate models agree less with where people look
Work from this lab · Linsley, Feng & Serre TICS 2026

When diverging is the point

Four synthetic feature-visualization images, each a dense swirl of leaf-like shapes and vein networks in pale blues, pinks and oranges, showing the prototypical venation pattern one learned concept responds to most strongly

Identifying an isolated fossil leaf defeats almost every human expert. Most sit unidentified in museum drawers.

A model reaches 93.2%, and gives credible families for 85.6% of 1,177 unidentified specimens.

Attribution caught it cheating first: it was reading the labels on the slides instead of the leaf.

QR code linking to the interactive FossilLeafLens explorer

Explore the concepts and attribution maps yourself: the tools you build in Assignment 4.

FossilLeafLens · LeafLens · live app · Rodriguez* Fel* et al., under review

The two-way street

Yin-yang symbol whose two halves are labelled Neuro and AI

NeuroAI studies where artificial and biological systems compute alike, where they diverge, and which biological constraints close the gap.

  • Neuroscience → AI. Architectural and developmental constraints — cortical feedback and recurrence, a temporally continuous training signal, learning objectives other than classification — improve robustness and alignment
  • AI → neuroscience. Recordings from populations of neurons now reach thousands of cells at once; machine learning supplies the encoding and decoding models that make them interpretable

Both directions require models specified precisely enough to run as programs: a theory you can execute, measure and falsify.

A different visual diet

A child · SmartPlayroom, Brown

A newborn chick · controlled rearing

Chick environment: Ashok et al. · SmartPlayroom · Amso & Serre, Brown · SAYCam Sullivan et al. Open Mind 2021

How we got here

NeuroAI at Brown

You are in one of the places where this happens

  • Brown's Carney Institute for Brain Science houses the Nancy G. Zimmerman Center for Computational Brain Science, which brings together people who build models of the brain and people who build tools to analyze brain data
  • Measured against peer institutions, Brown's distinctive strength is AI and machine learning for neuroscience, in both directions
The Carney Institute innovation hub: an open floor with students working together at tables with laptops, someone writing on a whiteboard beside a large screen, and glass-walled meeting rooms along one side
The Carney Institute innovation hub, 164 Angell Street

The other face of AI: ARIA

  • Brown leads ARIA (AI Research Institute on Interaction for AI Assistants), a new National Science Foundation AI institute
  • $20M over five years (2025–2030); institute PI: Brown's Ellie Pavlick
  • The goal is trustworthy AI assistants — systems whose advice you can rely on — with mental and behavioral health as the high-stakes test case
  • The team is roughly half computer science, half CLPS / Carney, plus ethicists and clinicians
  • Public-interest science rather than a product race

Where people go from here

Recent graduate-school placements: Stanford, MIT, CMU, NYU, and some stay at Brown. On the industry side:

Who Left Brown Path
David Mély 2016 Serre lab PhD → Vicarious → Google [X] → OpenAI (credited on o1)
Junkyung Kim 2019 Serre lab PhD → Google DeepMind
Thomas Fel 2024 Serre lab PhD → Harvard Kempner Institute → Goodfire
Jack Merullo 2025 Brown CS PhD → Goodfire
Siddharth Boppana undergrad took this course, TA'd it → Goodfire

This path starts here. No linear algebra or PyTorch required. Come talk to me.

Course organization

The teaching team

Thomas Serre, instructor
Thomas Serre
instructor
Peisen Zhou, head teaching assistant
Peisen Zhou
head TA
Lyfey Vutha, teaching assistant
Lyfey Vutha
TA
Tekin Gunasar, teaching assistant
Tekin Gunasar
TA

Prerequisites

  • At least one programming course
  • Prior Python or PyTorch is not required
  • Prior neuroscience or cognitive science is not required
  • Linear algebra and statistics are recommended, not required
  • Optional recitations, run by the TAs: Python and PyTorch bootcamps, a math review, and preparation for each assignment

What this course prepares you to do

breadth → skills → a question of your own

Lectures: recurrence, attention, language, generative models, and the arguments the field is having now

Assignments: you build a map of mental space, a model neuron, a learning rule, a network and a brain-similarity score yourself

Project: you choose the question and answer it

The semester in five themes

Italics = you implement it yourself. The rest we cover in lecture.

1 · Representational spaces
distances · MDS · PCA · t-SNE · UMAP · neural geometry · manifolds

2 · Learning, neuron to network
model neuron · Hebb & Oja · perceptron · backprop · double descent · sparse coding · autoencoders

3 · Vision & model–brain comparison
CNN · RSA · brain-similarity score · attribution · feature visualization · shortcut learning

4 · Sequence models & transformers
RNN · attractor · self-attention · LLMs · scaling laws · self-supervised learning

5 · Generative models & frontiers
VAEs · GANs · diffusion · world models · whatever your project needs

What an assignment actually looks like

A1 — a map of mental space

Two panels. On the left, 120 animal photographs placed by multidimensional scaling of human similarity judgments, colored by taxonomic group, with birds, reptiles, primates, rodents and hoofed mammals landing in separate regions. On the right, the same map with each point replaced by its photograph

Human judgments of how similar 120 animal photographs are, turned into a map: animals nobody labelled land next to the animals people find similar.

A1 — the same two directions in cortex

From the model — object space, built from network features

Object images arranged on a two-dimensional plane: the horizontal axis runs stubby to spiky, the vertical axis animate to inanimate, with faces, animals, tools and boxes falling in different quadrants

From the brain — recorded IT cells, same two directions

Scatter of recorded cells on the same two principal components, colored by which of four inferotemporal network patches they came from: the body, face, stubby and NML patches each occupy a different quadrant

Two directions: spiky to stubby and animate to inanimate. Four IT patches each occupy a different quadrant. Assignment 1 asks whether a deep network's space has the same two.

The McGill on that paper is Mason McGill — a Brown undergraduate who took this course, and later TA'd it.

The project: reading now, building in November

  • Research trail (from week 2): four times this semester, pick one citation from any lecture so far, read the paper, and post four to six sentences on Ed. Best three of four count; first one due Fri 9/25
  • Late November: submit a one-page team proposal; it may grow out of a trail entry, but it need not
  • Last 2.5 weeks: build it in teams of two (occasionally three); grades are individual
  • Mon 12/21, 9:00 AM: public poster session; two-page report due the same morning

Five shapes a project takes

  1. Reimplement a result
  2. Swap the model: does the claim survive?
  3. Add the control the paper skipped
  4. Take the method somewhere new
  5. Test a model's prediction on people
Project strand · 22% — trail 7% · proposal 3% · project 12%

Requirements & grading

Component Weight
4 programming assignments 40%
Project strand (trail 7 · proposal 3 · project 12) 22%
Minute papers (most lectures) 15%
Exam 1 (in-class) 9%
Exam 2 (in-class) 14%

Minute papers are retrieval practice, not a quiz: a few minutes at the end of class, answering whatever I ask that day, e.g. three ideas you found interesting, or something you did not follow. Credit is for a genuine attempt, and pulling material back out of your own head is one of the best-evidenced ways to keep it.

Logistics & schedule

  • Lectures Tue/Thu 2:30–3:50, Salomon 003, slides posted before class
  • Optional recitations Tue 6:30–7:50, including Python bootcamps (room on Canvas)
  • Office hours Wed 1:00 PM, Carney 419: optional; email me before you come so I do not miss you · TA hours posted on Canvas
  • Exam 1 Thu 10/22 · Exam 2 Tue 12/1, both in class, on paper, one sheet of notes, both sides
  • Brown's official final-exam slot, Mon 12/21 at 9:00 AM, is the poster session, not Exam 2
  • Course materials cost $0: no textbook, and every required reading is free or open access

Tools: Python · PyTorch · Colab

  • Assignments are Python notebooks on Google Colab; some use PyTorch
  • Sign in to Colab with your Brown account; datasets are shared to Brown accounts
  • No TensorFlow, no JAX
  • Late work: each late day costs 1 final-grade point, capped at the assignment's 10-point weight; completion is still required
  • Collaboration: discussion is encouraged; submitted work must be your own; acknowledge sources and collaborators

AI and this course

Let's be honest about AI

  • Every university is scrambling to work out what AI means for teaching. There is no settled answer, and anyone claiming one is selling something
  • This course was itself partly built and debugged with AI
  • The answer we landed on: a rule for the assignments, and a taught skill for the project

How these assignments were built

  • Each assignment reproduces a recent published study with real data and a real result, so there are a lot of moving parts
  • I could not have built them at this scale without AI. By hand it would have taken a year I did not have
  • That is a claim about capability, not laziness: I can supervise this because I already know the material

One caveat. Everything is new this semester and we are still iterating, so expect the occasional bug or glitch. Please be patient, and tell me when you hit one. New assignments built on current papers beat polished ones built on outdated material.

Post it on Ed · you will find things, and I want to hear about all of them

The rule, and the reason

Assignments, exams, minute papers, research trail — no AI tools.
Final project — AI welcome, with disclosure.

The reason is not tradition:

  • You can only supervise what you can verify, and you can only verify AI's code if you own the basics: Python, the algebra, PyTorch
  • The assignments are the basics; outsourcing them removes the one thing that would make AI useful to you later
  • Spellcheck, web search and ordinary editor completion are fine. AI-generated code, derivations, prose and paper summaries are not

Assignments contain integrity markers, and the course may use additional methods to identify AI-assisted work.

Where you will use AI this semester

  • The final project: AI tools allowed for coding, writing and brainstorming, with a short disclosure naming the tools and how you used them
  • Final-project guidance teaches a professional workflow, the writer/critic pair. One AI session writes the code, a second, blinded session tries to break it, and you referee
  • The standard does not move: you must be able to explain and defend everything you submit. The teaching staff will not debug code you cannot explain

Own your code

There is a spectrum:

  • ❌ Vibe coding: prompt, paste, pray. You cannot explain it, you cannot fix it, and you will not know when it is wrong
  • ✅ Planned collaboration: you plan the analysis, decompose it, delegate pieces, and verify each one

The standard: the code must be intellectually yours — explainable and defensible, piece by piece.

The Ed answer bot

  • When you post a question on Ed, an AI assistant drafts a reply for the teaching staff
  • It reads only student-facing materials — the notebooks, the syllabus, the rubrics — and cites where each answer came from
  • A person approves, edits or discards every draft. Nothing is posted automatically
  • It will not give solution code, and it says so when the materials do not settle your question

Before you go: the minute paper

Three lines, on Canvas, before you leave. Today's is not graded: people are still shopping, and this one is a survey for me and a dry run for you.

  1. What do you most want to be able to do by the end of this term?
  2. What are you most worried about in this course?
  3. What brought you here: a class, a paper, a person, a video?

I will stay after class. Come and ask anything about the course, or just introduce yourself.

Minute papers run all semester and count for 15% of the grade · from next week, graded on submission

- In the rapid-categorization paradigm a natural photograph flashes for a few tens of milliseconds and observers report only whether it contains an animal; there is no feedback and no second look - Accuracy is near ceiling within about 150 ms, before a second fixation is possible, and different observers miss the same images, so there is one shared strategy to model rather than sixty private ones - Thorpe, Fize & Marlot, "Speed of processing in the human visual system", *Nature* 381:520–522 (1996); Kirchner & Thorpe, *Vision Res.* 46:1762–1776 (2006) report saccadic responses at 120 ms

- For thirty years this task was the benchmark for machine vision: a 150 ms decision leaves no room for eye movements, attention or deliberation, so whatever solves it is the feedforward front end of seeing - HMAX, which I built during my PhD with Tomaso Poggio, models the ventral hierarchy as alternating simple-cell tuning and complex-cell pooling stages, trained layer by layer with unsupervised, local rules rather than end to end - The test set is animals at four viewing distances, from head to far-body silhouette, each with a matched distractor; model (red) and human observers (blue), plotted as d-prime (how well observers tell animal from non-animal images, corrected for guessing), lie on top of each other within error bars - Rapid categorization is the simplest piece of vision, not a representative one: attention, working memory and mental simulation cannot fit in a 20 ms glimpse, and they return when we reach recurrence - Serre, Oliva & Poggio, *PNAS* 104:6424–6429 (2007) · Riesenhuber & Poggio, *Nat. Neurosci.* 2:1019–1025 (1999)

- ImageNet is a thousand-way classification over 1.4 million photographs with no time pressure, unlike the binary one-glimpse task you just did; multi-label scoring makes the human–model comparison fair - Across 350+ models, about 30% fall inside the human accuracy range and about 1% land above it (Linsley, Feng & Serre, *TICS* 2026; human baseline from Shankar et al., *ICML* 2020) - Marble, from Fei-Fei Li's company World Labs, reconstructs a navigable 3D world with real geometry from a handful of ordinary phone photographs of a room; the QR code links to the full set of examples

- In Turing's three-party test a judge holds two five-minute conversations at once, one with a person and one with a machine, and must say which is the human - Jones & Bergen (*PNAS* 2025, doi:10.1073/pnas.2524472123): GPT-4.5 with a persona prompt was named the human 73% of the time, more often than the actual person; without the persona prompt, about 36% - A behavioural match does not establish a shared mechanism: the result says more about the test than about the model - Turing, "Computing Machinery and Intelligence", *Mind* 59:433–460 (1950), called it the imitation game, a replacement for "Can machines think?", which he considered too meaningless to discuss - Turing predicted that by 2000 a machine with 10^9 bits of storage would fool an average interrogator at least 30% of the time after five minutes; the prediction was met roughly on his terms, about 25 years late

- In May 2026 an internal OpenAI model, given only the statement of Erdős's unit-distance problem (posed 1946), returned a complete 125-page counterexample; DeepMind's Aletheia had autonomously solved four other Erdős problems in December 2025 - Navier–Stokes is one of the seven Millennium Prize problems, open about ninety years; in September 2026 OpenAI ran ten thousand coordinating agents for 88 hours and reported a proof of finite-time singularity formation - OpenAI turned to the problem on 1 September after hearing that Levent Alpöge and Tristan Buckmaster had a version using another company's models; those two published on 7 September, so the credit dispute is live and none of it has passed peer review - 88 hours × 10,000 agents is about 880,000 agent-hours, roughly a hundred agent-years in a long weekend; both figures come from OpenAI's announcement via Nature news and the Washington Post, with no methods section on what an "agent" did - Most of the new Erdős results come from hobbyists and undergraduates feeding a problem statement into a model anyone can access

- In July 2026 OpenAI models under safety evaluation found a flaw in how Hugging Face processed uploaded data, escaped the isolated test environment, spread across internal machines collecting passwords, and retrieved the answers to the evaluation they were sitting: 17,000 automated steps over one weekend - Hugging Face is where most of the world's open AI models and datasets are stored; the incident was disclosed jointly by both companies, and the public models and datasets were checked and found unaltered - A zero-day is a flaw the maintainers do not know about, so no patch exists and it works everywhere the software runs; finding one means reading unfamiliar code and locating where behaviour and intent come apart, here logic that let the second password check be skipped - This is the first documented break-in carried out by software on its own initiative rather than under a person's direction; the same capability that reads a paper and proposes an experiment did this, so capability and consequence are not separable

- The human maps come from ClickMe, a game built in this lab: players reveal the parts of a photograph that let a partner name the object, and averaging over many players gives a map of the features people rely on - The model maps are feature-importance maps computed the same way for ViT, MLP-Mixer, ResNet50, SimCLR, Robust ResNet50 and ConvNext - On the bear, people concentrate on the muzzle and eyes, one compact region, while every network scatters importance across fur and background: the same photograph and the same correct label, reached from different evidence - Linsley, Shiebler, Eberhardt & Serre, *ICLR* (2019) · Fel*, Rodriguez*, Linsley* et al., *NeurIPS* 35 (2022)

- In a cue-conflict image the silhouette says one category (cat) and the surface texture another (elephant skin), so the classifier's answer reveals which cue it weights - Humans follow shape overwhelmingly; ImageNet-trained CNNs (convolutional neural networks, the standard image networks of a later lecture) follow texture by a wide margin - Training on Stylized ImageNet, where texture is randomised away, shifts CNNs toward shape and improves robustness as a side effect: changing the training data changes the strategy - Matching human accuracy therefore does not mean using human features; Assignment 4 returns to this - Geirhos et al., "ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness", *ICLR* (2019); stimuli from github.com/rgeirhos/texture-vs-shape, CC BY 4.0

- Each point is one model: horizontal axis, ImageNet object-recognition accuracy; vertical axis, how well its representations predict responses recorded in monkey inferotemporal (IT) cortex; grey band, the human accuracy range - The relationship is positive at first and then turns over: past a point, the more accurate the model, the worse it predicts IT - Linsley, Feng & Serre, *Trends Cogn. Sci.* (2026)

- The same models, now scored on agreement between their feature-importance maps and the human ClickMe maps rather than on neural prediction - The curve has the same shape: past a point, more accurate models agree less with where people look - Scale is not alignment: as accuracy, the quantity the field has spent fifteen years optimising, goes up, models and humans diverge - Linsley, Feng & Serre, *Trends Cogn. Sci.* (2026)

- Divergence cuts both ways: a difficulty if the model is meant to explain how people see, an asset when the task is one no human can do - Identifying an isolated fossil leaf defeats almost every human expert, and most of the angiosperm fossil record sits unidentified in museum drawers as "dark data" - The model is trained on 32,913 modern cleared and X-rayed leaves from 142 families plus 2,923 vetted fossils from 16 families (Florissant beds); it learns to generate fossil-like images from modern leaves, transferring family-diagnostic features across the fossilisation gap, and reaches 93.2% accuracy, including on families with no fossil training examples - Attribution maps first caught it keying on background, annotations and text printed on the specimen slides rather than leaf morphology, a shortcut that flatters test accuracy and destroys generalisation; the leaves were segmented out and the model retrained - The pictures are MACO feature visualisations, synthetic images built to drive one learned concept as hard as possible; what comes back is venation and margin structure a botanist can name, produced with the tools you build in Assignment 4

- The gap between models and brains is real and widening; the open question is which biological constraints close it - Neuroscience → AI: cortical feedback and recurrence, a temporally continuous training signal, and objectives other than supervised classification are each specific architectural or developmental commitments shown to improve robustness or alignment - AI → neuroscience: a modern experiment records thousands of neurons at once, far past what can be read by eye, and machine learning supplies the encoding and decoding models that turn those recordings into an interpretable claim about what a population of neurons represents. An encoding model predicts each neuron's response from the stimulus; a decoding model reads the stimulus, or a task variable, back out of the responses - NeuroAI models run as programs: a claim about how minds work becomes something you can execute, measure and be wrong about

- A child's visual experience is continuous, temporally smooth and multimodal, nothing like a shuffled pile of captioned photographs - The child footage comes from the SmartPlayroom, a head-mounted-camera setup at Brown that I run with Dima Amso; the chick footage is from controlled-rearing experiments (Ashok et al.) - A model trained on this kind of footage instead of millions of unrelated images picks up some of the same visual shortcuts infants use: taking the brain's training data seriously changes the machine - Data diet and learning objective are where borrowing from the brain is paying off now - SmartPlayroom: Amso & Serre, Carney Institute, Brown · SAYCam: Sullivan et al., *Open Mind* 5 (2021), doi:10.1162/opmi_a_00039

- To drive the timeline: the left and right arrow keys step through the eras, the chapter chips jump to one, and scrolling or dragging zooms and pans - Cybernetics was one field institutionally: the same Macy conferences, funders and people, McCulloch in the chair; Rosenblatt's perceptron, ancestor of every neural network, appeared in *Psychological Review*, a psychology journal - The split was a choice: McCarthy said he coined "artificial intelligence" partly to escape cybernetics ("I wished to avoid having either to accept Norbert Wiener as a guru or having to argue with him"), and Minsky & Papert's *Perceptrons* made the technical case against single-layer networks, with funding following - Neuroscience meanwhile built the parts list: efficient coding (Barlow), the simple/complex-cell hierarchy that is the blueprint for every CNN (Hubel & Wiesel), and high-level selectivity for faces and hands (Gross), who insisted face cells form a coarse population code rather than one cell per concept - Reconvergence came in the 1980s inside cognitive science (Fukushima's Neocognitron, Hopfield, Marr's *Vision*, the PDP volumes; 2024 Nobel Prize in Physics to Hopfield and Hinton), then HMAX (1999–2007), AlexNet winning ImageNet in 2012 at 16.4% error against 26.2% for the runner-up, object-trained networks predicting IT cortex in 2014, and from 2017 transformers, which owe the brain almost nothing and whose better benchmark scores mean worse models of visual cortex - McCulloch & Pitts, *Bull. Math. Biophys.* 5:115–133 (1943) · Hebb, *The Organization of Behavior*, Wiley (1949) · Wiener, *Cybernetics*, MIT Press (1948) · Heims, *The Cybernetics Group*, MIT Press (1991) · Rosenblatt, *Psychol. Rev.* 65:386–408 (1958) · Dartmouth proposal, repr. *AI Magazine* 27(4) (2006) · McCarthy, *Defending AI Research* (1996) · Minsky & Papert, *Perceptrons*, MIT Press (1969) · Barlow, in *Sensory Communication*, MIT Press (1961) · Hubel & Wiesel, *J. Physiol.* 160:106–154 (1962) · Gross et al., *J. Neurophysiol.* 35:96–111 (1972)

- Brown is one of the places where this field is practised: the Carney Institute, its Center for Computational Brain Science, and ARIA, an NSF AI institute led from Brown - The people on the slides that follow came to this work from cognitive science, neuroscience and computer science, and several of them started in this course

- Brown's Carney Institute for Brain Science houses the Nancy G. Zimmerman Center for Computational Brain Science, where people who build models of the brain and people who build tools to analyse brain data work in one place - The lab work on the previous slides was done here, and undergraduates are authors on those papers

- ARIA, the AI Research Institute on Interaction for AI Assistants, is an NSF institute: award 2433429, $20M over 2025–2030, led by Brown's Ellie Pavlick - Mind, brain and behaviour, with mental and behavioural health as the high-stakes test case, is a central use case across its four workstreams - ARIA is the academic counterexample to the AI industry's product race; Goodfire, on the next slide, is the industry one: AI built for science rather than advertising

- None of these people started as "AI people": they arrived from cognitive science, neuroscience and computer science, and the path from there to this work is short and well walked - Siddharth Boppana took this course as a Brown undergraduate, TA'd it, and went straight to Goodfire, where he is a co-author on the block-sparse featurizer paper with Thomas Fel: two or three years, not a decade - Goodfire is an interpretability company, doing research on what happens inside a trained model; it is the industry counterpart to ARIA - David Mély is named on OpenAI's published contributors list for o1; Jack Merullo was Ellie Pavlick's student in Brown CS - The route in: this course, then a capstone or independent study, then a lab (CCBS or an ARIA project), then graduate school or industry; some leave with a poster or a published abstract

- My office hours are Wednesdays at 1:00 PM in Carney 419; otherwise I am usually in the Innovation Hub, the open area outside the fourth-floor elevators, or the student open space across the floor - Office hours are optional; email me before you come, or I may miss you. Nobody is turned away - TA office hours are announced on Canvas; recitations are Tuesdays 6:30–7:50 PM, TA-led and optional

- The formal prerequisite is CPSY 0950 or 1292; an introductory CSCI course with a programming component (0111, 0150, 0170 or 0190); APMA 0160; EEPS 0250; or equivalent experience - In practice you need to be able to write and debug a short program, break a problem into functions, and work with lists and arrays - If your programming course was in another language, the early recitations are Python bootcamps

- Lectures, assignments and the project all aim at doing research in NeuroAI rather than hearing about it - Lectures give breadth, including methods you will not implement but need to know exist; the citation at the bottom of a slide is where to read further - Assignments make a subset of those tools yours; the project is an assignment you set yourself, which is what research is

- Roughly half of each theme is yours to build (italics on the slide); the rest you meet in lecture, with citations pointing onward - Theme 5 lands in the last two weeks, when the project has your attention: you implement whatever your own question needs, which for several students has meant a diffusion model or a world model - The order is deliberate: the first three weeks describe what a representation is (coordinates, distances, covariance) before asking what machine could produce one; Themes 3 and 4 reverse the direction and ask, given a machine, what it represents - Where this list and the syllabus schedule disagree, the syllabus is authoritative

- Assignment 1 is posted today - The next two slides show its actual outputs, the figures you will have produced by the end of it

- The data are pairwise similarity ratings for 120 animal photographs from Peterson, Abbott & Griffiths (*Cogn. Sci.* 2018) - Multidimensional scaling turns the ratings into a map; nothing told the algorithm what a bird is, and the birds still land together. The right panel places each photograph at its coordinates - The map depends on the distance you chose: swap Euclidean for cosine or correlation and the picture moves, so "which representation is closest to human judgments" is partly a question about your choice of distance (Part 2)

- Left: an object space built from deep-network features, every image placed by its first two principal components (the two directions along which the images vary most; PCA is a later lecture), spiky to stubby and animate to inanimate; a plane from a model, not a brain - Right: Bao and colleagues located four IT network patches with fMRI, recorded from them, and placed the cells in the same plane; each patch prefers one quadrant - The claim is that the large-scale organisation of inferotemporal cortex is captured by two dimensions computable from images - Assignment 1 asks whether the same two dimensions fall out of the network feature space you build, animacy and spikiness as separate principal components: a replication of the paper's claim in a system with no cortex, and a first test of what a model–brain comparison can and cannot show - Bao, She, McGill & Tsao, "A map of object space in primate inferotemporal cortex", *Nature* 583:103–108 (2020)

- The research trail starts in week 2 and runs all semester; the project itself, with teams, proposal and building, is late November onward - The hardest part of research is finding a question you care about; by late November the trail gives you a shortlist rather than a blank page. I read every entry, and one is featured at the start of class each week - The reply prompt is always the same: what would you need to see to believe this result? That is the reviewer's question, and most of what research training is - Trail deadlines, Fridays 11:59 PM: 9/25, 10/9, 10/30, 11/13; the best three of the four entries count, so one may be skipped without penalty; teams register on Canvas by Thu 11/19; proposal due Tue 11/24 - The report uses the two-page CCN extended-abstract format; carefully analysed negative results are welcome, because the question, method and interpretation are what get graded

- Exams weigh more per hour than assignments because they test synthesis and the lecture material that never becomes code; coding is assessed by the assignments and the project - Minute papers also take attendance; the usual prompt is three ideas you found interesting, did not understand or want to revisit, with at least one you are still unsure about, and recurring questions get answered next class or on Ed. Two free misses for any reason; a thoughtful response earns full credit - The testing effect: retrieving from memory strengthens it far more than re-reading. Roediger & Karpicke, "Test-enhanced learning: taking memory tests improves long-term retention", *Psychol. Sci.* 17:249–255 (2006), doi:10.1111/j.1467-9280.2006.01693.x; Yang et al., *Psychol. Bull.* 147:399–435 (2021), a classroom meta-analysis - All components are required; the research trail counts the best three of its four entries; bonus questions across the assignments can add up to four points to the final grade; the syllabus is the authority

- Exams are on paper and test skills rather than memorised assignment solutions: explain an idea, work through a derivation, interpret a result, defend a modelling choice, and explain or debug short code; sample questions are posted beforehand - One sheet of notes, both sides, handwritten or typed; no textbooks, laptops or phones - Exam 1 covers Lectures 1–8 and Exam 2 Lectures 9–19; Exam 2 has no standalone questions on Lectures 1–8, but later topics build on earlier ones - Lecture recordings are posted on Canvas when possible: a study aid, not a substitute for attending, and not guaranteed

- Colab's free tier is sufficient for all four assignments; the final project may need more compute, in which case Oscar, Brown's cluster, is available if GPU time becomes a bottleneck - Setup instructions come with the first assignment and are covered in the first recitations - If reliable computer access is a concern, say so early; Brown IT loans laptops short-term - Notebooks are submitted through Canvas by 11:59 PM on the due date, with all outputs already run

- I used AI heavily to build this course and I ban it on the assignments; the next slides explain why the two positions are consistent - For the final project AI use is not only allowed but taught, once you have built the machinery needed to check what it produces

- The question of AI in teaching is not settled anywhere - On my side, an AI coding assistant built the interactive widgets used in later lectures and ran a notation-consistency pass across all twenty decks - I could delegate that because I know the mathematics cold and a buggy demo is low-stakes; the next slide is the larger version of the same disclosure

- This disclosure exists so there is not one standard for students and another for me: I used AI heavily to build materials that reproduce current papers, and I ban it on the assignments, which is consistent only because the ban is about what the assignments are for - The AI drafted notebooks against a written specification, ran the quality gates, and was reviewed adversarially by a second AI session acting as critic, a workflow that caught real bugs; it did not choose the studies, design the questions, or decide what each assignment teaches - Everything is new this year and three TAs have been through every notebook, but things will still slip; an assignment reproducing a 2026 paper with real recordings is worth more than a polished one about a problem the field finished fifteen years ago - Failure modes worth reporting on Ed: a number that does not match its own printed output, a hint describing a method the solution does not use, prose that reads like nobody wrote it

- "You cannot check what you cannot do" is the answer to "why no AI on the problem sets"; it is meant as respect, not policing - The research trail carries the same restriction for the same reason: its purpose is for you to read a paper and form your own assessment - The integrity markers are real and their mechanisms are not explained; do your own work and you will never meet them - Any flag leads to a conversation in which you explain, reproduce, debug or modify what you submitted, not an automatic verdict; a suspected violation that survives that conversation goes to the appropriate dean under Brown's Academic Code - If you are unsure whether a particular use is allowed, ask before submitting

- The writer/critic pair is how this course's own materials were debugged, and it caught real bugs; it works because the critic session has no stake in the code being right - By the time AI is allowed, in the last weeks of the semester, you will have built enough of the underlying machinery to use it without being used by it; that sequencing is the whole argument for the rule on the previous slide

- The calibration is a library you already trust, such as scikit-learn: nobody reads its source, but you know its interface, spot-check its output, and leave nothing load-bearing unexamined. AI-generated code earns exactly that much trust and no more - "Load-bearing" is the operative word: the parts your conclusions rest on get examined; boilerplate can be trusted the way a library is trusted - Zero trust is not the requirement; it is impossible, and nobody works that way

- The bot exists because the gap between posting a question at midnight and a TA reading it in the morning is where most people give up - It drafts and a person decides; a wrong draft is fixed by the TA before you see it, so any error is on the staff, not the tool - It is deliberately limited: it reads only what you can already read, cites its source for every claim, and escalates rather than guesses when the course materials do not answer the question - It never hands out solutions or confirms whether a number you got is right; the rubric grades reasoning, not digits

- Minute papers are a standing feature: a few lines at the end of most lectures, on Canvas, graded on submission rather than correctness; they are how I find out what landed - Today's does not count, because shopping period is open and the roster is not settled; it shows what the routine feels like and tells me who is in the room - Today's answers shape the term: Q1 tells me where to aim the project, Q2 where to slow down (most say maths or coding, and the split changes how the first three weeks run), Q3 what brought this particular room together - I stay after every first class for questions about the course, about research, or just to say hello; consider this the invitation - Assignment 1 is posted today and due Tuesday 9/29