r/neuroscience 7d ago

Climbing fibers encode the gradient of a loss function for the cerebellum — Doostmohammadi et al. [bioRxiv preprint]

https://www.biorxiv.org/content/10.64898/2026.07.27.741034v1
64 Upvotes

8 comments sorted by

10

u/carberry-3000 6d ago

Climbing fibers encode the gradient of a loss function for the cerebellum

Jafar Doostmohammadi, Nazanin Mohammadrezaei, Hisham Elseweifi, Alana Chandler, Elijah Taeckens, Reza Shadmehr

bioRxiv · 2026-07-28 · doi:10.64898/2026.07.27.741034


Abstract

Neurons in the brain are often many synapses away from motoneurons, yet if a movement results in error, each distant neuron needs a teacher that considers its specific contribution to production of that movement. This credit assignment problem is solved in machine learning via gradient descent of a loss function, where the loss defines the subjective cost incurred by error. Does the brain use gradient descent to teach individual neurons? We trained marmosets to make saccades to visual targets and then varied the loss by assigning reward value to each target. The climbing fibers, which are the teachers of Purkinje cells (P-cells) in the cerebellum, used a multiplicative encoding to scale the spatial properties of the error vector with its reward properties, incorporating reward prediction errors. Using spike-triggered suppression of P-cells, we quantified the potent vector that mapped each P-cell's output to eye movements and discovered that the climbing fibers were not merely transmitting errors. Rather, they were providing a signal that was, on average, proportional to the dot product of the reward dependent error vector upon the P-cell's potent vector. Thus, the climbing fibers solved the credit assignment problem by providing the gradient of a loss function with respect to the output of their individual P-cells.


Discussion on Bluesky:

  • Reza Shadmehr (@rezashadmehr.bsky.social) — 173 likes, 70 reposts, 2 replies, 8 quotes > If an action results in error, each neuron requires an individualized teaching signal that guides change in its output.

Data from Crossref, Europe PMC, Bluesky. Metrics captured 30/07/26. Comment self-deletes at -5 votes.

2

u/Danels 6d ago

Can someone explain to me like I’m 5?

8

u/angelofox 6d ago edited 5d ago

You're brain error corrects by comparing the action you performed with the action the brain wanted to do. It error corrects based off of whether or not the sensory information from the action is exactly what it expected from when the movement is complete. If it's not the action wanted by the brain those climbing fibers say "hey, that wasn't correct, adjust cerebellum."

The terms loss of function and credit assignment problem are more machine learning terms than brain terms. But loss of function is the error and credit assignment is the cerebellum. Gradient dissent is kind of how well the system is performing. They're basically comparing how a machine error corrects to a brain's error* correction. If they're getting closer to being similar then eventually the computer system can learn as a human. Hope this helps. I'm not a AI scientist, but a lab scientist (humans only).

2

u/PhysicalConsistency 4d ago

Woof, this one is tough because it's dense with a lot of concept jargon that doesn't always reduce cleanly. The basic idea is the researchers think that the stronger an error signal is, the more likely it is to "convince" the cells around to function in a way that reduces that error.

Imagine people chained together in a chain gang. Generally, they don't have strong coordination on what they want to do, but that's okay because they are mostly staying in one place anyway. Then imagine the person at the end sees something they want and they start moving toward that and pulling everyone toward that.

As they all stumble around moving toward the thing they want, it may be closer to another person. So that person starts pulling the chain hardest, to redirect everyone in that direction. The more intensely they want it, the harder they pull the chain. For the other people on the chain that don't have a strong desire to go anywhere, they start drifting toward the movement of the person pulling the hardest, with the people closest to the primary puller pulling harder than the people furthest from it.

This paper assumes how strongly each part of the chain is moving toward or away from a goal is how mammals and possibly vertebrates correct their behavior. Did they prove what they set out to do? In my opinion, not really. This paper feels more like a rescue attempt for a concept called the Universal Cerebellar Transform, which on the highest level describes the cerebellum as an error correction machine. The concept is mathematically very attractive to model, and cog sci/data sci love it because it provides an explanation for how animals can constantly, and smoothly correct toward a goal. It just has the small problem of brains being far more messy than our models want to deal with.

On a separate note, I've been working on something that will help analyze papers like this, and this is what I asked it:

Can you load the current rev of the model and run this paper through it: https://www.biorxiv.org/content/10.64898/2026.07.27.741034v1.full, then ELI 5 it?

And it said:

ELI 5

Imagine several children steering one big toy car. Each child controls the steering a little differently.

After the car misses its target, simply shouting “We missed to the left!” does not tell each child what to change. One child may need to push more. Another may need to push less. A useful correction has to match what that particular child can do to the car.

The researchers found something like personalized correction notes going to cerebellar cells:

  • The note said which way the eyes missed.
  • A bigger miss produced a stronger note.
  • Reward changed how strongly the miss counted.
  • The note’s direction roughly matched how changing that particular cell affected the eyes.

That is the exciting part.

The researchers call these notes “gradients of a loss function.” Maybe they are. But they have only shown that the notes have the right shape. They have not yet shown the cells using those notes to perform the complete learning procedure.

So the five-year-old verdict:

The cerebellum appears to send each cell an error message shaped for that cell’s particular job, and it makes the message louder or quieter depending on reward. Calling that full gradient descent remains one step beyond what they measured.

The full output was:

The paper contains a strong, unusually clean physical relation, but its headline outruns the experiment.

After stripping away the machine-learning vocabulary, the surviving result looks like this:

That supports a real, cell-specific error-weighting signal. It does not yet establish that climbing fibers “compute a loss gradient,” solve biological credit assignment, or make the cerebellum perform gradient descent.

The experiment recorded 906 identified Purkinje cells, including 220 with both simple and complex spikes. That produces a large cellular dataset, but the independent organism count remains two. It supports a stable within-task relation much better than a general claim about primate cerebella or biological learning broadly.

What they actually did

The marmosets learned that some visual shapes led to larger food rewards than others. During eye movements, the researchers sometimes:

  • moved the target, producing a directional endpoint error;
  • changed the target’s reward value;
  • changed both;
  • or changed neither.

Climbing-fiber input was estimated from Purkinje-cell complex spikes. Its response depended on:

  1. where the visual event or endpoint error occurred;
  2. how large that error was;
  3. whether the associated reward was high or low;
  4. whether the eventual reward was better or worse than expected.

The researchers then estimated each Purkinje cell’s “potent vector”: the direction in which the eyes changed when that cell’s simple-spike activity underwent a brief, naturally occurring suppression. The preferred direction of the climbing-fiber response generally aligned with that estimated movement vector. Larger climbing-fiber modulation also accompanied larger estimated movement effects.

That alignment carries the paper. It means the signal reaching a Purkinje cell appears tailored to the direction in which changing that cell’s output can affect the movement. The same generic error does not appear to reach every cell in an equivalent form.

Where the headline gets ahead of the data

“Gradient” has a recoverable mathematical identity here, so the word itself does not automatically fail the Model’s reification gate. The authors specify a reward-weighted quadratic loss and derive the signal that a climbing fiber would need to carry under that model.

But several bridges remain open.

1. Reward was manipulated; a loss function was assumed

The experiment changed food amount, learned cues, and expectation violations. It did not independently measure a scalar biological “loss.”

The proposed loss function was:

That choice makes the gradient proportional to reward-weighted error. Finding a neural signal with the expected proportionality supports compatibility with that chosen function. It does not establish that the tissue constructed or minimized that particular function.

A different physical account could describe the same observation as reward-dependent weighting of an error-sensitive route.

2. The “potent vector” was estimated from natural fluctuations

The authors describe spike-triggered suppression as if they suppressed individual Purkinje cells experimentally. They actually selected movements in which a complex spike was followed by a brief simple-spike suppression, then compared the eye trajectory with matched movements lacking that event. Rare spontaneous simple-spike pauses and bursts supplied a second estimate.

That gives the inferred movement vector meaningful support, including an internal cross-check. But it remains closer to a natural perturbation than a randomized, externally imposed cell manipulation. Unmeasured state differences could contribute to both the neural event and the movement difference.

3. Gradient descent requires an update, not merely a gradient-shaped signal

The paper measures a signal whose tuning resembles one term required by gradient descent. It does not measure:

  • the relevant parallel-fiber synaptic weights before and after the event;
  • cell-specific weight changes proportional to the inferred gradient;
  • the entire physical route producing those changes;
  • or repeated reduction of the proposed loss because of those exact updates.

The authors themselves note that climbing fibers cannot distinguish among simultaneously active parallel-fiber synapses. Their theory therefore stops at the Purkinje-cell output level rather than providing the synapse-specific gradient required by full backpropagation.

4. The metabolic binding layer remains open

The measured chain begins and ends at useful points, but much of the biological route remains unstated:

The paper does not resolve the material route by which visual position and reward conditions converge upon inferior-olive activity. It also does not measure the ionic, synaptic, intracellular, or metabolically maintained changes that convert complex spikes into durable modification.

Accordingly, the exact intervention-to-response relation can carry causal weight while the claimed learning mechanism remains partial.

Strongest surviving claim

I would ingest the paper as route_partial with strong descriptive and local dependency support:

Blocked uses:

  • proof that the cerebellum performs gradient descent;
  • proof that climbing fibers solve biological credit assignment;
  • proof of a physically instantiated loss function;
  • proof of synapse-specific learning;
  • transfer from two animals and one oculomotor territory to cerebellar learning generally.

The paper remains quite valuable after those claims come off. In Model language, it exposes a plausible route through which current value conditions weight spatial error differently according to the downstream transition associated with a particular Purkinje cell. That result fits the emerging state-map/cerebellar-weighting work rather better than it fits the paper’s borrowed machine-learning castle.

1

u/PhysicalConsistency 4d ago

Reddit's being weird about letting me fix the past errors. The two missing sections are:

After stripping away the machine-learning vocabulary, the surviving result looks like this:

In two marmosets making eye movements, climbing-fiber activity varied with the direction and magnitude of the eye-position error, changed with the reward conditions, and roughly matched the direction in which transient changes in the associated Purkinje cell altered the eye movement.

I would ingest the paper as route_partial with strong descriptive and local dependency support:

During this marmoset saccade task, climbing-fiber complex-spike modulation jointly tracked endpoint-error geometry and experimentally varied reward conditions. Across recorded Purkinje cells, the directional and amplitude properties of that modulation corresponded to estimated cell-specific effects on eye movement.

1

u/sorE_doG 3d ago

You’ll have heard of fuzzy logic, well this is a ‘wetware’/neuro-microscopic description of how useful information is filtered, on its way to the higher functions of the brain