Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

2.1 Credit Assignment in SNNs

Technical University of Denmark

To improve any neural network, we need to solve the credit assignment challenge: when the network makes an error, how do we determine which components to adjust? And by how much? This chapter explains the intuition and mathematics behind credit assignment with an emphasis on time because, unfortunately, credit assignment becomes harder when you work with stateful neurons and networks, which are described in the previous topic on Foundations of SNNs.

2.1.1The Credit Assignment Problem

The core question in this chapter is: Given an output error, which neurons and synapses contributed to that error? The term credit assignment was coined by Marvin Minsky in a 1960 paper Minsky, 1961, but he went the other way around: how do we reward the good parts of the network. Whether you think about credit as “reward” or “punishment”, the essence of the problem is the same:

The credit for a working program can only be assigned to [...] subroutines, and as these operate in hierarchies we should not expect individual instruction reinforcement to work well. - Minsky (1961)

What Minsky is saying, is that our networks consist of subcomponents that, together, make up the whole. So, if we are to assign credit and improve the network, we first need to understand how each component contributes to the whole. What is the contribution of, say, a synapse to the network output? Only when we understand that contribution, do we know how to change the synapse. That is the fundamental problem of credit assignment: how do we assign credit to the individual components?

2.1.1.1Credit assignment in space

The structural credit assignment problem is to determine which internal decisions are responsible - Minsky (1961)

In a feedforward network, credit assignment reduces to a structural question: which connections along the path from input to output are responsible for the error?

Consider a single linear layer with a single weight, and no bias term, as shown in Figure 1. Following the example in the figure, the layer receives some input (1), and the objective is to output 2. In this example, ww is set to 2{2}, so the output 1∗w1=1∗2=21 * w_1 = 1 * 2 = 2 is off by one. We can solve this by adjusting the linear weight from 2→12 \to 1, which gives 1∗1=11 * 1 = 1: we’ve reached our goal!

Animation showing spatial credit assignment.

Figure 1:Spatial credit assignment. A linear weight gets readjusted from 2→12 \to 1 to ensure that the output is the same as the goal: 1.

Now, consider the case shown in Figure 2, where we have two linear layers, both with w=2w = 2, and a goal of 2{2}. The output becomes 4{4}, since 1∗w1=1∗2=21 * w_1 = 1 * 2 = 2 and 2∗w2=2∗2=42 * w_2 = 2 * 2 = 4. Our error is now 2{2}, but which parts of the network do we update? And by how much? The simplest solution would probably be to set w1w_1 or w2w_2 to 1{1}, but there are many more solutions.

Animation showing spatial credit assignment.

Figure 2:Spatial credit assignment with two layers. Which layer should receive the update? With multiple components, this becomes less clear.

Solution to Exercise 1 #

We could set w1=0.5w_1 = 0.5 and w2=4w_2 = 4, which yields 1∗0.5=0.51 * 0.5 = 0.5 and 0.5∗4=20.5 * 4 = 2.

In general, we need to solve the equation 1∗w1∗w2=21 * w_1 * w_2 = 2. But we have more than one variable, so there are infinitely many solutions! This is why credit assignment is hard.

Even in this minimal example, there is no single correct answer: many combinations of weights produce the same output. This ambiguity is at the heart of spatial credit assignment, and it only grows worse as networks get deeper and more complex. A caveat on the word ambiguity: that many weight settings solve the task does not make credit undefined. Given a loss, the gradient still prescribes one well-defined update. It is the solution set, not the credit signal, that is degenerate: many distinct weight settings collapse onto the same output, so the task alone singles out none of them. Now, imagine adding time to the equation.

2.1.1.2Credit assignment in time

In networks without memory, credit assignment only needs to trace errors backward through layers. But neurons with state carry information across time, and a spike at timestep tt may cause an error at timestep t+100t + 100. Now we must assign credit not just to the right component, but to the right component at the right moment.

To see why, take the single-weight example from Figure 1 and unroll it in time: the same neuron, with the same weight ww, but now using state. Acting at t=0t=0 and again at t=1t=1 carries state forward from one step to the next (Figure 3). Notice that this is the same picture as the two-layer case in Figure 2, only now the “depth” is time rather than space. The output is wrong by the same amount, and the same question returns: which timestep should receive the update?

The catch is that ww is shared: the very same parameter acts at every timestep. When we ask how a small change in ww affects the final error, the answer is not a single number but a sum: one contribution from the role ww played at t=0t=0, another from its role at t=1t=1, and so on for every step it was active. Unrolling the network in time and differentiating makes this explicit:

∂L∂w=∑t∂L∂w∣t\frac{\partial L}{\partial w} = \sum_{t} \left.\frac{\partial L}{\partial w}\right|_{t}

This is exactly the backpropagation through time (BPTT) computation Werbos, 1990: we unroll the recurrence into a deep feedforward graph, one layer per timestep, and propagate the error backward through it.

Bundling every moment into a single gradient is convenient for computing an update, but it hides the same ambiguity we met in space. The total error could be blamed on what the neuron did at t=0t=0, on what it did at t=1t=1, or split between the two in any proportion; the sum in (1) constrains only the whole, never the per-step shares. Where the spatial two-layer case had infinitely many pairs (w1,w2)(w_1, w_2) that yield the same output, the temporal case has infinitely many ways to distribute one shared weight’s error across the moments it acted. Time, in other words, is just another axis of depth - and credit assignment is ambiguous along it for exactly the same reason.

Animation showing temporal credit assignment across two timesteps.

Figure 3:Temporal credit assignment. The same neuron (weight w=2w = 2) acts at t=0t=0 and t=1t=1, carrying its state forward. The output is off by the same error as the two-layer case, and the dashed feedback shows the ambiguity: which timestep should receive the update? As with multiple spatial weights, there are infinitely many solutions.

The gradient in (1) really flows through the chain of state updates ∂V[t]/∂V[t−1]\partial V[t]/\partial V[t-1]. Carrying the membrane potential step by step through the entire chain of state updates makes SNNs especially hard:

2.1.1.3Credit assignment parameter updates

Once we know which components contributed to an error, we face a second question: how do we translate that knowledge into concrete parameter changes?

2.1.1.4A note on scalability

The ambiguity we have discussed so far is a question of correctness: which component, at which moment, deserves the credit. Scalability is a separate, practical question: even when we know how to compute the credit, can we afford to?

Backpropagation through time answers the credit question by unrolling the network into one layer per timestep and storing every intermediate state, so it can later walk the errors backward. That storage is the bottleneck. Memory grows with the sequence length TT multiplied by the size of the network, because every activation along the way must be kept until the backward pass reaches it; compute grows with the same product, since each stored step must also be revisited. A long recording or a deeply recurrent loop can exhaust memory long before the method itself runs into any limit.

This cost is a large part of why the coming chapters exist. It motivates truncating the unroll to a short window, propagating sensitivities forward in time instead of backward (real-time recurrent learning (RTRL), which stores nothing extra as the sequence grows but scales poorly in network size Marschall et al., 2020) or abandoning the global backward pass entirely in favor of local rules (such as eligibility traces, plasticity, and reward signals) that update each synapse from information it already has on hand. Correctness tells us what credit assignment should compute. Scalability decides which approximation we can actually run.

2.1.1.5Why not just use backpropagation?

Beyond its memory cost, BPTT sits awkwardly with biology and neuromorphic hardware Lillicrap et al., 2020: the backward pass needs the transpose of the forward weights (the weight-transport problem), it requires separate locked forward and backward phases, and its updates are non-local. These objections motivate the alternatives that follow, some of which keep backprop’s form but swap the transposed weights for fixed random ones (feedback alignment Lillicrap et al., 2016, direct feedback alignment Nøkland, 2016) while others drop the global backward pass entirely.

2.1.2Approaches to solving credit assignment

The rest of this topic explores three families of approaches, each developed fully in the chapters that follow:

2.1.2.1Choosing a method: a simple heuristic

Now that you have multiple options, the practical question is when to reach out for which. The choice depends on what is available (a differentiable loss? a reward signal? a compute budget?) and on what you need (biological plausibility? on-chip locality? sample efficiency?) Here are a few questions to guide you:

The table below summarizes the trade-offs. Plasticity- and eligibility-based rules are closely related (both are local) and evolution is one instance of the broader direct-optimization family. They are separated here only where the distinction changes when you would reach out for them.

MethodUse when...Avoid when...Typical cost
Gradient-based (surrogate)You have a differentiable loss and can store the unrolled graphSequences are very long, memory is tight, or no gradient existsHigh memory, growing with sequence length
Plasticity-based (local/Hebbian)You need an on-chip, biologically plausible, online ruleA precise global error must be minimizedLow; local and online
Eligibility-basedLearning is local but the reward or error is delayedA dense, immediate gradient is availableLow–moderate; one trace per synapse
Reinforcement-basedFeedback is a sparse, delayed, or non-differentiable rewardA dense supervised target is availableModerate–high; high variance, sample-hungry
Evolution-basedNo gradient exists but many forward passes are cheapYou have a good gradient and limited computeHigh compute, but embarrassingly parallel
Direct optimizationThe parameter space is tiny or no gradient existsThe parameter space is largeScales poorly with increasing dimensionality

These families are not mutually exclusive. In practice they are often combined. Eligibility traces carry a reward signal, surrogate gradients seed an evolutionary search, and plasticity rules can themselves be meta-learned by evolution. Treat the heuristic as a starting point, not a verdict.

2.1.3Summary

It is my conviction that no scheme for learning, or for pattern-recognition, can have very general utility unless there are provisions for recursive, or at least hierarchical, use of previous results. We cannot expect a learning system to come to handle very hard problems without preparing it with a reasonably graded sequence of problems of growing difficulty. - Minsky, 1961

As Minsky anticipated, credit assignment in SNNs requires exactly this kind of hierarchical reasoning—tracing contributions through layers of space and steps of time.

Here is a list of resources sorted by topic (please help expand the list):

References
  1. Minsky, M. (1961). Steps toward Artificial Intelligence. Proceedings of the IRE, 49(1), 8–30. 10.1109/JRPROC.1961.287775
  2. Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press. https://mitpress.mit.edu/9780262039246/reinforcement-learning/
  3. Werbos, P. J. (1990). Backpropagation Through Time: What It Does and How to Do It. Proceedings of the IEEE, 78(10), 1550–1560. 10.1109/5.58337
  4. Amari, S.-I. (1998). Natural Gradient Works Efficiently in Learning. Neural Computation, 10(2), 251–276.
  5. Marschall, O., Cho, K., & Savin, C. (2020). A Unified Framework of Online Learning Algorithms for Training Recurrent Neural Networks. Journal of Machine Learning Research, 21(1), 1–34.
  6. Lillicrap, T. P., Santoro, A., Marris, L., Akerman, C. J., & Hinton, G. (2020). Backpropagation and the Brain. Nature Reviews Neuroscience, 21(6), 335–346. 10.1038/s41583-020-0277-3
  7. Lillicrap, T. P., Cownden, D., Tweed, D. B., & Akerman, C. J. (2016). Random Synaptic Feedback Weights Support Error Backpropagation for Deep Learning. Nature Communications, 7(1), 13276. 10.1038/ncomms13276
  8. Nøkland, A. (2016). Direct Feedback Alignment Provides Learning in Deep Neural Networks. Advances in Neural Information Processing Systems 29 (NeurIPS), 29.
  9. Gerstner, W., Lehmann, M., Liakoni, V., Corneil, D., & Brea, J. (2018). Eligibility Traces and Plasticity on Behavioral Time Scales: Experimental Support of neoHebbian Three-Factor Learning Rules. Frontiers in Neural Circuits, 12, 53. 10.3389/fncir.2018.00053
  10. Bellec, G., Scherr, F., Subramoney, A., Hajek, E., Salaj, D., Legenstein, R., & Maass, W. (2020). A Solution to the Learning Dilemma for Recurrent Networks of Spiking Neurons. Nature Communications, 11(1), 3625. 10.1038/s41467-020-17236-y
  11. Rumelhart, D. E., Hinton, G. E., & Williams, R. J. (1986). Learning representations by back-propagating errors. Nature, 323(6088), 533–536. 10.1038/323533a0
  12. Williams, R. J., & Zipser, D. (1989). A learning algorithm for continually running fully recurrent neural networks. Neural Computation, 1(2), 270–280. 10.1162/neco.1989.1.2.270
  13. Bengio, Y., Simard, P., & Frasconi, P. (1994). Learning long-term dependencies with gradient descent is difficult. IEEE Transactions on Neural Networks, 5(2), 157–166. 10.1109/72.279181
  14. Martens, J. (2020). New insights and perspectives on the natural gradient method. Journal of Machine Learning Research, 21(146), 1–76.
  15. Singh, S. P., & Sutton, R. S. (1996). Reinforcement Learning with Replacing Eligibility Traces. Machine Learning, 22(1–3), 123–158. 10.1007/BF00114726