Continual learning with no offline phase: local replay reaches 91.6% on split-MNIST
Most replay-based continual learning consolidates old knowledge either in a dedicated offline phase or by mixing replayed samples into the incoming data stream. The authors of this arXiv paper point out that brains also consolidate while awake, through local sleep: brief, use-dependent off-periods of individual circuits. They ask whether a network trained only by local, biologically constrained rules can consolidate with no offline phase at all.
Their scheme has three parts. An isolation rule confines replay updates to hidden synapses that are invisible to the current input under k-winner-take-all dynamics, and the optimiser state advances only inside that mask. A refractory rotation rule makes units that have just fired sit out the next competition, which widens the set of units that can be consolidated. A homeostatic pressure and a relative-novelty gate decide when replay bursts fire and when rotation runs.
The authors say this inverts the usual direction of non-interfering continual learning. Instead of protecting old memories from new input, the hidden computation on the current input is held invariant, exactly on the proven channels and for all but 0.3% of waking samples per update elsewhere, while past memories are written into the degrees of freedom the current batch leaves unused.
On class-incremental split-MNIST the system reaches 91.6+-0.3% with no offline phase. That is at or above the best offline-night schedule on two held-out splits, tied with DER++ and above experience replay, ER-ACE, A-GEM and unmasked local replay. In a single pass it leads DER++ (91.8% against 90.1%), while the offline-night schedule falls to 76.9%. The advantage is largest at small buffers and gives way to the backpropagation references at large ones. On split CIFAR-10 the system leads offline rehearsal and experience replay but trails ER-ACE and DER++.
An ablation finding: rotation carries most of the gain, and isolation adds the invariance guarantee. The authors also report that the mechanism is not tied to the local learning rule. Under the same schedule, a backpropagation network with k-WTA hidden layers gains from rotation, and isolation is again free on top of it.
Key facts
- The paper asks whether a network trained by local, biologically constrained rules can consolidate memories with no offline phase, mimicking the local sleep that brains use while awake.
- Three mechanisms do the work: an isolation rule that confines replay to synapses invisible to the current input, a refractory rotation rule, and homeostatic pressure plus a relative-novelty gate that time the replay bursts.
- On class-incremental split-MNIST the system reaches 91.6+-0.3% with no offline phase; in a single pass it scores 91.8% against 90.1% for DER++, while the offline-night schedule falls to 76.9%.
- On split CIFAR-10 it leads offline rehearsal and experience replay but trails ER-ACE and DER++.
- Rotation carries most of the gain; isolation adds the invariance guarantee, and rotation also helps a backpropagation network with k-WTA hidden layers.
Why it matters
Continual learning has to keep old knowledge while absorbing new data. Replay methods usually need a separate offline phase or replayed samples mixed into the live stream. This paper tests a different route taken from biology: consolidating during wakefulness, in the parts of the network the current input does not use. The authors describe it as inverting the usual direction of non-interfering continual learning, since the current computation is held invariant and old memories are written into unused degrees of freedom.
Who it affects
Mainly researchers working on continual learning, replay methods and biologically plausible learning rules. The results are on class-incremental split-MNIST and split CIFAR-10 only, so this is a research result rather than something aimed at practitioners today.
How to use it
No code release, compute cost or training time is mentioned in the abstract, so there is nothing to run yet. A researcher could take the recipe as described: k-winner-take-all hidden layers, an isolation mask for replay updates, refractory rotation, and homeostatic and novelty gates. The authors report that rotation, and isolation on top of it, also work with a backpropagation network that has k-WTA hidden layers.
How solid is it
This is an arXiv preprint, and the numbers are the authors' own. On split-MNIST the system is reported at 91.6+-0.3% with no offline phase, at or above the best offline-night schedule on two held-out splits and tied with DER++. The metric behind the split-MNIST percentages is not named. The ablation says rotation carries most of the gain, and the finding that the mechanism transfers to a k-WTA backpropagation network suggests it is not an artefact of the local rule.
Risks and caveats
The picture is mixed. The advantage is largest at small buffers and gives way to the backpropagation references at large ones. On split CIFAR-10 the system trails ER-ACE and DER++, leading only offline rehearsal and experience replay; no numbers are given for CIFAR-10. Buffer sizes are not specified, only small and large. Invariance is exact only on the proven channels; elsewhere it holds for all but 0.3% of waking samples per update. No claim is made about datasets beyond split-MNIST and split CIFAR-10.
“We ask whether a network trained by local, biologically constrained rules can consolidate with no offline phase at all.”
— From the paper's abstract