Teacher-guided corrective residuals support reflex-like recovery in simulated locomotion
Abstract¶
Animals do not recover from every stumble by relearning a motor policy. Recovery is more naturally described as layered control: an ongoing motor command is combined with faster corrective pathways shaped by error and practice. We tested this idea in BrainCAD, a biologically inspired neural-control toolkit built around MuJoCo locomotion experiments. Several reinforcement-learning-only corrective circuits, including spinal gain gating and residual modules, did not learn reliable fast recovery under PPO. A cerebellar-inspired residual became effective only when trained with a direct teacher signal. This is close to residual policy learning and policy distillation Johannink et al., 2018Rusu et al., 2015Schmitt et al., 2018; the contribution is not a new distillation algorithm, but a controlled demonstration that this learning signal was necessary in our biologically inspired locomotor-control architecture. In Humanoid-v4, the same trained checkpoint was evaluated with the residual switched on and off under identical perturbation schedules. Across 20 seeds, the residual reduced fall AUC by 0.361 on average and increased return AUC by 74.6. The correction could also be distilled into a cortex-only policy, preserving most of the fall reduction. Cross-body experiments showed strong effects in Humanoid, Walker2d, and Ant, with Hopper as a boundary case where stability gains were less clean and return often decreased. These results suggest that biologically inspired corrective modules require an appropriately structured learning signal, not just an anatomically motivated architecture.
Claim. In BrainCAD/ReflexBench locomotion experiments, cerebellar-inspired corrective modules improved reflex-like recovery only when trained with a structured teacher signal.
Reflexes are often described as simple loops, but real movement is not simple. A body with many joints and contacts has many ways to fail: a foot can slip, sensory input can become unreliable, or a shove can arrive while the controller is already committed to a step. Nervous systems solve this with layered circuits. Spinal circuits provide fast sensorimotor pathways, cortex supports flexible voluntary control, and the cerebellum is strongly associated with internal models and error-based motor adaptation Wolpert et al., 1998Ito, 2008Seidler et al., 2013.
Motivated by this organization, we implemented several corrective circuits in BrainCAD BrainCAD Project, 2026 and tested them in ReflexBench, a perturbation suite for MuJoCo locomotion. The first lesson was negative: PPO alone Schulman et al., 2017 often failed to discover a small, fast corrective action, even when the architecture was biologically plausible. The successful variant used a teacher. A robust policy was first trained under perturbations, and a cerebellar residual was then trained to imitate the teacher’s correction relative to a nominal cortex, following the broad logic of residual policy learning and distillation Johannink et al., 2018Rusu et al., 2015Schmitt et al., 2018. We therefore frame the result as a learning-signal result, not as evidence for a new principle of cerebellar biology.
The model has three pieces. A cortex policy proposes a base action,
where is the observation. A teacher policy, trained separately with perturbations, produces . During imitation, the cerebellar residual is trained toward
At evaluation time, the residual controller executes
where is an optional engagement gate. A later consolidation step removes the residual machinery by training a new cortex to imitate the final action directly.
Figure 1:Teacher-guided residual control and main results. A, a perturbation-trained teacher supplies a corrective target for a cerebellar residual. B, in Humanoid-v4, residual ON vs OFF evaluation with the same checkpoint reduces fall AUC across pushes, sensor corruption, and slips (N=20). C, the same comparison improves return AUC. D, cross-body results are suggestive of a morphology effect: contact-rich bodies benefit strongly, while Hopper remains a boundary condition. Points are mean estimates without confidence regions and are shown only as an exploratory morphology summary.
RL-only correction was not enough. We first tested spinal gain gating, residual correction trained only by PPO, curriculum perturbations, basal-ganglia-like engagement gates, and stability auxiliary terms. These variants were useful diagnostics, but none gave a reliable causal reduction in fall AUC. Thus, the circuit diagram alone was not the main explanation. The key variable was whether the module received enough information to learn a useful correction at the right time.
Teacher-guided residuals improved survival. The teacher-guided residual was tested causally: the same checkpoint and the same perturbation schedule were evaluated with the residual enabled or disabled. In the N=20 Humanoid campaign, fall AUC decreased for push (-0.313, 95% CI [-0.491, -0.147]), sensor (-0.447, 95% CI [-0.553, -0.338]), and slip (-0.325, 95% CI [-0.511, -0.155]) perturbations, where negative numbers mean fewer falls. Return AUC increased for push (58.8, 95% CI [27.6, 88.4]), sensor (103.6, 95% CI [82.7, 123.5]), and slip (61.4, 95% CI [27.4, 91.6]). Perturbation sanity checks verified that pushes, friction changes, and sensor corruption were actually applied, not only logged.
Teacher quality mattered. A strong teacher produced stabilizing residuals. A deliberately weak teacher produced residuals that could become harmful. This argues against a simple capacity explanation: the student learned a direction of correction from the teacher, and a bad teacher could teach the wrong direction.
Corrections could be consolidated. We trained a cortex-only policy to imitate the final action of the residual controller. This consolidated policy no longer used an online cerebellar residual, but it retained most of the fall reduction in Humanoid and generalized under a long-push schedule. This suggests that the residual pathway can act as a training scaffold: useful during learning, but not always necessary at inference.
Exploratory morphology analysis should be interpreted cautiously. Ant and Humanoid showed strong fall reductions, Walker2d replicated, and Hopper exposed a stability-performance tradeoff. This supports a modest morphology-dependent control hypothesis: layered corrections may become more useful as the body creates more recovery states for the controller to handle, but four simulated bodies cannot establish an evolutionary scaling law.
Methods in brief¶
Experiments used BrainCAD with MuJoCo locomotion environments Todorov et al., 2012Brockman et al., 2016. ReflexBench applied external pushes, friction reductions, and observation corruption. Fall AUC was computed as the area under the fall-probability curve across the predefined perturbation-severity grid: for Humanoid, push magnitudes 0-1 crossed with duration, slip severities 0-1.25 crossed with patch settings, and sensor dropout bursts 0-0.4. Return AUC was computed analogously from mean return across the same severity grid. Curves were computed per seed before paired ON-OFF deltas were summarized with bootstrap confidence intervals. Lower fall AUC is better, while higher return AUC is better. Source notebooks and validation summaries are included in the micropublication package.
Acknowledgments¶
This work was developed as part of the Neuromatch Impact Scholars Program by Team RNNematode. We thank Raymond Chua for senior mentorship, scientific guidance, and feedback on motor-control framing. We also acknowledge the BrainCAD/NCAP software infrastructure used for circuit templates, MuJoCo wrappers, evaluation, and video export.
Data Availability¶
The submitted code package includes notebooks that reproduce the main figure from saved result artifacts, validation tables, and instructions for accessing the saved BrainCAD/ReflexBench results. Published via Impact Scholars; original development repository.
- Wolpert, D. M., Miall, R. C., & Kawato, M. (1998). Internal Models in the Cerebellum. Trends in Cognitive Sciences, 2(9), 338–347. 10.1016/S1364-6613(98)01221-2
- Ito, M. (2008). Control of Mental Activities by Internal Models in the Cerebellum. Nature Reviews Neuroscience, 9(4), 304–313. 10.1038/nrn2332
- Seidler, R. D., Kwak, Y., Fling, B. W., & Bernard, J. A. (2013). Neurocognitive Mechanisms of Error-Based Motor Learning. Advances in Experimental Medicine and Biology, 782, 39–60. 10.1007/978-1-4614-5465-6_3
- BrainCAD Project. (2026). BrainCAD: Biologically Inspired Neural Control Toolkit. Software package included in the project repository.
- Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal Policy Optimization Algorithms. https://arxiv.org/abs/1707.06347
- Johannink, T., Bahl, S., Nair, A., Luo, J., Kumar, A., Loskyll, M., Ojea, J. A., Solowjow, E., & Levine, S. (2018). Residual Reinforcement Learning for Robot Control. https://arxiv.org/abs/1812.03201
- Rusu, A. A., Colmenarejo, S. G., Gulcehre, C., Desjardins, G., Kirkpatrick, J., Pascanu, R., Mnih, V., Kavukcuoglu, K., & Hadsell, R. (2015). Policy Distillation. https://arxiv.org/abs/1511.06295
- Schmitt, S., Hudson, J. J., Zidek, A., Osindero, S., Doersch, C., Czarnecki, W. M., Leibo, J. Z., Kuttler, H., Zisserman, A., Simonyan, K., & Eslami, S. M. A. (2018). Kickstarting Deep Reinforcement Learning. https://arxiv.org/abs/1803.03835
- Todorov, E., Erez, T., & Tassa, Y. (2012). MuJoCo: A Physics Engine for Model-Based Control. 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, 5026–5033. 10.1109/IROS.2012.6386109
- Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., & Zaremba, W. (2016). OpenAI Gym. https://arxiv.org/abs/1606.01540