CLBF-Constrained Curious Actor-Critic: Safe and Exploratory Reinforcement Learning for Embodied Robotic Systems

Authors

  • Malathy Batumalay Faculty of Data Science and IT, INTI International University, Nilai Author
  • Harikrishnan A/L Ramiah Department of Electrical Engineering, Faculty of Engineering, Universiti Malaya, Kuala Lumpur Author
  • Nithesh Naik Jesselton University College, Kota Kinabalu, Sabah Author

Keywords:

Safe Reinforcement Learning, Control Lyapunov-Barrier Functions (CLBF), Embodied Robotics, Curiosity-Driven Exploration, Soft Actor-Critic (SAC)

Abstract

We propose a novel reinforcement learning framework for embodied robotic systems that integrates safety guarantees with curiosity-driven exploration through a Control Lyapunov-Barrier Function (CLBF) constrained actor-critic architecture. The proposed method, termed C3AC, addresses the fundamental tension between safe operation and exploratory behavior in physical robots. The critic network is augmented to output three distinct value components: a standard action-value function, a Lyapunov certificate that enforces convergence to goal states, and a barrier certificate that maintains forward invariance of a predefined safe set. These certificates are learned jointly with a forward dynamics model, and their satisfaction is enforced through differentiable penalty terms in the critic loss. To encourage exploration of novel yet safe regions, an ensemble of five forward dynamics models is maintained, and the epistemic uncertainty across this ensemble serves as an intrinsic reward signal. A dual Lagrangian network adaptively weights this curiosity bonus based on the current safety margin, ensuring that exploration is only encouraged when Lyapunov constraints are satisfied. The actor network outputs nominal actions that are subsequently projected onto a provably safe set via a differentiable quadratic program solved in real time at 50 Hz. This projection step explicitly enforces the Lyapunov decrease condition and barrier constraint, thereby guaranteeing safety during execution. The entire system operates in a closed loop with off-policy training using Soft Actor-Critic updates. The primary contributions of this work are threefold: first, the embedding of dual safety certificates directly into the critic architecture; second, the coupling of ensemble-based intrinsic motivation with safety-constrained action projection; and third, the real-time differentiable optimization that backpropagates through the safety projection to update the actor. Experimental evaluations on a UR5 robotic arm demonstrate that C3AC achieves superior task completion rates while maintaining zero safety violations during training, outperforming both standard SAC and prior safety-constrained methods in terms of sample efficiency and exploration quality.

References

J Hughes, A Abdulali, R Hashem, et al. Embodied artificial intelligence: Enabling the next intelligence revolution. In IOP Conference Series: Materials Science and Engineering, 2022.

RS Sutton and AG Barto. Reinforcement learning: An introduction. Technical report, cambridge.org, 1998.

N Roy, I Posner, T Barfoot, P Beaudoin, et al. From machine learning to robotics: Challenges and opportunities for embodied intelligence. Technical report, arXiv preprint arXiv:2110.15245, 2021.

P Wang, H Li, and CY Chan. Continuous control for automated lane change behavior based on deep deterministic policy gradient algorithm. In 2019 IEEE Intelligent Vehicles Symposium (IV), 2019.

T Haarnoja, A Zhou, P Abbeel, et al. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In International Conference on Machine Learning, 2018.

J Achiam, D Held, A Tamar, et al. Constrained policy optimization. In International Conference on Machine Learning, 2017.

A Stooke, J Achiam, and P Abbeel. Responsive safety in reinforcement learning by pid lagrangian methods. In International Conference on Machine Learning, 2020.

D Pathak, P Agrawal, AA Efros, et al. Curiosity-driven exploration by self-supervised prediction. In 2017 IEEE Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2017.

B Lakshminarayanan, A Pritzel, et al. Simple and scalable predictive uncertainty estimation using deep ensembles. In Advances in Neural Information Processing Systems, 2017.

AI Doban and M Lazar. Computation of lyapunov functions for nonlinear differential equations via a massera-type construction. IEEE Transactions on Automatic Control, 2017.

AD Ames, JW Grizzle, and P Tabuada. Control barrier function based quadratic programs with application to adaptive cruise control. In 53rd IEEE Conference on Decision and Control, 2014.

X Xu, P Tabuada, JW Grizzle, and AD Ames. Robustness of control barrier functions for safety critical control. IFAC-PapersOnLine, 2015.

R Quirynen, S Safaoui, et al. Real-time mixed-integer quadratic programming for vehicle decision-making and motion planning. IEEE Transactions on Intelligent Vehicles, 2024.

E Altman. Constrained markov decision processes. Technical report, taylorfrancis.com, 2021.

G Dalal, K Dvijotham, M Vecerik, T Hester, et al. Safe exploration in continuous action spaces. Technical report, arXiv preprint arXiv:1801.08757, 2018.

R Cheng, G Orosz, RM Murray, and JW Burdick. End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, 2019.

Y Burda, H Edwards, D Pathak, A Storkey, et al. Large-scale study of curiosity-driven learning. Technical report, arXiv preprint arXiv:1808.04355, 2018.

W Xiao, L Yuan, T Ran, L He, et al. Csper: Curiosity-driven safety-prioritized experience replay for mapless navigation using deep reinforcement learning. IEEE Transactions on Cognitive and Developmental Systems, 2026.

YC Chang, N Roohi, and S Gao. Neural lyapunov control. In Advances in Neural Information Processing Systems, 2019.

B Tearle, KP Wabersich, A Carron, et al. A predictive safety filter for learning-based racing control. IEEE Robotics and Automation Letters, 2021.

G Huang, Y Li, G Pleiss, Z Liu, JE Hopcroft, et al. Snapshot ensembles: Train 1, get m for free. Technical report, arXiv preprint arXiv:1704.00109, 2017.

Y Gal and Z Ghahramani. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In International Conference on Machine Learning, 2016.

JJ Choi, D Lee, K Sreenath, CJ Tomlin, et al. Robust control barrier-value functions for safety-critical control. In 2021 60th IEEE Conference on Decision and Control (CDC), 2021.

H Rahimian and S Mehrotra. Distributionally robust optimization: A review. Technical report, arXiv preprint arXiv:1908.05659, 2019.

C Zhang, S Bengio, M Hardt, B Recht, et al. Understanding deep learning requires rethinking generalization. Technical report, arXiv preprint arXiv:1611.03530, 2016.

A ArjomandBigdeli, A Mata, and S Bak. Verification of neural network control systems in continuous time. Lecture Notes in Computer Science, 2024.

A Abate, C David, P Kesseli, D Kroening, et al. Counterexample guided inductive synthesis modulo theories. Lecture Notes in Computer Science, 2018.

International Organization for Standardization. Iso 10218: Robots and robotic devices: Safety requirements for industrial robots. Technical report, International Organization for ..., 2011.

A Porras-Vázquez and JA Romero-Pérez. A new methodology for facilitating the design of safety-related parts of control systems in machines according to iso 13849: 2006 standard. Reliability Engineering & System Safety, 2018.

X Zhou, B Chen, Y Gui, and L Cheng. Conformal prediction: A data perspective. ACM computing surveys, 2025.

GC Calafiore and MC Campi. The scenario approach to robust control design. IEEE Transactions on Automatic Control, 2006.

D Panagou, DM Stipanović, et al. Distributed coordination control for multi-robot networks using lyapunov-like barrier functions. IEEE Transactions on Automatic Control, 2015.

Downloads

Published

2026-09-30

Issue

Section

Articles