IEEE Transactions on Industrial Informatics (under review)
Certifiable robustness is a critical prerequisite for robotic autonomy in complex real-world environments, where systems inevitably suffer from diverse inherent uncertainties including unmodeled dynamics and sensor noise. To address the lack of formal stability guarantees in existing learning-based locomotion controllers, we propose a unified framework integrating neural Lyapunov functions with reinforcement learning.
Different from prior deterministic Lyapunov methods and Gaussian-only stochastic stability analysis, we model both dynamic mismatch and state estimation noise as sub-Gaussian random disturbances and derive rigorous probabilistic Lyapunov conditions with an explicit provable lower bound on system convergence probability. These stability constraints are embedded as regularizers into an end-to-end Actor–Critic pipeline to jointly optimize task tracking performance and certifiable robustness without requiring accurate analytical system models.
Our framework yields an expanded certified region of attraction and quantitative robustness metrics. Extensive validations are carried out across systems of rising complexity: inverted pendulum simulation, quadruped robot simulation, and full-sized Unitree bipedal humanoid with both simulation and physical hardware tests. Comparative results against classical control, deterministic neural Lyapunov, and state-of-the-art imitation RL baselines demonstrate superior disturbance resilience and stable gait recovery under locomotion tasks.
The proposed pipeline jointly trains a task-oriented policy and a neural Lyapunov certificate (Twin Control Lyapunov Function, TCLF) inside an Actor–Critic RL loop. A state-aware gating mechanism balances task rewards and Lyapunov robustness penalties, enlarging the certified region of attraction while preserving locomotion tracking performance.
Quick preview of stabilization rollouts under different controllers.
3D Lyapunov Landscape
Example 1
Example 2
Example 3
Example 4
Example 5
3D Lyapunov Landscape
Example 1
Example 2
Example 3
Example 4
Example 5
3D Lyapunov Landscape
Example 1
Example 2
Example 3
Example 4
Example 5
Simulation snapshots of robust locomotion on the Doggo platform.
Snapshot 1
Snapshot 2
Snapshot 3
Snapshot 4
Short video previews on Unitree G1 (simulation and hardware).
Nominal Walk
Push Example 1
Push Example 2
Gait Recovery
Nominal Walk
Push Example 1
Push Example 2
Gait Recovery
Full-length humanoid experiment reel (YouTube).
youtubeVideoId at the bottom of this page to embed it.
@article{robust_nlfrl2026,
title = {Certifiably Robust Reinforcement Learning via Probabilistic Neural Lyapunov Functions},
journal = {IEEE Transactions on Industrial Informatics},
year = {2026},
note = {under review}
}