Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Physics-informed resilient autonomous racing via deep reinforcement learning

Cai, Boliang 2025. Physics-informed resilient autonomous racing via deep reinforcement learning. PhD Thesis, Cardiff University.
Item availability restricted.

[thumbnail of Physics_informed_resilient_autonomous_racing_via_deep_reinforcement_learning.pdf]
Preview
PDF - Accepted Post-Print Version
Download (44MB) | Preview
[thumbnail of Cardiff University Electronic Publication Form] PDF (Cardiff University Electronic Publication Form) - Supplemental Material
Restricted to Repository staff only

Download (265kB)

Abstract

While robust model-based control methods can effectively handle sys tem uncertainties, they often face a strict trade-off between safety and performance, becoming overly conservative when confronted with highly nonlinear, time-variant dynamics such as high-speed racing under un predictable friction. Conversely, learning-based alternatives offer power ful data-driven adaptability to complex environments, but struggle with data inefficiency, poor generalisation to out-of-distribution dynamics, lack of formal safety guarantees, and high computational costs. This thesis addresses these limitations by bridging physics-based principles and data driven methods through four targeted contributions. First, for decentralised multi-robot navigation in unstructured envi ronments, a Multi-Featured Policy Gradient (MFPG) architecture is pro posed. By decomposing rewards and incorporating a “social-norm reward bias,” this method significantly enhances training stability and safety dur ing robot interaction. Second, to address the challenges of dynamic stability and computa tional efficiency in high-speed autonomous racing under varying friction coefficients, this work proposes a Deep-Dynamics-Mediated Reinforce ment Learning (DDM-RL) framework. To improve the data efficiency of RL algorithms, a Physics-Informed Dynamics Estimator and Predictor (PIDEP) for state estimation and a Preference-Informed Value Estima tion (PrI-VE) module are proposed. To alleviate the computing burden of the onboard controller, a sparse control strategy is proposed. By combining these components, the framework ensures real-time resilience and stability in uncertain environments while effectively reducing computa tional load. Third, to address safety issues during the reinforcement learning ex ploration phase under unpredictable friction, a unified framework named Acting on the Tangent Space of the Probabilistic Constraint Manifold (ATAPCM) is proposed. Within this framework, the components func tion as a complementary pipeline. First, a friction coefficient estimator combining a Bayesian Physics-Informed Neural Network (B-PINN) and a Kalman filter evaluates time-variant environmental dynamics to pro vide a worst-case friction assessment. Second, these probabilistic esti mates are fed into the ATAPCM command correction module. As the RLagent generates exploratory control commands, this module evaluates them against the estimated friction limits. If an action risks instability, ATAPCM mathematically projects it onto a safe probabilistic manifold. Simulation results demonstrate that this integrated pipeline significantly reduces training costs and constraint violations compared to existing rein forcement learning-based methods, providing a provable safety guarantee for autonomous racing. Finally, based on the above results, another trial of physics-informed RL with differentiable physics, namely, Differentiable Dynamic Predicted Policy Gradient (DDPPG), is introduced. A differentiable dynamics pre dictor is incorporated into an RL framework using differentiable planning to optimise control policies for high-speed manoeuvres. Numerical simulations confirm that this approach improves training efficiency and policy resilience to parameter variations compared to standard RL methods.

Item Type: Thesis (PhD)
Date Type: Completion
Status: Unpublished
Schools: Schools > Engineering
Uncontrolled Keywords: 1. Deep Reinforcement Learning; 2. Autonomous Mapless Navigation; 3. Physics-Informed Machine Learning; 4. Autonomous Racing; 5. Constraint Manifold; 6. Safety-critical control; 7. Resilient Control; 8. Varying friction coefficient; 9. Differentiable simulation;
Funders: China Scholar Council (CSC202107610026)
Date of First Compliant Deposit: 19 May 2026
Last Modified: 19 May 2026 14:08
URI: https://orca.cardiff.ac.uk/id/eprint/187110

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics