Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Accelerated sim-to-real deep reinforcement learning: learning collision avoidance from human player

Niu, Hanlin, Ji, Ze ORCID: https://orcid.org/0000-0002-8968-9902, Arvin, Farshad, Lennox, Barry, Yin, Hujun and Carrasco, Joaquin 2021. Accelerated sim-to-real deep reinforcement learning: learning collision avoidance from human player. Presented at: 2021 IEEE/SICE International Symposium on System Integration (SII2021), Iwaki, Fukushima, Japan, 11-14 January 2021. Proceedings of the IEEE/SICE International Symposium on System Integration. IEEE, pp. 144-149. 10.1109/IEEECONF49454.2021.9382693
Item availability restricted.

[thumbnail of Ji Z - Accelerated Sim-to-Real Deep Reinforcement ....pdf] PDF - Accepted Post-Print Version
Restricted to Repository staff only

Download (4MB)

Abstract

This paper presents a sensor-level mapless collision avoidance algorithm for use in mobile robots that map raw sensor data to linear and angular velocities and navigate in an unknown environment without a map. An efficient training strategy is proposed to allow a robot to learn from both human experience data and self-exploratory data. A game format simulation framework is designed to allow the human player to tele-operate the mobile robot to a goal and human action is also scored using the reward function. Both human player data and self-playing data are sampled using prioritized experience replay algorithm. The proposed algorithm and training strategy have been evaluated in two different experimental configurations: Environment 1, a simulated cluttered environment, and Environment 2, a simulated corridor environment, to investigate the performance. It was demonstrated that the proposed method achieved the same level of reward using only 16% of the training steps required by the standard Deep Deterministic Policy Gradient (DDPG) method in Environment 1 and 20% of that in Environment 2. In the evaluation of 20 random missions, the proposed method achieved no collision in less than 2 h and 2.5 h of training time in the two Gazebo environments respectively. The method also generated smoother trajectories than DDPG. The proposed method has also been implemented on a real robot in the real-world environment for performance evaluation. We can confirm that the trained model with the simulation software can be directly applied into the real-world scenario without further fine-tuning, further demonstrating its higher robustness than DDPG. The video and code are available: https://youtu.be/BmwxevgsdGc https://github.com/hanlinniu/turtlebot3_ddpg_collision_avoidance

Item Type: Conference or Workshop Item - published (Paper)
Date Type: Publication
Status: Published
Schools: Schools > Engineering
Additional Information: "© 2021 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works."
Publisher: IEEE
ISBN: 978-1-7281-7659-8
ISSN: 2474-231
Date of First Compliant Deposit: 8 February 2021
Date of Acceptance: 5 October 2020
Last Modified: 04 Aug 2026 09:30
URI: https://orca.cardiff.ac.uk/id/eprint/138349

Citation Data

Cited 11 times in Scopus. View in Scopus. Powered By Scopus® Data

Actions (repository staff only)

Edit Item Edit Item

Downloads

Downloads per month over past year

View more statistics