Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Multi-agent policy sharing and safety-aware action correction for multi-USV cooperative encirclement

Wu, Yixian, Tian, Shunyu, Tao, Weiyu, Ji, Ze ORCID: https://orcid.org/0000-0002-8968-9902 and Wei, Changyun 2026. Multi-agent policy sharing and safety-aware action correction for multi-USV cooperative encirclement. IEEE Internet of Things Journal 10.1109/JIOT.2026.3709862

Full text not available from this repository.

Abstract

Multi-USV cooperative encirclement faces complex dynamic environments and high collision risks caused by dense interactions among agents. Although multi-agent reinforcement learning (MARL) provides a promising solution to this problem, existing methods still suffer from imbalanced learning among agents, slow convergence under sparse rewards, and insufficient safety guarantees during execution. In encirclement tasks, these issues are often mutually coupled. Unstable learning signals may weaken policy coordination, while insufficiently coordinated behaviors may further increase collision risks during dense interactions. To alleviate these problems, this paper proposes a novel framework named PSAC-MATD3 (Policy Sharing and Action Correction MATD3). For cooperative learning, an adaptive recovery mechanism (ARM)-based policy sharing method is introduced to mitigate learning imbalance among agents and avoid performance degradation caused by negative transfer. For learning stability, a phase-aware reward shaping strategy is designed to address the sparse-reward problem and is further combined with a hybrid prioritized experience replay mechanism. This mechanism integrates temporal-difference (TD) errors with a cooperative encirclement priority (CEP) metric to balance short-term value correction and long-term formation stability. To reduce collision risks in the simulated encirclement process, a control barrier function (CBF) is incorporated at the execution layer to correct velocity commands in real time. Simulation results show that the proposed method achieves better overall performance than the baseline methods under the tested settings. PSAC-MATD3 obtains a task success rate of 90.80% and reduces collision events in the simulated dense-interaction scenarios. In addition, the learning-layer improvements exhibit relatively stable encirclement performance under different maximum evader speeds and initial configurations. The source code is available at https://github.com/changyunwei/multi usv PSAC.

Item Type: Article
Date Type: Published Online
Status: In Press
Schools: Schools > Engineering
Publisher: Institute of Electrical and Electronics Engineers
ISSN: 2327-4662
Last Modified: 13 Jul 2026 11:45
URI: https://orca.cardiff.ac.uk/id/eprint/188161

Actions (repository staff only)

Edit Item Edit Item