| Liu, Jing, Wei, Donglai, Liu, Yang, Zhang, Sipeng, Yang, Tong, Zhou, Wei, Ding, Weiping and Leung, Victor C.M. 2027. SCMM: Calibrating cross-modal representations for text-based person search. Pattern Recognition 182 , 114775. 10.1016/j.patcog.2026.114775 |
Abstract
Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale database using natural language descriptions, serving as a critical task in multimodal perception and visual pattern recognition. Bridging the semantic gap between heterogeneous modalities while capturing fine-grained correspondences remains a fundamental challenge, especially when discriminating visually similar individuals based on complex textual semantics. To address these challenges, we propose Sew Calibration and Masked Modeling (SCMM), a unified framework that calibrates cross-modal representations for effective multimodal visual-textual pattern matching. Concretely, SCMM introduces two principal components: a sew calibration loss that dynamically aligns image-text features via a quality-guided adaptive margin governed by textual information density, and a masked caption modeling loss that establishes fine-grained semantic correspondences through transformer-based masked prediction. The sew calibration mechanism imposes bidirectional constraints to compactly cluster same-identity features in a shared embedding space. Simultaneously, the masked modeling component acts as a cross-modal decoder that learns word-level representations, effectively discriminating subtle attribute differences. Importantly, our dual-encoder architecture strikes an optimal balance between representation expressiveness and computational efficiency by adopting a training-only decoder design. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReID datasets demonstrate that SCMM achieves state-of-the-art performance with Rank-1 accuracies of 73.81%, 64.25%, and 57.35%, respectively. Thorough ablation studies confirm the efficacy of each proposed mechanism in establishing robust cross-modal patterns for multimodal perception and recognition.
| Item Type: | Article |
|---|---|
| Date Type: | Publication |
| Status: | Published |
| Schools: | Schools > Computational & Mathematical Sciences Schools > Computer Science & Informatics |
| Publisher: | Elsevier |
| ISSN: | 0031-3203 |
| Date of Acceptance: | 18 August 2026 |
| Last Modified: | 14 Sep 2026 12:19 |
| URI: | https://orca.cardiff.ac.uk/id/eprint/189556 |
Actions (repository staff only)
![]() |
Edit Item |




Altmetric
Altmetric