Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

SCMM: Calibrating cross-modal representations for text-based person search

Liu, Jing, Wei, Donglai, Liu, Yang, Zhang, Sipeng, Yang, Tong, Zhou, Wei, Ding, Weiping and Leung, Victor C.M. 2027. SCMM: Calibrating cross-modal representations for text-based person search. Pattern Recognition 182 , 114775. 10.1016/j.patcog.2026.114775

Full text not available from this repository.

Abstract

Text-Based Person Search (TBPS) aims to retrieve target person images from a large-scale database using natural language descriptions, serving as a critical task in multimodal perception and visual pattern recognition. Bridging the semantic gap between heterogeneous modalities while capturing fine-grained correspondences remains a fundamental challenge, especially when discriminating visually similar individuals based on complex textual semantics. To address these challenges, we propose Sew Calibration and Masked Modeling (SCMM), a unified framework that calibrates cross-modal representations for effective multimodal visual-textual pattern matching. Concretely, SCMM introduces two principal components: a sew calibration loss that dynamically aligns image-text features via a quality-guided adaptive margin governed by textual information density, and a masked caption modeling loss that establishes fine-grained semantic correspondences through transformer-based masked prediction. The sew calibration mechanism imposes bidirectional constraints to compactly cluster same-identity features in a shared embedding space. Simultaneously, the masked modeling component acts as a cross-modal decoder that learns word-level representations, effectively discriminating subtle attribute differences. Importantly, our dual-encoder architecture strikes an optimal balance between representation expressiveness and computational efficiency by adopting a training-only decoder design. Extensive experiments on CUHK-PEDES, ICFG-PEDES, and RSTPReID datasets demonstrate that SCMM achieves state-of-the-art performance with Rank-1 accuracies of 73.81%, 64.25%, and 57.35%, respectively. Thorough ablation studies confirm the efficacy of each proposed mechanism in establishing robust cross-modal patterns for multimodal perception and recognition.

Item Type: Article
Date Type: Publication
Status: Published
Schools: Schools > Computational & Mathematical Sciences
Schools > Computer Science & Informatics
Publisher: Elsevier
ISSN: 0031-3203
Date of Acceptance: 18 August 2026
Last Modified: 14 Sep 2026 12:19
URI: https://orca.cardiff.ac.uk/id/eprint/189556

Actions (repository staff only)

Edit Item Edit Item