Semantic modality compensation improves person re-identification across visible and infrared images

Semantic Modality Compensation for Unsupervised Visible-Infrared Person Re-identification under Unpaired Settings

Computer Vision and Pattern Recognition

Summary

Recognizing the same person in images taken by visible light and infrared cameras without knowing who they are is hard, especially when the images don’t match one-to-one. The paper proposes a new way to separate what makes a person unique from the differences caused by the type of camera. This method creates a shared understanding space and uses learned clues to fill in missing links between the two camera types, helping to better identify people even if the images are unpaired. The authors show their approach works better than existing methods, particularly when there are many unmatched identities across camera types.

What this means in practice

  • For security system engineers: Improve cross-camera person matching in surveillance systems using visible and infrared sensors, even when identities are incomplete or unpaired.
  • For robotics developers: Enhance person recognition capabilities of robots operating in diverse lighting by compensating for modality differences without labeled identity data.

Authors

Duanning Chen, Ke He, Bin Yang, Yongxiang Yao

Abstract

Unsupervised visible-infrared person re-identification (USL-VI-ReID) learns person representations that can be compared across modalities without identity annotations. In the unpaired setting, however, identity correspondences between modalities are often incomplete, leaving many identities without an observed counterpart in the other modality. Existing unpaired methods bridge this gap by generating or mapping features for the other modality, mainly by exploiting the statistics of visual features without explicitly separating content that is discriminative for identity from style that is specific to modality. Consequently, the generated features may distort identity cues or inherit bias from the source modality, undermining the reliability of supervision across modalities. We formulate unpaired learning across modalities as a semantic compensation problem and propose Semantic Modality Compensation (SMC), a framework based on prompt composition that decouples identity semantics from modality style within a shared visual semantic space. SMC first constructs a discriminative ReID space through augmented dual contrastive learning, yielding pseudo labels, cluster prototypes, and memory banks for each modality. It then learns visible and infrared modality prompts in the CLIP semantic space and maps clusters obtained from pseudo labels to identity semantic tokens. For each cluster lacking a reliable match in the other modality, SMC combines its identity token with the prompt for the target modality to synthesize a semantic counterpart in the missing modality. The synthesized counterpart is then projected back into the ReID space and injected into a compensation memory through confidence gating. Extensive experiments under both paired and unpaired settings demonstrate that SMC consistently outperforms state-of-the-art methods, with particularly large gains when identity mismatch is severe.