[1] viXra:2608.0006 [pdf] submitted on 2026-08-02 00:34:14
Authors: Xiaohao Xie, Wei Meng, Wenhua Jiao
Comments: 16 Pages.
Person Re-Identification (ReID) under severe occlusion remains a formidable challenge, primarily because existing discriminative models lack the capacity to infer missing topological structures, inevitably leading to fragmented feature representations and compromised metric spaces. To break this perceptual limitation, we propose a novel end-to-end Diffusion-Driven Dual-stream Framework ($text{D}^3text{F}$), which seamlessly integrates generative structural priors from Diffusion Transformers (DiT) into vision-language ReID. To overcome the high computational overhead and semantic gaps inherent in diffusion models, we first design a truncated prior extraction mechanism alongside a Cascaded Inverted Modality Shuffle Network (CIMSN) to achieve lightweight and deep cross-modal interaction. Furthermore, recognizing that divergent generative features can severely contaminate the strict distance metric space, we propose an Asymmetric Decoupled Feature Guidance (ADFG) strategy. ADFG strictly utilizes the fused generative prior as a contextual lens to update text prompts, while reverting to pure discriminative features for mask localization and final metric pooling. This restrained decoupling strategy strikes an optimal balance between occlusion-resistant inference and metric space purity. Extensive experiments demonstrate that the proposed $text{D}^3text{F}$ significantly outperforms existing methods, achieving state-of-the-art (SOTA) performance on both occluded and holistic ReID benchmark datasets.
Category: Artificial Intelligence