Cardiff University | Prifysgol Caerdydd ORCA
Online Research @ Cardiff 
WelshClear Cookie - decide language by browser settings

Uncertainty-guided spatiotemporal consistency fusion network for infrared-visible video fusion under extremely low-light conditions

Zhao, Cheng, Song, Tianyun, Wu, Zhiliang, Wang, Tianfu, Gabbouj, Moncef, Yue, Guanghui, Lei, Baiying and Zhou, Wei 2026. Uncertainty-guided spatiotemporal consistency fusion network for infrared-visible video fusion under extremely low-light conditions. IEEE Transactions on Image Processing 35 , pp. 8894-8909. 10.1109/tip.2026.3719477

Full text not available from this repository.

Abstract

Infrared–visible video fusion under extremely low-light conditions is critically important yet remains underexplored, largely due to the scarcity of high-quality datasets and challenges posed by spatiotemporal uncertainty and modality bias. To address the dataset shortage, we built a dataset of 4,739 infrared and visible registration video pairs captured under extremely low-light conditions, spanning 5 scene types and 17 subcategories. Further, we proposed an Uncertainty-guided Spatiotemporal Consistency Fusion Network, termed USCFNet, for the infraredvisible video fusion. At each layer of the encoder, an Entropy-Gated SpatioTemporal Attention (EGSTA) module is introduced to capture temporal instability and spatial reliability variations through entropy-aware attention modulation, thereby enhancing feature spatiotemporal consistency. The refined infrared and visible features are then fused via a Difference-Guided Fusion (DGF) module, which adaptively exploits their content and edge differences to improve structural integrity and detail clarity. By progressively connecting DGF modules from shallow to deep layers, the network achieves the synergistic fusion of shallow textures and deep semantics. Subsequently, the output of the last DGF module is fused with the modality features of the last layer through a hierarchical mixture-of-experts fusion module. This module enables the balanced integration of modality information while preserving fine local details. Finally, the fusion feature is fed into the decoder to produce the final fused video. Extensive experiments on our dataset and two public datasets show that USCFNet outperforms competing methods, achieving lower distortion and stronger spatiotemporal consistency. The source code and dataset are available at https://github.com/Zhaocheng1/ELVID.

Item Type: Article
Date Type: Published Online
Status: Published
Schools: Schools > Computational & Mathematical Sciences
Schools > Computer Science & Informatics
Publisher: Institute of Electrical and Electronics Engineers (IEEE)
ISSN: 1057-7149
Last Modified: 02 Sep 2026 14:39
URI: https://orca.cardiff.ac.uk/id/eprint/189032

Actions (repository staff only)

Edit Item Edit Item