This is the official implementation of the SETFusion model proposed in the paper (SETFusion: A Semantic Transformer for Infrared and Visible Image Fusion) with Pytorch.
| Methods | End-to-End | Convolutional Operation | Pyramid Semantic Transformer | Multi-scale Semantic Transformer | VIF Loss | Unsupervised | Generalization Ability |
|---|---|---|---|---|---|---|---|
| GTF | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ |
| RFN-Nest | ✔ | ✔ | ✘ | ✘ | ✘ | ✘ | ✔ |
| PMGI | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ |
| FusionGAN | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ |
| MFEIF | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ |
| GANMcC | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ |
| PPT Fusion | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ | ✘ |
| DATFuse | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ |
| TCCFusion | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ |
| CrossFuse | ✘ | ✔ | ✘ | ✘ | ✘ | ✔ | ✘ |
| MMDRFuse | ✔ | ✔ | ✘ | ✘ | ✘ | ✔ | ✔ |
| SETFusion | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ | ✔ |
| Methods | Time (s) | Parameters (M) | FLOPs (G) |
|---|---|---|---|
| GTF | 3.4207 | / | / |
| RFN-Nest | 2.3096 | 19.17 | 7.68 |
| PMGI | 0.1934 | 0.04 | 1.69 |
| FusionGAN | 2.6796 | 0.93 | 0.55 |
| MFEIF | 0.3181 | 4.94 | 25.30 |
| GANMcC | 5.6752 | 1.86 | 1.02 |
| PPT Fusion | 0.4126 | 1.23 | 25.08 |
| DATFuse | 0.0254 | 0.01 | 1.21 |
| TCCFusion | 0.1220 | 0.19 | 27.08 |
| CrossFuse | 1.0636 | 20.64 | 11.32 |
| MMDRFuse | 0.0644 | 0.0004 | 0.14 |
| OmniFuse | 0.0742 | 18.08 | 13.50 |
| PIDFusion | 0.0462 | 0.05 | 208.60 |
| DDBFusion | 0.9775 | 5.86 | 184.93 |
| DCEvo | 0.2505 | 2.23 | 2336.42 |
| SETFusion | 0.2069 | 0.32 | 18.14 |
If this work is helpful to you, please cite it as:
@ARTICLE{Tang_2026_SETFusion,
author={Tang, Wei and He, Fazhi and Zhang, Lin and Zhao, Shengjie },
journal={Pattern Recognition},
title={SETFusion: A Semantic Transformer for Infrared and Visible Image Fusion},
year={2026},
volume={175},
number={},
pages={113130},
doi={10.1016/j.patcog.2026.113130}}
If you have any questions, feel free to contact me (weitang@tongji.edu.cn).













