FSD-Net: FSD-Net: Multi-Scale Outdoor Fire and Smoke Detection via Swin Transformer-Enhanced YOLOv11s with Lightweight Feature Fusion and Inner-CIoU Loss
Abstract
In this paper, a multi-scale fire and smoke detection model, named FSD-Net, is established upon the YOLOv11s model to overcome the challenges in outdoor fire scenarios, such as large variations in target shapes, complicated backgrounds, and the difficulty of detecting small objects. First, a multi-scale feature extraction module, C3Swin, is designed by integrating the Swin Transformer into the backbone network. This strengthens the model’s capability to capture fine-grained fire and smoke features. Second, we incorporate shallow features into the FPN-PAN architecture to improve small-object detection performance. Third, we developed a lightweight fusion module, GS-A2C2f, to enhance the model’s feature fusion capabilities.. Furthermore, GhostConv is employed to decrease computational complexity. An auxiliary bounding-box loss function, Inner-CIoU, is introduced to better balance detection performance and inference speed. Experimental results on a self-built dataset comprising 11,672 images evaluated following the COCO metrics, suggest that, compared to baseline model, FSD-Net achieves improvements of 2.5%, 2.3%, and 1.1% in precision, recall, and mAP@50, respectively, while lessening the number of parameters by 1.3 million and increasing FPS by 3.8%. These results confirm that FSD-Net achieves accurate fire smoke detection and localization, providing an effective solution for outdoor fire detection in complex environments.References
[1] J. P. Rafferty. (2025). Los Angeles wildfires of 2025.
Available: https://www.britannica.com/event/Los-
Angeles-wildfires-of-2025
[2] C. Jin et al., "Video Fire Detection Methods Based
on Deep Learning: Datasets, Methods, and Future D
irections," Fire, vol. 6, no. 8, 2023.https://doi.org/10
.3390/fire6080315
[3] S. Ren, K. He, R. Girshick, and J. Sun, "Faster R-C
NN: Towards Real-Time Object Detection with Region Proposal Networks," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 39, no. 6, pp. 1137-1149, 2017. https://doi.org/10.1109/TPAMI.2016.2577031
[4] K. He, G. Gkioxari, P. Dollár, and R. Girshick, "Mask R-CNN," in 2017 IEEE International Conference on Computer Vision (ICCV), 2017, pp. 2980-2988. https://doi.org/10.1109/ICCV.2017.322
[5] J. Redmon, S. Divvala, R. Girshick, and A. Farhadi, "You Only Look Once: Unified, Real-Time Object Detection," in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 779-788. https://doi.org/10.1109/CVPR.2016.91
[6] B. W. Xiao and C. M. Yan, "A lightweight global awareness deep network model for flame and smoke detection," (in English), Optoelectronics Letters, vol. 19, no. 10, pp. 614-622, Oct 2023. https://doi.org/10.1007/s11801-023-3041-x
[7] F. Y. Chen, M. Yang, and Y. Wang, "Recognition of Forest Fire Smoke Based on Improved YOLOv8n Model," FIRE TECHNOLOGY, vol. 61, no. 5, pp. 3351-3374, SEP 2025. https://doi.org/10.1007/s10694-025-01733-x
[8] C. M. Zhao, L. K. Zhao, K. Zhang, Y. H. Ren, H. Chen, and Y. H. Sheng, "Smoke and Fire-You Only Look Once: A Lightweight Deep Learning Model for Video Smoke and Flame Detection in Natural Scenes," Fire-Switzerland, vol. 8, no. 3, Mar 4 2025, Art. no. 104. https://doi.org/10.3390/fire8030104
[9] M. H. Ashraf, F. Jabeen, H. Alghamdi, M. S. Zia, and M. S. Almutairi, "HVD-Net: A Hybrid Vehicle Detection Network for Vision-Based Vehicle Tracking and Speed Estimation," Journal of King Saud University - Computer and Information Sciences, vol. 35, no. 8, p. 101657, 2023/09/01/ 2023. https://doi.org/https://doi.org/10.1016/j.jksuci.2023.101657
[10] Y. Liu et al., "A Survey of Visual Transformers," IEEE Transactions on Neural Networks and Learning Systems, vol. 35, no. 6, pp. 7478-7498, 2024. https://doi.org/10.1109/TNNLS.2022.3227717
[11] Y. Zheng, G. Zhang, S. Q. Tan, Z. G. Yang, D. X. Wen, and H. S. Xiao, "A forest fire smoke detection model combining convolutional neural network and vision transformer," FRONTIERS IN FORESTS AND GLOBAL CHANGE, vol. 6, APR 17 2023, Art. no.1136969. https://doi.org/10.3389/ffgc.2023.1136969
[12] B. S. Sun and X. Cheng, "Smoke Detection Transformer: An Improved Real-Time Detection Transformer Smoke Detection Model for Early Fire Warning," FIRE-SWITZERLAND, vol. 7, no. 12, DEC 2024, Art. no. 488.https://doi.org/10.3390/fire7120488
[13] L. Wang, H. Li, F. Siewe, W. Ming, and H. Li, "Forest fire detection utilizing ghost Swin transformer with attention and auxiliary geometric loss," Digital Signal Processing, vol. 154, 2024. https://doi.org/10.1016/j.dsp.2024.104662
[14] M. E. Qureshi, M. H. Ashraf, M. W. Arshad, A. Khan, H. Ali, and Z. U. Abdeen, "Hierarchical Feature Fusion With Inception V3 for Multiclass Plant Disease Classification," Informatica, vol. 49, no. 27, 2025. https://doi.org/10.31449/inf.v49i27.8208
[15] D. X. Yin, P. L. Cheng, and Y. Huang, "YOLO-EPF: Multi-scale smoke detection with enhanced pool former and multiple receptive fields," DIGITAL SIGNAL PROCESSING, vol. 149, JUN 2024, Art. no. 104511.https://doi.org/10.1016/j.dsp.2024.104511
[16] Y. X. Li et al., "Improving Fire and Smoke Detection with You Only Look Once 11 and Multi-Scale Convolutional Attention," FIRE-SWITZERLAND, vol. 8, no. 5, APR 22 2025, Art. no. 165. https://doi.org/10.3390/fire8050165
[17] D. Gragnaniello, A. Greco, C. Sansone, and B. Vento, "Fire and smoke detection from videos: A literature review under a novel taxonomy," Expert Systems with Applications, vol. 255, 2024. https://doi.org/10.1016/j.eswa.2024.124783
[18] R. Khanam and M. Hussain, "YOLOv11: An Overview of the Key Architectural Enhancements," arXiv preprint arXiv:2410.17725, 2024.
[19] H. Zhao et al., "Psanet: Point-wise spatial attention network for scene parsing," in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 267-283. https://doi.org/10.1007/978-3-030-01240-3_17
[20] T. Y. Lin, P. Dollár, R. Girshick, K. He, B. Hariharan, and S. Belongie, "Feature Pyramid Networks for Object Detection," in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 936-944. https://doi.org/10.1109/CVPR.2017.106
[21] H. Li, P. Xiong, J. An, and L. Wang, “Pyramid Attention Network for Semantic Segmentation,” in Proc. British Machine Vision Conference (BMVC), Newcastle, UK, Sep. 2018, p. 285. https://doi.org/10.2139/ssrn.5008171
[22] F. Chollet, "Xception: Deep learning with depthwise separable convolutions," in Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 1251-1258. https://doi.org/10.1109/cvpr.2017.195
[23] Z. Liu et al., "Swin transformer: Hierarchical vision transformer using shifted windows," in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 10012-10022. https://doi.org/10.1109/iccv48922.2021.00986
[24] D. M. Wang, Y. Qian, J. Y. Lu, P. Wang, Z. R. Hu, and Y. K. Chai, "Fs-yolo: fire-smoke detection based on improved YOLOv7," (in English), Multimedia Systems, vol. 30, no. 4, Aug 2024. https://doi.org/ARTN 215.10.1007/s00530-024-01359-z
[25] S. Choi, Y. Song, and H. Jung, "Study on Improving Detection Performance of Wildfire and Non-Fire Events Early Using Swin Transformer," IEEE ACCESS, vol. 13, pp. 46824-46837, 2025. https://doi.org/10.1109/ACCESS.2025.3528983
[26] C. Jin, A. Zheng, Z. Wu, and C. Tong, "Real-Time Fire Smoke Detection Method Combining a Self-Attention Mechanism and Radial Multi-Scale Feature Connection," Sensors (Basel), vol. 23, no. 6, Mar 22 2023. https://doi.org/10.3390/s23063358
[27] J. Wang, X. Zhang, and C. Zhang, "A lightweight smoke detection network incorporated with the edge cue," Expert Systems with Applications, vol. 241, 2024. https://doi.org/10.1016/j.eswa.2023.122583
[28] Y. Tian, Q. Ye, and D. Doermann, "Yolov12: Attention-centric real-time object detectors," arXiv preprint arXiv:2502.12524, 2025.
[29] K. Han, Y. Wang, Q. Tian, J. Guo, C. Xu, and C. Xu, "Ghostnet: More features from cheap operations," in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 1580-1589. https://doi.org/10.1109/cvpr42600.2020.00165
[30] Z. Zheng et al., "Enhancing Geometric Factors in Model Learning and Inference for Object Detection and Instance Segmentation," IEEE Transactions on Cybernetics, vol. 52, no. 8, pp. 8574-8586, 2022. https://doi.org/10.1109/TCYB.2021.3095305
[31] H. Zhang, C. Xu, and S. Zhang, "Inner-IoU: more effective intersection over union loss with auxiliary bounding box," arXiv preprint arXiv:2311.02877, 2023.
[32] P. Almeida, T. Rezende, A. Lisboa, and A. Barbosa, "Fire Detection based on Two-Dimensional Convolutional Neural Network and Temporal Analysis," 7th IEE LA-CCI, Temuco, 2021. https://doi.org/10.1109/la-cci48322.2021.9769824
[33] W. Liu et al., "SSD: Single Shot MultiBox Detector," in Computer Vision – ECCV 2016, Cham, 2016, pp. 21-37: Springer International Publishing. https://doi.org/10.1007/978-3-319-46448-0_2
[34] Y. Zhao et al., "DETRs Beat YOLOs on Real-time Object Detection," in 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024, pp. 16965-16974. https://doi.org/10.1109/CVPR52733.2024.01605
[35] R. Khanam and M. Hussain, "What is YOLOv5: A deep look into the internal features of the popular object detector," arXiv preprint arXiv:2407.20892, 2024.
[36] R. Varghese and M. S, "YOLOv8: A Novel Object Detection Algorithm with Enhanced Performance and Robustness," in 2024 International Conference on Advances in Data Engineering and Intelligent Computing Systems (ADICS), 2024, pp. 1-6. https://doi.org/10.1109/ADICS58448.2024.10533619
[37] C.-Y. Wang, I. H. Yeh, and H.-Y. Mark Liao, "YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information," in Computer Vision – ECCV 2024, Cham, 2025, pp. 1-21: Springer Nature Switzerland. https://doi.org/10.1007/978-3-031-72751-1_1
[38] A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding, “YOLOv10: Real-Time End-to-End Object Detection,” in Advances in Neural Information Processing Systems (NeurIPS), 2024. https://doi.org/10.52202/079017-3429
DOI:
https://doi.org/10.31449/inf.v50i14.14408Keywords:
Fire smoke detection, YOLOv11s, Multi-scale feature extraction, Feature Fusion, Attention Mechanism, Lightweight NetworkDownloads
Published
Issue
Section
License
Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.
All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.
Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.







