YOLO-SCEMA: Efficient Multi-scale Attention Fusion Module Combining Spatial and Channel reconstruction Convolution

Abstract

Convolutional Neural Networks (CNNs) have achieved significant performance in various computer visiontasks, but at the cost of enormous computing resources. To alleviate this dilemma, this paper proposes alightweight and efficient multiscale attention fusion module (SCEMA) by combining spatial, and channelreconstruction convolution with an efficient attention mechanism and adds the module to the YOLOV8 networkstructure named YOLO-SCEMA. SCEMA adopts a parallel processing strategy, with the left branchperforming feature refinement operations through spatial and channel reconstruction units to reduce redundantcalculations, and the right branch effectively integrating features of different scales through featuregrouping and cross-spatial learning, enhancing the model’s understanding of multiscale image contentand improving its performance in handling complex image structures. The experimental results on the opensource datasets ExDark,VisDrone2019 and FYP show that YOLO-SCEMA has increased the mAP(50) scoreby 7.37%,3.24% and 1.5% , and compared to the YOLOv8 benchmark while reducing the parameter andcomputational complexity by 36.9% and 8.6%, respectively. Compared with the latest YOLO series, YOLOSCEMAperforms better in detection accuracy and parameter quantity.

Author Biography

  • Zufeng Fu, Nanyang Normal University;Collaborative Innovation Center of Intelligent Explosion-proof Equipment
    Corresponding Author

References

Authors

  • Zhuang Yang Nanyang Normal University image/svg+xml
  • Haiying Wang Nanyang Normal University image/svg+xml
  • Xiaolong Fu Nanyang Normal University image/svg+xml
  • Ming Hui Nanyang Normal University;Collaborative Innovation Center of Intelligent Explosion-proof Equipment
  • Xiaolong Gao Nanyang Normal University image/svg+xml
  • Kun Ning Nanyang Normal University image/svg+xml
  • Zufeng Fu Nanyang Normal University;Collaborative Innovation Center of Intelligent Explosion-proof Equipment

DOI:

https://doi.org/10.31449/inf.v50i14.12887

Keywords:

Complex scene object detection, lightweight, attention mechanisms, YOLO, multiscale fusion

Downloads

Published

08/06/2026

How to Cite

Yang, Z., Wang, H., Fu, X., Hui, M., Gao, X., Ning, K., & Fu, Z. (2026). YOLO-SCEMA: Efficient Multi-scale Attention Fusion Module Combining Spatial and Channel reconstruction Convolution. Informatica, 50(14). https://doi.org/10.31449/inf.v50i14.12887