Debiasing Visual Question Answering via Ensemble Gradient Detection and Iterative Attention Forgetting

Qiuying Han; Shaohui Zhang; Peng Wang; Boyuan Li; Qiwen Lu

doi:10.31449/inf.v49i24.9013

Debiasing Visual Question Answering via Ensemble Gradient Detection and Iterative Attention Forgetting

Abstract

In Visual Question Answering (VQA), the model’s ability to understand and reason across different modalities—language, visual, and multimodal—is crucial for accurate predictions. However, recent studies have identified a significant challenge of language bias in VQA models, where the model’s reasoning often hinges on incorrect linguistic associations rather than genuine multimodal understanding. To addressthis issue, We propose a novel Ensemble Bias Gradient Debiasing Approach (EBGDA) that combines bias detection with a dynamic forgetting mechanism. Our method utilizes a bias detector to identify and score biases across linguistic, visual, and multimodal data, enabling the model to focus on unbiased informationfor more accurate predictions. Additionally, inspired by human reasoning, we introduce the Forgotten Attention Algorithm (FAA), which iteratively “forgets” irrelevant visual content, progressively concentrating attention on the image regions most relevant to the question. This combination of bias mitigation andattention focusing enhances the model’s ability to make multimodal inferences, reducing bias and improving overall performance. Extensive experiments on the VQA-CP v2, VQA v2, VQA-VS, GQA-OOD, and VQA-CE datasets demonstrate the effectiveness of our approach, showing state-of-the-art performance inmitigating biases and excelling in complex multimodal scenarios. Our approach achieves a 21.32% improvement over UpDn on VQA-CP v2, establishing a new state-of-the-art among methods without data augmentation.

Authors

Qiuying Han School of Computer Science and Technology, Zhoukou Normal University, Zhoukou 466001, China
Shaohui Zhang School of Artificial Intelligence, Zhoukou Normal University, Zhoukou 466001, China
Peng Wang School of Artificial Intelligence, Zhoukou Normal University, Zhoukou 466001, China
Boyuan Li School of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China
Qiwen Lu School of Computer and Information Engineering, Henan University, Zhengzhou 450046, China

DOI:

https://doi.org/10.31449/inf.v49i24.9013

Downloads

Additional Files

Latex Source File_2025.4.25

Published

12/18/2025

How to Cite

Han, Q., Zhang, S., Wang, P., Li, B., & Lu, Q. (2025). Debiasing Visual Question Answering via Ensemble Gradient Detection and Iterative Attention Forgetting. Informatica, 49(24). https://doi.org/10.31449/inf.v49i24.9013

Download Citation

Issue

Vol. 49 No. 24 (2025): Online-only issue

Section

Online-only

License

Authors retain copyright in their work. By submitting to and publishing with Informatica, authors grant the publisher (Slovene Society Informatika) the non-exclusive right to publish, reproduce, and distribute the article and to identify itself as the original publisher.

All articles are published under the Creative Commons Attribution license CC BY 3.0. Under this license, others may share and adapt the work for any purpose, provided appropriate credit is given and changes (if any) are indicated.

Authors may deposit and share the submitted version, accepted manuscript, and published version, provided the original publication in Informatica is properly cited.

Debiasing Visual Question Answering via Ensemble Gradient Detection and Iterative Attention Forgetting

Abstract

Authors

DOI:

Downloads

Additional Files

Published

How to Cite

Issue

Section

License

Developed By

Information