The Lab BenchAccepted · IEEE BigData 2026

Cross-Modal Attention for Visual Question Answering

Research with Prof. Hajar Homayouni at San Diego State University's Data Science Lab, January 2026 to now. I'm the author along with Prof Homayouni on the resulting paper, which was accepted to the IEEE BigData2026 Conference. I've been invited to present my research to an international audience at one of the premier data mining conferences.

Diagram of shared cross-attention: image and text features go through two shared-weight experts, then fuse into a prediction
Shared cross-attention architecture.

the paper

Header of the accepted paper: title, authors, affiliations, and the opening lines of the abstract

The question

In visual question answering, a model looks at an image and answers a question about it, like "what color is the umbrella?" To do that, the image features and the text features have to talk to each other. Most models let both sides attend to each other the same way (symmetric). We wanted to know what happens if the attention is asymmetric, so one side leads.

What I did

  • Compared symmetric and asymmetric cross-modal attention on the VQA v2 benchmark, using frozen and unfrozen CLIP-L/14 (vision) and RoBERTa (language) encoders with a trainable fusion module of about 307M parameters.
  • Engineered and evaluated 9+ models. Our best reached about 0.85× the accuracy of a state-of-the-art Google model.
  • Ranked 175 of 550 on the VQA Challenge leaderboard on EvalAI.
  • Built and ran the whole training pipeline on Google Cloud: L4 and A100 GPU VMs, a Cloud Storage data pipeline, and Papermill plus tmux so runs kept going without me watching.
  • Did four code reviews for the team and presented our progress to a graduate Data Science class twice.

What we found

The symmetric model reached about 72% validation accuracy, and the asymmetric one reached about 67%. The interesting part was how the asymmetric model failed. Its training accuracy kept climbing (to about 71%) while validation accuracy flattened out. That widening gap pointed to overfitting.

I traced it to a structural gradient imbalance in the asymmetric design and proposed fixes: differential learning rates, training for more epochs, and warm-starting.

How I got here

I cold-emailed professors asking about AI research. My interview with Dr. Homayouni went really well. She liked my background, and I was excited about her work, so I joined the lab.