Improving model adversarial robustness in Extractive Question Answering via Wasserstein-Guided feature Representations
Gang Huang, Lu Zhang, Hailun Wang · Alexandria Engineering Journal · 2025
Extractive Question Answering (EQA) models aim to locate accurate answers from passages given a question but are highly susceptible to adversarial attacks. Existing adversarial training methods improve robustness by generating perturbed passages, yet they remain computationally expensive and prone to producing noisy examples, limiting their effectiveness. To address these challenges, a novel adversarial training algorithm, W asserstein- D istance- G uided feature R epresentations ( WDGR ), is proposed, which operates directly on perturbed feature representations instead of generating full adversarial passages. Specifically, the proposed method employs a feature generator to perturb the original sample’s feature space, crafting challenging adversarial examples characterized by a large Wasserstein distance from the original features. These adversarial samples are designed to maximize the model’s noise tolerance while ensuring extracted answers remain consistent with the correct ones using original samples. This approach not only reduces the time complexity of adversarial training but also produces more effective and targeted adversarial examples, improving the model’s robustness against adversarial attacks. Extensive evaluations across multiple datasets and attack scenarios demonstrate that WDGR consistently outperforms existing adversarial defense mechanisms, recovering up to 20% of performance lost under adversarial conditions. These findings highlight its potential applications in real-world question-answering systems, including search engines, customer support automation, and intelligent tutoring systems, where robustness against adversarial manipulations is critical.