CBT: Corrective boosting training approach for multi-choice commonsense question answering

Page view(s)
0
Checked on
CBT: Corrective boosting training approach for multi-choice commonsense question answering
Title:
CBT: Corrective boosting training approach for multi-choice commonsense question answering
Journal Title:
Expert Systems with Applications
Keywords:
Publication Date:
16 December 2025
Citation:
Fan, Y., Zou, B., Chen, Y., & Hong, Y. (2026). CBT: Corrective boosting training approach for multi-choice commonsense question answering. Expert Systems with Applications, 305, 130851. https://doi.org/10.1016/j.eswa.2025.130851
Abstract:
Multiple-Choice Commonsense Question Answering (MC-CQA) requires models to select the correct answer from given options for commonsense questions. Our investigation reveals a critical limitation in Large Language Models (LLMs) for this task: they exhibit significant inconsistency in their predictions. When the order of options is permuted, these models produce inconsistent predictions for the same question, revealing a dependence on superficial linguistic cues rather than robust commonsense reasoning abilities. To address this problem, we propose Corrective Boosting Training (CBT), a novel framework that mimics human progressive learning by alternating between task-specific question answering and contrastive optimization phases. The framework introduces two enhancement mechanisms: it automatically generates contrastive samples from mispredicted instances through question rewriting and chain-of-thought expansion, while also incorporating adversarial training to improve robustness against option-order perturbations. For precise consistency evaluation, we propose the Consistency Gap (CG) metric, which quantifies accuracy differences between the original test set and its option-permuted variants. We conduct comprehensive evaluations of the CBT framework on three benchmark datasets: CODAH, Discosense, and ComSense. Experimental results demonstrate that CBT achieves 4.88%and 2.32% improvements in standard accuracy over RoBERTa and ELECTRA baselines on CODAH/Discosense, respectively, outperforming state-of-the-art models. Notably, CBT exhibits remarkable stability in the Consistency Gap (CG) metric across all three datasets, reducing the CG value from 6%-8% to below 2% and showing clear advantages over mainstream models, including RoBERTa, ELECTRA, and LLaMA. These findings collectively validate CBT’s effectiveness in enhancing both model accuracy and prediction consistency for commonsense question answering.
License type:
Attribution-NonCommercial-NoDerivatives 4.0 International (CC BY-NC-ND 4.0)
Funding Info:
There was no specific funding for the research done
Description:
ISSN:
0957-4174
Files uploaded:

File Size Format Action
eswa-cbt-fyf-20251213.pdf 653.07 KB PDF Request a copy