Cheng Shen, Yew-Soon Ong, and Joey Tianyi Zhou. 2025. CondenseLM: LLMs-driven Text Dataset Condensation via Reward Matching. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 1237–1252, Suzhou, China. Association for Computational Linguistics.
Abstract:
Dataset condensation has emerged as a promising technique to improve data efficiency under
limited data budgets. However, when applied to
the text level, existing methods struggle to compress more information into samples through
optimization. Thus, these methods provide no
obvious advantage over simpler coreset selection despite their high computational cost. In
this paper, we introduce CondenseLM, a novel
paradigm for both effective and efficient textlevel dataset condensation. Our framework employs an LLMs-driven approach to sidestep the
inherent limitations of existing methods, successfully generating more informative and less
biased samples. In addition, it incorporates reward matching to align the LLMs-condensed
dataset with the original dataset, maximizing
representability and coverage. We conducted
extensive experiments on SST-2, MNLI, AG
News, and IMDB. Our approach outperforms
both coreset selection and existing dataset condensation methods by large margins while also
substantially reducing the computational cost.
License type:
Publisher Copyright
Funding Info:
This research / project is supported by the National Research Foundation - National Large Language Models Funding Initiative
Grant Reference no. : AISG-NMLP-2024-003
This research / project is supported by the National Research Foundation - Digital Trust Centre Innovation Grant
Grant Reference no. : DTC-IGC-02