Vu, M., Wei, Z., Bhattarai, B., Teh, K. K., & Huy Dat, T. (2024). VietSing: A High-quality Vietnamese Singing Voice Corpus. In (Editor), 2024 Asia Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). https://doi.org/10.1109/apsipaasc63619.2025.10849014
Abstract:
This paper introduces a comprehensive Vietnamese dataset designed for singing voice synthesis (SVS). While there are extensive datasets available for widely spoken languages such as English and Chinese, resources for less common languages like Vietnamese are still scarce. The dataset, VietSing, comprises high-quality audio recordings and corresponding phonetic annotations, meticulously curated to support the development and evaluation of SVS systems. Detailed phonetic transcriptions and alignment with musical scores are provided to facilitate precise modeling of Vietnamese phonetics and prosody in song. We outline the data collection process, annotation methodology, and the challenges faced in ensuring linguistic and musical accuracy. Initial experiments using popular SVS model demonstrate the potential of VietSing to enhance the naturalness and intelligibility of synthesized Vietnamese singing voices.
License type:
Publisher Copyright
Funding Info:
This research / project is supported by the Industry funded - EC-2023-105
Grant Reference no. : NA