Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR

Page view(s)
14
Checked on Sep 09, 2025
Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR
Title:
Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR
Journal Title:
Interspeech 2020
Keywords:
Publication Date:
27 October 2020
Citation:
Zhou, X., Lee, G., Yılmaz, E., Long, Y., Liang, J., Li, H. (2020). Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSR. Interspeech 2020, 5016–5020. https://doi.org/10.21437/interspeech.2020-2556
Abstract:
Transformer has shown impressive performance in automatic speech recognition. It uses an encoder-decoder structure with self-attention to learn the relationship between high-level representation of source inputs and embedding of target outputs. In this paper, we propose a novel decoder structure that features a self-and-mixed attention decoder (SMAD) with a deep acoustic structure (DAS) to improve the acoustic representation of Transformer-based LVCSR. Specifically, we introduce a self-attention mechanism to learn a multi-layer deep acoustic structure for multiple levels of acoustic abstraction. We also design a mixed attention mechanism that learns the alignment between different levels of acoustic abstraction and its corresponding linguistic information simultaneously in a shared embedding space. The ASR experiments on Aishell-1 show that the proposed structure achieves CERs of 4.8% on the dev set and 5.1% on the test set, which are the best reported results on this task to the best of our knowledge.
License type:
Publisher Copyright
Funding Info:
This research / project is supported by the National Research Foundation, and Agency for Science, Technology and Research (A*STAR) - AI Singapore Programme; Human Robot Collaborative AI for Advanced Manufacturing and Engineering (AME)
Grant Reference no. : A18A2b0046

This research / project is supported by the National Research Foundation Singapore - National Robotics Programme: Human-Robot Interaction Phase 1
Grant Reference no. : 1922500054

This research / project is supported by the National Research Foundation Singapore - National Robotics Programme; AI Speech Lab
Grant Reference no. : AISG100E-2018-006

This research / project is supported by the National Natural Science Foundation of China - NA
Grant Reference no. : 61701306
Description:
ISSN:
2958-1796