TY - GEN
T1 - From speech to letters - using a novel neural network architecture for grapheme based ASR
AU - Eyben, Florian
AU - Wöllmer, Martin
AU - Schuller, Björn
AU - Graves, Alex
PY - 2009
Y1 - 2009
N2 - Main-stream Automatic Speech Recognition systems are based on modelling acoustic sub-word units such as phonemes. Phonemisation dictionaries and language model based decoding techniques are applied to transform the phoneme hypothesis into orthographic transcriptions. Direct modelling of graphemes as sub-word units using HMM has not been successful. We investigate a novel ASR approach using Bidirectional Long Short-Term Memory Recurrent Neural Networks and Connectionist Temporal Classification, which is capable of transcribing graphemes directly and yields results highly competitive with phoneme transcription. In design of such a grapheme based speech recognition system phonemisation dictionaries are no longer required. All that is needed is text transcribed on the sentence level, which greatly simplifies the training procedure. The novel approach is evaluated extensively on the Wall Street Journal 1 corpus.
AB - Main-stream Automatic Speech Recognition systems are based on modelling acoustic sub-word units such as phonemes. Phonemisation dictionaries and language model based decoding techniques are applied to transform the phoneme hypothesis into orthographic transcriptions. Direct modelling of graphemes as sub-word units using HMM has not been successful. We investigate a novel ASR approach using Bidirectional Long Short-Term Memory Recurrent Neural Networks and Connectionist Temporal Classification, which is capable of transcribing graphemes directly and yields results highly competitive with phoneme transcription. In design of such a grapheme based speech recognition system phonemisation dictionaries are no longer required. All that is needed is text transcribed on the sentence level, which greatly simplifies the training procedure. The novel approach is evaluated extensively on the Wall Street Journal 1 corpus.
UR - https://www.scopus.com/pages/publications/77949404053
U2 - 10.1109/ASRU.2009.5373257
DO - 10.1109/ASRU.2009.5373257
M3 - Conference contribution
AN - SCOPUS:77949404053
SN - 9781424454792
T3 - Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition and Understanding, ASRU 2009
SP - 376
EP - 380
BT - Proceedings of the 2009 IEEE Workshop on Automatic Speech Recognition and Understanding, ASRU 2009
PB - IEEE Computer Society
T2 - 2009 IEEE Workshop on Automatic Speech Recognition and Understanding, ASRU 2009
Y2 - 13 December 2009 through 17 December 2009
ER -