The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge - Télécom Paris Accéder directement au contenu
Communication Dans Un Congrès Année : 2018

The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge

Nazim Dugan
  • Fonction : Auteur
Cornelius Glackin
  • Fonction : Auteur
Nigel Cannings
  • Fonction : Auteur

Résumé

This paper describes the system developed by the Empathic team for the open set condition of the Iberspeech 2018 Speech to Text Transcription Challenge. A DNN-HMM hybrid acoustic model is developed, with MFCC's and iVectors as input features, using the Kaldi framework. The provided ground truth transcriptions for training and development are cleaned up using customized clean-up scripts and then realigned using a two-step alignment procedure which uses word lattice results coming from a previous ASR system. 261 hours of data is selected from train and dev1 subsections of the provided data, by applying a selection criterion on the utterance level scoring results. The selected data is merged with the 91 hours of training data used to train the previous ASR system with a factor 3 times data augmentation by reverberation using a noise corpus on the total training data, resulting a total of 1057 hours of final …
Fichier non déposé

Dates et versions

hal-02288554 , version 1 (14-09-2019)

Identifiants

Citer

Nazim Dugan, Cornelius Glackin, Gérard Chollet, Nigel Cannings. The intelligent voice ASR system for the iberspeech 2018 speech to text transcription challenge. IberSPEECH 2018, Nov 2018, Barcelone, Spain. pp.272-276, ⟨10.21437/IberSPEECH.2018-57⟩. ⟨hal-02288554⟩
54 Consultations
0 Téléchargements

Altmetric

Partager

Gmail Facebook X LinkedIn More