Recently with the development of aritifial intelligence, Computer-aided practice training system has been extensively studied to evaluate the English pronunciation automatically for improving conversation skills.
We describe mispronunciation detection and feedback generation system for non-native second language learners by using deep acoustic model based on factorized TDNN and language model with phoneme error model.
Unlike the previous method, our method is based on the phoneme error pattern and the language model, it can provide more efficient feedback to learners. Deep acoustic model consists of TDNN-F with grouped fully-connected layers and shuffle operation. This network architecture maintains recognition accuracy like traditional TDNN and costs less then it. Also, our system evaluates pronunciation proficiency of utterance in word level and phoneme level based on confidence from Minimum Bayesian Risk decoder, feedback is generated on it.
The proposed TDNN-F neural network costs 50% less than the TDNN-F and the model size is reduced by halves. It can be deployed in mobile devices. Feedback with details of error and multimedia can help learner to correct mispronunciation.
The Results were published in the "Int. J. Advanced Networking and Applications"(Vol 17, Issue 01,Pages 6719-6727(2025)) under the title of "Automatic Pronunciation Evaluation and Feedback Generation System based on Resource-efficient factorized TDNN and Phoneme Error Pattern Model".