Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents
Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents

Application of machine learning techniques to automate speech mispronunciation detection in children and adolescents

J Appl Oral Sci. 2026 Sep 7;34:e20260078. doi: 10.1590/1678-7765-2026-0078. eCollection 2026.

ABSTRACT

INTRODUCTION: Speech sound disorders (SSDs) are common in children and can affect academic and psychosocial outcomes. Clinical identification relies on expert auditory-perceptual assessment, which is time-intensive and may vary across raters. Automated screening tools could support triage when access to specialists is limited.

OBJECTIVE: This study aimed to develop and evaluate a deep learning (DL) classifier for detecting mispronunciation errors in standardized pediatric speech recordings obtained using a structured fricative-focused word elicitation protocol, using expert-adjudicated labels as the reference standard.

METHODOLOGY: In this cross-sectional study, we analyzed 100 participants (6-18 years) providing 1,800 standardized word recordings. Two expert speech-language pathologists (SLPs) labeled recordings as mispronunciation present vs. absent; disagreements were adjudicated by a third SLP. Inter- and intra-rater reliability were assessed. Audio was denoised and standardized to 16 kHz. A pretrained transformer speech model (WavLM Base+) was fine-tuned for recording-level binary classification. Class imbalance was addressed using augmentation, weighted sampling, and focal loss. Performance was assessed on a held-out test set at the recording level using accuracy, sensitivity, specificity, precision, F1-score, and area under the ROC curve (AUC).

RESULTS: The overall recording-level prevalence of speech sound errors was 16.15%, with the “th” category (/θ/ + /ð/) being the most frequently mispronounced phoneme group. No significant differences were observed by sex, age group (6-9, 10-13, 14-18 years), or malocclusion status (p > 0.05). Labeling showed strong agreement (Cohen’s κ = 0.84; intra-rater reliability 0.89-0.92). On the test set (253 recordings), the model achieved 90.9% accuracy, 77.8% sensitivity, 98.2% specificity, 95.9% precision, and AUC = 0.936.

CONCLUSIONS: Fine-tuned pretrained speech representations demonstrated promising performance for screening pediatric mispronunciation errors under standardized recording conditions using expert labels. The high-specificity profile supports its use as a screening and triage decision-support tool, while future studies are needed to validate performance across broader clinical and real-world settings.

PMID:42715461 | DOI:10.1590/1678-7765-2026-0078