J Appl Oral Sci. 2026 Sep 7;34:e20260078. doi: 10.1590/1678-7765-2026-0078. eCollection 2026.
ABSTRACT
INTRODUCTION: Speech sound disorders (SSDs) are common in children and can affect academic and psychosocial outcomes. Clinical identification relies on expert auditory-perceptual assessment, which is time-intensive and may vary across raters. Automated screening tools could support triage when access to specialists is limited.
OBJECTIVE: This study aimed to develop and evaluate a deep learning (DL) classifier for detecting mispronunciation errors in standardized pediatric speech recordings obtained using a structured fricative-focused word elicitation protocol, using expert-adjudicated labels as the reference standard.
METHODOLOGY: In this cross-sectional study, we analyzed 100 participants (6-18 years) providing 1,800 standardized word recordings. Two expert speech-language pathologists (SLPs) labeled recordings as mispronunciation present vs. absent; disagreements were adjudicated by a third SLP. Inter- and intra-rater reliability were assessed. Audio was denoised and standardized to 16 kHz. A pretrained transformer speech model (WavLM Base+) was fine-tuned for recording-level binary classification. Class imbalance was addressed using augmentation, weighted sampling, and focal loss. Performance was assessed on a held-out test set at the recording level using accuracy, sensitivity, specificity, precision, F1-score, and area under the ROC curve (AUC).
RESULTS: The overall recording-level prevalence of speech sound errors was 16.15%, with the “th” category (/θ/ + /ð/) being the most frequently mispronounced phoneme group. No significant differences were observed by sex, age group (6-9, 10-13, 14-18 years), or malocclusion status (p > 0.05). Labeling showed strong agreement (Cohen’s κ = 0.84; intra-rater reliability 0.89-0.92). On the test set (253 recordings), the model achieved 90.9% accuracy, 77.8% sensitivity, 98.2% specificity, 95.9% precision, and AUC = 0.936.
CONCLUSIONS: Fine-tuned pretrained speech representations demonstrated promising performance for screening pediatric mispronunciation errors under standardized recording conditions using expert labels. The high-specificity profile supports its use as a screening and triage decision-support tool, while future studies are needed to validate performance across broader clinical and real-world settings.
PMID:42715461 | DOI:10.1590/1678-7765-2026-0078