Total: 1
Accentedness and comprehensibility scales are widely used to evaluate pronunciation development in second language learners. However, such assessments rely heavily on human rater evaluations. This study investigates whether a speech large language model (LLM) can approximate human judgments of accentedness and comprehensibility. We first compare correlations between LLM scores and human ratings. We then apply linear mixed-effects models to examine whether LLM scores capture learner progress across pre-and post-test conditions. Finally, by combining segmental and suprasegmental measures with Lasso regression, we analyze whether the LLM employs acoustic cues similar to those human raters rely on when assigning scores. The results show that the LLM scores are moderately correlated with human ratings, capture pre-post test progress, and exhibit overlapping features with human evaluations. However, future research could explore fine-tuning the LLM and incorporating linguistic knowledge.