Replication Data for: Modeling monophthongal versus diphthongal /aɪ/ in sung vocal performance with interpretable machine learning

DOI

Dataset description

This dataset contains replication data for the study "Modeling monophthongal versus diphthongal /aɪ/ in sung vocal performance with interpretable machine learning" (De Timmerman & Verbeke, 2026). The study investigates how binary perceptual annotations of monophthongal versus diphthongal /aɪ/ relate to measurable acoustic variation in sung vocal performance. Using a corpus of studio-recorded blues music processed with vocal-instrumental source separation, F1 and F2 formant trajectories were extracted for 1,004 /aɪ/ tokens, each perceptually categorized by a trained annotator as either monophthongal or diphthongal. These data were used to train and interpret two machine-learning models predicting the perceptual labels: a gradient boosted decision tree (XGBoost) trained on a set of engineered acoustic features, and a multilayer perceptron (MLP) trained directly on the raw, time-aligned formant trajectories.

Article abstract

This study combines modern source-separation techniques with interpretable machine learning methods to investigate how binary perceptual annotations of diphthongal and monophthongal /aɪ/ relate to measurable acoustic variation in sung vocal performance. Using a corpus of studio-recorded music processed with vocal-instrumental source separation, we extract F1 and F2 trajectories for 1,004 /aɪ/ tokens, each perceptually categorized as either monophthongal or diphthongal. These trajectories were modeled using two approaches: (i) a gradient boosted decision tree trained on engineered acoustic features, and (ii) a multilayer perceptron trained directly on raw formant trajectories. Both models achieved high accuracy and AUC-ROC, indicating that perceptual labels can reliably be predicted from both raw and engineered acoustic input. Moreover, a SHAP-based feature importance analysis showed that features such as Delta F1/Delta F2, cubic spline coefficients and trajectory derivatives captured systematic differences between perceived monophthongal and diphthongal tokens, highlighting the value of dynamic representations over static distance-based measures. The results indicate that perceptual annotations of diphthongal and monophthongal /aɪ/ correspond with quantifiable acoustic information and demonstrate how explainable machine learning can help map gradient vowel dynamics onto binary perceptual categories. The study further emphasizes how recent advances in source separation make sung performance a valuable new domain for phonetic research.

Python, 3.9.18

Keras, 3.10.0

XGBoost, 2.1.4

SHAP, 0.48.0

UMAP, 0.5.7

Praat, 6.4.18

Identifier
DOI https://doi.org/10.18710/RDU8M2
Related Identifier IsSupplementTo https://doi.org/10.1121/10.0044380
Metadata Access https://dataverse.no/oai?verb=GetRecord&metadataPrefix=oai_datacite&identifier=doi:10.18710/RDU8M2
Provenance
Creator De Timmerman, Romeo ORCID logo; Verbeke, Gil ORCID logo
Publisher DataverseNO
Contributor De Timmerman, Romeo; TROLLing curator; Ghent University; Stef Slembrouck; The Tromsø Repository of Language and Linguistics (TROLLing)
Publication Year 2026
Rights CC0 1.0; info:eu-repo/semantics/openAccess; http://creativecommons.org/publicdomain/zero/1.0
OpenAccess true
Contact De Timmerman, Romeo (Ghent University); TROLLing curator
Representation
Resource Type tabular; Dataset
Format text/plain; text/comma-separated-values
Size 15477; 570497; 3459369
Version 1.1
Discipline Acoustics; Engineering Sciences; Humanities; Mechanical and industrial Engineering; Mechanics and Constructive Mechanical Engineering