Publications
David Mortensen's publications including all publications with other members of ChangeLing Lab.
2026
- InterspeechAn Empirical Recipe for Universal Phone RecognitionIn Interspeech 2026
- InterspeechScaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific ClusterIn Interspeech 2026
- InterspeechAdapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson’s DiseaseIn Interspeech 2026
- Self-supervised speech models encode phonetic context via position-dependent orthogonal subspaces
- LM4UCHarnessing Linguistic Dissimilarity for Language Generalization on Unseen Low-Resource VarietiesIn Second Workshop on Language Models for Underserved Communities (LM4UC)
- ACLCommunicating in Emergent Language with an Induced Morphological PhrasebookIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- ACLPRiSM: Benchmarking Phone Realization in Speech ModelsIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- ACLPOWSM: A Phonetic Open Whisper-Style Speech Foundation ModelIn Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- ACL Findings[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector ArithmeticIn Findings of the Association for Computational Linguistics: ACL 2026
- ACL FindingsLinear Script Representations in Speech Foundation Models Enable Zero-Shot TransliterationIn Findings of the Association for Computational Linguistics: ACL 2026
- ACL FindingsPBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical LinguisticsIn Findings of the Association for Computational Linguistics: ACL 2026
- LChangeFrom sunblock to softblock: Analyzing the correlates of neology in published writing and on social mediaIn The Proceedings for the 6th International Workshop on Computational Approaches to Language Change (LChange’26)
- EACLHappiness is Sharing a Vocabulary: A Study of Transliteration MethodsIn Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers)
2025
- IJCNLP-AACLThe Translation Barrier Hypothesis: Multilingual Generation with Large Language Models Suffers from Implicit Translation FailureIn Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics
- EMNLPMorpheme Induction for Emergent LanguageIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
- EMNLPSearching for the Most Human-like Emergent LanguageIn Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing
- EMBCDomain-Specific Multilingual Strategies for Medical NLP: A Cross-Lingual Analysis of Orthographic and Phonemic RepresentationsIn 2025 47th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC)
- ACLZIPA: A family of efficient models for multilingual phone recognitionIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- ACLDialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to ModelsIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- ACLProgramming by Example meets Historical Linguistics: A Large Language Model Based Approach to Sound Law InductionIn Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- NAACLLeveraging Allophony in Self-Supervised Speech Models for Atypical Pronunciation AssessmentIn Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
- PNASDerivational morphology reveals analogical generalization in large language modelsProceedings of the National Academy of Sciences
- InterspeechTowards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource LanguagesIn Interspeech 2025
2024
- InterspeechSelf-supervised Speech Representations Still Struggle with African American Vernacular EnglishIn Proc. INTERSPEECH 2024
- Can Large Language Models Code Like a Linguist?: A Case Study in Low Resource Sound Law Induction
- Neural Proto-Language Reconstruction
- EMNLPZero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource LanguagesIn Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
- ACLSemisupervised Neural Proto-Language ReconstructionIn Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)
- TMLRA Review of the Applications of Deep Learning-Based Emergent CommunicationTransactions on Machine Learning Research
- LREC-COLINGConstructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong ReasonsIn Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
- LREC-COLINGImproved Neural Protoform Reconstruction via Reflex PredictionIn Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
- LREC-COLINGPhonotactic Complexity across DialectsIn Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
- LREC-COLINGPWESuite: Phonetic Word Embeddings and Tasks They FacilitateIn Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
- LREC-COLINGVerbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMsIn Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)
- NAACLXferBench: a Data-Driven Benchmark for Emergent LanguageIn Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers)
2023
- Kuki-Chin Phonology: An OverviewHimalayan Linguistics
- ASRUEvaluating self-supervised speech models on a Taiwanese Hokkien corpusIn 2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
- AfricaNLPAfrican Substrates Rather Than European Lexifiers to Augment African-diaspora Creole TranslationIn 4th Workshop on African Natural Language Processing
- LChangeAutomating Sound Change Prediction for Phylogenetic Inference: A Tukanoan Case StudyIn Proceedings of the 4th Workshop on Computational Approaches to Historical Language Change
- WMTChatGPT MT: Competitive for High- (but Not Low-) Resource LanguagesIn Proceedings of the Eighth Conference on Machine Translation
- EMNLPDo All Languages Cost the Same? Tokenization in the Era of Commercial Language ModelsIn Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
- EMNLPCounting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language ModelIn Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing
- EMNLP FindingsCalibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity TypingIn Findings of the Association for Computational Linguistics: EMNLP 2023
- CxGs+NLPConstruction Grammar Provides Unique Insight into Neural Language ModelsIn Proceedings of the First International Workshop on Construction Grammars and NLP (CxGs+NLP, GURT/SyntaxFest 2023)
- ACLTransformed Protoform ReconstructionIn Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)
- SIGMORPHONGeneralized Glossing Guidelines: An Explicit, Human- and Machine-Readable, Item-and-Process Convention for Morphological AnnotationIn Proceedings of the 20th SIGMORPHON workshop on Computational Research in Phonetics, Phonology, and Morphology
- SIGMORPHONSigMoreFun Submission to the SIGMORPHON Shared Task on Interlinear GlossingIn Proceedings of the 20th SIGMORPHON workshop on Computational Research in Phonetics, Phonology, and Morphology
- TSDMultilingual TTS Accent Impressions for Accented ASRIn International Conference on Text, Speech, and Dialogue
2022
- LoResMTData-adaptive Transfer Learning for Translation: A Case Study in Haitian and JamaicanIn Proceedings of the Fifth Workshop on Technologies for Machine Translation of Low-Resource Languages (LoResMT 2022)
- COLINGWikiHan: A New Comparative Dataset for Chinese LanguagesIn COLING 2022
- InterspeechWhen Is TTS Augmentation Through a Pivot Language Useful?In Interspeech 2022
- InterspeechSpeech Recognition for Around 2000 Languages without AudioIn Interspeech 2022
- InterspeechASR2K: Speech Recognition for Around 2000 Languages without AudioIn Interspeech 2022
- LRECA Hmong Corpus with Elaborate Expression AnnotationsIn Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)
- LRECPhone Inventories and Recognition for Every LanguageIn Proceedings of the Thirteenth International Conference on Language Resources and Evaluation (LREC 2022)
- Large-Scale Computerized Forward Reconstruction Yields New Perspectives in French Diachronic PhonologyDiachronica
- Wfst4StrRust/Python library for working with strings using weighted finite state transducers
- ACL FindingsZero-shot Learning for Grapheme to Phoneme Conversion with Language EnsembleIn Findings of the Association for Computational Linguistics: ACL 2022
- NAACLLearning the Ordering of Coordinate Compounds and Elaborate Expressions in Hmong, Lahu, and ChineseIn Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2021
- TACLQuantifying Cognitive Factors in Lexical DeclineTransactions of the Association for Computational Linguistics
- ICASSPMultilingual phonetic dataset for low resource speech recognitionIn ICASSP 2021
- East Tusom: A phonetic and phonological sketch of a largely undocumented Tangkhulic language
- InterspeechTusom2021: A Phonetically Transcribed Speech Dataset from an Endangered Language for Universal Phone Recognition ExperimentsIn Proc. Interspeech 2021
- InterspeechPhoneme Recognition Through Fine Tuning of Phonetic Representations: A Case Study on Luhya Language VarietiesIn Proc. Interspeech 2021
- InterspeechDifferentiable Allophone Graphs for Language-Universal Speech RecognitionIn Proc. Interspeech 2021
- EMNLPEvaluating the Morphosyntactic Well-formedness of Generated TextsIn Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing
- EACLRanking Transfer Languages with Pragmatically-Motivated Features for Multilingual Sentiment AnalysisIn EACL 2021
2020
- EMNLPAutomatic Extraction of Rules Governing Morphological AgreementIn Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing
- AAAITowards Zero-shot Learning for Automatic Phonemic TranscriptionIn Proceedings of the Thirty-Fourth AAAI Conference on Artificial Intelligence
- ICASSPUniversal Phone Recognition with a Multilingual Allophone SystemIn ICASSP 2020
- Computerized Forward Reconstruction for Analysis in Diachronic Phonology, and Latin to French Reflex PredictionIn 1st Workshop on Language Technologies for Historical and Ancient LAnguages (LT4HALA)
- Characterizing Sociolinguistic Variation in the Competing Vaccination CommunitiesIn Proceedings of the International Conference SBP-BRiMS 2020
- LRECAlloVera: A Multilingual Allophone DatabaseIn Proceedings of the Twelfth International Conference on Language Resources and Evaluation (LREC 2020)
- SCiLWhere New Words Are Born: Distributional Semantic Analysis of Neologisms and Their Semantic NeighborhoodsIn Proceedings of the Society for Computation in Linguistics
2019
- SIGMORPHONCMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in MorphologyIn Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology
- Hmong (Mong Leng)In The Mainland Southeast Asia Linguistic Area
- IndoMorphCollection of Foma FST morphological analyzers for languages of the Indian subcontinent
2018
- EMNLPAdapting Word Embeddings to New Languages with Morphological and Phonological Subword RepresentationsIn Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing
- The ARIEL-CMU situation frame detection pipeline for LoReHLT16: a model translation approachMachine Translation
- EpitranPrecision orthography-to-IPA conversion for 65 languages
- MStemPython multilingual morphological stemming framework and stemmer collection
- LRECParser combinators for Tigrinya and Oromo morphologyIn Proceedings of the 11th Language Resources and Evaluation Conference
- LRECEpitran: Precision G2P for Many LanguagesIn Proceedings of the 11th Language Resources and Evaluation Conference
2017
- EACLURIEL and lang2vec: Representing languages as typological, geographical, and phylogenetic vectorsIn Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers
- Hmong-Mien LanguagesIn Oxford Research Encyclopedia of Linguistics
- Hmong-Mien LanguagesIn Oxford Bibliographies in Linguistics
2016
- EMNLPPhonologically Aware Neural Model for Named Entity Recognition in Low Resource Transfer SettingsIn Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing
- LRECBridge-Language Capitalization Inference in Western Iranian: Sorani, Kurmanji, Zazaki, and TajikIn Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016)
- COLINGNamed Entity Recognition for Linguistic Rapid Response in Low-Resource Languages: Sorani Kurdish and TajikIn Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers
- COLINGPanPhon: A Resource for Mapping IPA Segments to Articulatory Feature VectorsIn Proceedings of COLING 2016, the 26th International Conference on Computational Linguistics: Technical Papers
- PanPhonArticulatory feature extractor and library
- NAACLPolyglot Neural Language Models: A Case Study in Cross-Lingual Phonetic Representation LearningIn Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies
2013
- Lexical Prefixes and Tibeto-Burman Laryngeal ContrastsIn Proceedings of the Thirty-Seventh Annual Meeting of the Berkeley Linguistics Society (BLS 37)
- A Reconstruction of Proto-Tangkhulic RhymesLinguistics of the Tibeto-Burman Area
- Tonally Conditioned Vowel Raising in Shuijingping MangJournal of East Asian Linguistics
2012
- Sorbung, an Undescribed Language of Manipur: Its Phonology and Place in Tibeto-BurmanJournal of the Southeast Asian Linguistics Society
- The Emergence of Obstruents after High VowelsDiachronica
- NetSPEA web-based application for exploring rule-based analyses of phonological data
- A Classification of Compounding in American Sign Language: an Evaluation of the Bisetto and Scalise FrameworkMorphology
2011
- HsSPEHaskell library implementing SPE-style rule-based phonology
- Web ComparatorA web-based application for organizing and analyzing comparative lexical databases. Used in the production of “Emergence of Obstruents” and “Proto-Tangkhulic Rhymes”
2004
- The emergence of dorsal stops after high vowels in HuishuIn Proceedings of the Thirtieth Annual Meeting of the Berkeley Linguistics Society (BLS 30)
2003
- Review of \emphBaheng Yu Yanjiu [research on the Pa-Hng language] by Mao Zongwu and Li YunbingLinguistics of the Tibeto-Burman Area
2002
- Review of \emphLes langue Hmong-Mjen: phonologie historique by Barbara NiedererLinguistics of the Tibeto-Burman Area