Yue Chen, Ph.D.

About Me

I am an applied machine learning researcher with a Ph.D. in Computational Linguistics from Indiana University, where I worked with Dr. Sandra Kübler. My experience spans academic research and production-scale machine learning systems at Microsoft.

My research focuses on representation learning and analysis, model generalization, interpretability, and evaluation, particularly for speech, language, and human-centered AI. I am interested in understanding what models actually learn, how well they generalize across real-world distribution shifts, and how to design rigorous evaluations that distinguish genuine capability from dataset- or representation-specific effects.

At Microsoft, I have worked on production ML and data systems involving forecasting, anomaly detection, telemetry analysis, and AI-assisted diagnostics in large-scale cloud environments. I am especially interested in research that combines careful experimentation with real-world AI systems.

Before Microsoft, I worked as a software engineer at Oracle and SUSE Linux, and worked on dialogue systems during an internship at Interactions LLC.

Research Focus

  • Representation Learning & Representation Analysis — understanding which aspects of learned or engineered representations carry discriminative signal and how representation structure relates to downstream behavior.
  • Model Generalization, Robustness & Evaluation — controlled evaluation across speakers, linguistic content, and other sources of distribution shift; designing experiments that separate genuine generalization from shortcut learning or leakage.
  • Speech, Language & Human-Centered AI — modeling complex human signals, including emotion and empathy, with an emphasis on reliable and interpretable evaluation.
  • Applied AI Research in Real-World Systems — connecting research questions to production-scale ML, telemetry, forecasting, anomaly detection, and AI-assisted diagnostics.

Selected Research

Generalization in Speech Emotion Recognition

I study how speech emotion models generalize beyond the conditions represented in their training data. My work uses controlled experiments to examine performance across speakers, lexical content, repeated observations, and other sources of variation rather than treating aggregate accuracy as the only measure of success.

Representation Analysis and Interpretability

I use dimensionality reduction and feature/sub-feature analysis to investigate what information different representations capture, how discriminative structure changes across conditions, and when visual or statistical evidence is strong enough to support conclusions about model behavior.

Production Machine Learning

At Microsoft, I have worked on ML and data systems for forecasting, anomaly detection, telemetry analysis, and AI-assisted diagnostics. This experience informs my interest in research problems where evaluation quality, reliability, and deployment constraints matter as much as model performance in isolation.

Publications

Please see my Google Scholar profile, Semantic Scholar profile, Microsoft Research profile, and arXiv profile.

One representative project is our WASSA 2022 work on empathy detection and emotion classification, where we investigated how demographic attributes affected model performance and evaluated the tradeoff between additional features and generalization.

News

  • 05/27/2026: I am reviewing for WiNLP 2026.
  • 03/25/2026: I am reviewing for NeurIPS 2026.
  • 01/13/2026: I am reviewing for COLM 2026.
  • 01/08/2026: I am reviewing for LoResLM 2026.
  • 09/08/2025: I am reviewing for AISTATS 2026.
  • 07/14/2025: I am joining the program committee of AAAI 2026.
  • 07/10/2025: I am joining the program committee of AAAI 2027.
  • 06/30/2025: I joined the WiNLP pre-submission mentorship program, supporting authors and early-career researchers.
  • 05/29/2025: I am reviewing for WiNLP 2025.
  • 03/31/2025: I am joining the program committee of RANLP 2025.

Older news

Research Community Service

Journal Reviewer

  • Natural Language Processing
  • IEEE Transactions on Multimedia
  • IULC Working Papers

Conference Program Committee / Reviewer

  • NeurIPS (2021–2026)
  • ICLR (2020–2025)
  • ICML (2022, 2024, 2025)
  • AAAI (2025–2027)
  • AISTATS (2025–2026)
  • ACL / EMNLP / NAACL and ACL Rolling Review
  • COLM (2024–2026)
  • RepL4NLP (2020–2025)
  • WiNLP (2024–2026)
  • RANLP (2017–2025)
  • AACL, COLING, IJCNLP, LoResLM, SocialNLP, and RANLP SRW

Conference Organizing

  • Publicity Chair, DeepCBR at IJCAI 2021

Invited Talks

Professional Organizations

ClingDing

I used to (not anymore) organize a weekly Computational Linguistics talk series called ClingDing. Per university policy, I am not allowed to post the Zoom link here, so if you are interested, please email me for more details.

Fun