Mountain View
California
37.39°N / 122.08°W
Yichen LuEason
Research Scientist at Anuttacon, building real-time voice agents — voice interaction models, harness systems, and agent self-evolution.
easonlu1017@gmail.com
Teaching machines to listen.
I am a Research Scientist at Anuttacon, working on real-time voice agents — voice interaction models and large audio language models (LALMs). I have hands-on ownership across pretraining, post-training (SFT & RL), agent interaction design, and evaluation.
I hold an M.S. in Artificial Intelligence from CMU and a B.S. in Statistics & Computer Science from UIUC. At CMU I worked at WAVLab under Prof. Shinji Watanabe, on efficient inference for speech language models and unified audio-visual understanding.
My research interests have since shifted toward voice interaction models, harness systems, and agent self-evolution — how a model learns to drive its own scaffolding, and how that scaffolding in turn defines what the model needs to become. Earlier work centred on general speech and audio language models, audio-visual fusion, and multimodal language models.
During my undergraduate studies I was fortunate to work with Prof. Tarek Abdelzaher, Prof. Kris Hauser, and Prof. Yuxiong Wang, which greatly shaped my research journey.
- Sep 2025 EMNLP 2025 System Demo accepted: ViDove, a multimodal translation agent — now productized as vidove.ai.
- Aug 2025 We released Whispers from the Star, Anuttacon's AI-native game. I built the on-device ASR system.
- Jan 2025 Started as Research Scientist at Anuttacon.
- 2025 One paper accepted to Interspeech 2025: GALAXY, a large-scale open-domain multimodal dataset.
- 2024 Oral presentation at EMNLP 2024: FastAdaSP — multitask-adapted efficient inference for large speech LMs.
- Jan 2025 — Present Research Scientist, Anuttacon — Mountain View, CA I own post-training for the company's voice agent model, from recipe design through evaluation, delivering roughly 5–20% gains on audio benchmarks. I designed the agent harness that defines how the real-time speech model coordinates with the backend agent and the user, and led development of our general audio understanding model across speech, audio and music. I also built the evaluation infrastructure and shipped the on-device ASR behind Whispers from the Star.
- Jul 2023 — Jan 2025 Research Assistant, CMU WAVLab — Pittsburgh, PA Advised by Prof. Shinji Watanabe. Proposed FastAdaSP for efficient speech LM inference (EMNLP 2024 Oral) and SynesLM for unified audio-visual speech recognition and translation. Core contributor to ESPnet, working on discrete speech unit modeling.
- May 2023 — Aug 2023 Machine Learning Engineer Intern, TrovaAI — Champaign, IL Designed a RAG retrieval pipeline combining vector and keyword indices over heterogeneous documents for an AI agent platform, and built LLM-based and MT-metric (BLEU/COMET) evaluation for the agent system.
- Jul 2021 — Jan 2022 Machine Learning Engineer Intern, VMware, Inc. — Beijing, China Built an automated data analysis and reporting system for BERT-based large-scale machine translation, improving query and analysis throughput via the Elastic Stack.



SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation
Interspeech 2024 Workshop
- Aug 2023 — Dec 2024 Carnegie Mellon University — M.S. in Artificial Intelligence Coursework: Multimodal ML, LLM Systems, System Tool Chains for AI, Speech Recognition and Understanding.
- Aug 2019 — May 2023 University of Illinois Urbana-Champaign — B.S. in Statistics & Computer Science GPA 3.93/4.00. Graduated with Highest Honors; Dean's List.
- Founder ViDove — multimodal translation agent (vidove.ai) An end-to-end multimodal translation agent for video subtitle generation, with multimodal context and memory-augmented reasoning (EMNLP 2025 System Demo). Led a 10-person engineering team and productized it as vidove.ai. Deployed in production with the StarCraft II World Team League, one of the largest SC2 e-sports tournaments worldwide. vidove.ai GitHub
- Feb 2014 — Dec 2018 Sugar Masses Creative — China Minecraft Construction Summit Founder, developer and director of the largest official Minecraft tournament in China, drawing over 1,000 contestants and more than a million online viewers a year with a team of 30. Secured long-term cooperation with NetEase, Qihoo 360, Tencent and Youku, and investment from NetEase and JoyMe.com. Presented at ChinaJoy 2016 with more than 325,000 entries. Video 1 Video 2 NetEase coverage
- Technical · 28 May 2026 Some Thoughts on Agents — Model as Harness System
- Personal · 01 Nov 2025 Notes, 1 Nov 2025 — On Growth, Self-Consistency and Love of the Work
Reviewer / Program Committee — IEEE ASRU 2025, IEEE T-ASLP, AAAI 2026, ICASSP.
- A tuxedo cat named Brann.
- Favourite musical — Hamilton.
- Formerly an AMVer, anime music video creator.