Mountain View
California
37.39°N / 122.08°W
Yichen LuEason
Research Scientist at HoYoverse (Prev. Anuttacon), building real-time voice agents — voice interaction models, harness systems, and agent self-evolution.
easonlu1017@gmail.com
Teaching machines to listen.
I am a Research Scientist at HoYoverse (Prev. Anuttacon), working on real-time voice agents — voice interaction models and large audio language models (LALMs). I have hands-on ownership across pretraining, post-training (SFT & RL), agent interaction design, and evaluation.
I hold an M.S. in Artificial Intelligence from CMU and a B.S. in Statistics & Computer Science from UIUC. At CMU I worked at WAVLab under Prof. Shinji Watanabe, on efficient inference for speech language models and unified audio-visual understanding.
My research interests have since shifted toward voice interaction models, harness systems, and agent self-evolution — how a model learns to drive its own scaffolding, and how that scaffolding in turn defines what the model needs to become. I build Proteus, an open-source framework for measuring exactly that. Earlier work centred on general speech and audio language models, audio-visual fusion, and multimodal language models.
During my undergraduate studies I was fortunate to work with Prof. Tarek Abdelzaher, Prof. Kris Hauser, and Prof. Yuxiong Wang, which greatly shaped my research journey.
- Aug 2026 Released Proteus v0.2.0 — a harness-agnostic framework for measuring agent self-evolution, now on PyPI.
- Sep 2025 EMNLP 2025 System Demo accepted: ViDove, a multimodal translation agent — now productized as vidove.ai.
- Aug 2025 We released Whispers from the Star, Anuttacon's AI-native game. I built the on-device ASR system.
- Jan 2025 Started as Research Scientist at Anuttacon.
- 2025 One paper accepted to Interspeech 2025: GALAXY, a large-scale open-domain multimodal dataset.
- 2024 Oral presentation at EMNLP 2024: FastAdaSP — multitask-adapted efficient inference for large speech LMs.
- Jan 2025 — Present Research Scientist, HoYoverse (Prev. Anuttacon) — Mountain View, CA I own post-training for the company's voice agent model, from recipe design through evaluation, delivering roughly 5–20% gains on audio benchmarks. I designed the agent harness that defines how the real-time speech model coordinates with the backend agent and the user, and led development of our general audio understanding model across speech, audio and music. I also built the evaluation infrastructure and shipped the on-device ASR behind Whispers from the Star.
- Jul 2023 — Jan 2025 Research Assistant, CMU WAVLab — Pittsburgh, PA Advised by Prof. Shinji Watanabe. Proposed FastAdaSP for efficient speech LM inference (EMNLP 2024 Oral) and SynesLM for unified audio-visual speech recognition and translation. Core contributor to ESPnet, working on discrete speech unit modeling.
- May 2023 — Aug 2023 Machine Learning Engineer Intern, TrovaAI — Champaign, IL Designed a RAG retrieval pipeline combining vector and keyword indices over heterogeneous documents for an AI agent platform, and built LLM-based and MT-metric (BLEU/COMET) evaluation for the agent system.
- Jul 2021 — Jan 2022 Machine Learning Engineer Intern, VMware, Inc. — Beijing, China Built an automated data analysis and reporting system for BERT-based large-scale machine translation, improving query and analysis throughput via the Elastic Stack.



SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation
Interspeech 2024 Workshop
- Aug 2023 — Dec 2024 Carnegie Mellon University — M.S. in Artificial Intelligence Coursework: Multimodal ML, LLM Systems, System Tool Chains for AI, Speech Recognition and Understanding.
- Aug 2019 — May 2023 University of Illinois Urbana-Champaign — B.S. in Statistics & Computer Science GPA 3.93/4.00. Graduated with Highest Honors; Dean's List.
-
Creator
Proteus — measuring agent self-evolution
(proteus-evolve.github.io)
A harness-agnostic framework for the science of self-evolving agents:
plug in any agent harness × model, let it evolve its own memory, skills,
tools — or its own source code — under goal or no-goal conditions, and
measure how the harness changes with one structural and behavioural
ruler. v0.2.0 ships transactional self-evolution (staged activation with
failed-candidate repair), condition-locked reproducible sweeps, sandboxed
benchmark grading (SWE-bench, Polyglot), and a contract-checked adapter
on-ramp.
pip install proteus-evolve. GitHub PyPI Docs - Founder ViDove — multimodal translation agent (vidove.ai) An end-to-end multimodal translation agent for video subtitle generation, with multimodal context and memory-augmented reasoning (EMNLP 2025 System Demo). Led a 10-person engineering team and productized it as vidove.ai. Deployed in production with the StarCraft II World Team League, one of the largest SC2 e-sports tournaments worldwide. vidove.ai GitHub
- Feb 2014 — Dec 2018 Sugar Masses Creative — China Minecraft Construction Summit Founder, developer and director of the largest official Minecraft tournament in China, drawing over 1,000 contestants and more than a million online viewers a year with a team of 30. Secured long-term cooperation with NetEase, Qihoo 360, Tencent and Youku, and investment from NetEase and JoyMe.com. Presented at ChinaJoy 2016 with more than 325,000 entries. Video 1 Video 2 NetEase coverage
- Technical · 28 May 2026 Some Thoughts on Agents — Model as Harness System
- Personal · 01 Nov 2025 Notes, 1 Nov 2025 — On Growth, Self-Consistency and Love of the Work
Reviewer / Program Committee — IEEE ASRU 2025, IEEE T-ASLP, AAAI 2026, ICASSP.
- A tuxedo cat named Brann.
- Favourite musical — Hamilton.
- Formerly an AMVer, anime music video creator.