News

Jul 08, 2026 AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs’ Contextual and Cultural Knowledge and Thinking Accepted to COLM 2026
Jun 03, 2026 My co-authored MeanFlow-TSE: One-Step Generative Target Speaker Extraction with Mean Flow Accepted to INTERSPEECH 2026
May 31, 2026 My contributing Video Generation Models: A Survey of Post-Training and Alignment Accepted to Transactions on Machine Learning Research (TMLR)
Jan 27, 2026 AVMeme Exam is out: A Multimodal Multilingual Multicultural Benchmark for LLMs’ Contextual and Cultural Knowledge and Thinking
Jan 17, 2026 My mentored paper SightSound-R1: Cross-Modal Reasoning Distillation from Vision to Audio Language Models Accepted to ICASSP 2026
Jan 13, 2026 My MSR intern paper Sci-Phi: A Large Language Model Spatial Audio Descriptor Accepted to IEEE Open Journal of Signal Processing
Nov 07, 2025 DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Accepted to AAAI 2026
Oct 14, 2025 Bridging Ears&Eyes cross audio&visual LLM distill Won the Best Paper🥇 in WASPAA 2025