SMC-ITA: Sequential Monte Carlo Inference-Time Alignment for Video-to-Audio Generation
APSIPA ASC 2026 [PDF]
I am a Ph.D. student in the Department of Technology Management for Innovation at the University of Tokyo, Japan, supervised by Prof. Yutaka Matsuo.
My research focuses on reinforcement learning for audio understanding and reward-guided alignment for audio generation. I am also interested in unified generation of speech and non-speech audio, as well as joint audio–video generation.
APSIPA ASC 2026 [PDF]
Interspeech 2026 [PDF]
ACL 2026 (Findings) [PDF]
ACM Multimedia 2023 [PDF]
Interspeech 2023 [PDF]
–Present
Ph.D. student in Technology Management for Innovation
–
Speech Algorithm Engineer
–
Master’s in Electrical Engineering and Information Systems