
Pengjun Fang
Incoming MPhil Student, The Hong Kong University of Science and Technology
Advised by Prof. Qifeng Chen
I work on multimodal generation, video-to-audio synthesis, and controllable video generation.
About
I am an incoming MPhil student in Computer Science at HKUST, advised by Prof. Qifeng Chen. I received my B.Eng. in Computer Science from HKUST, and spent an exchange semester at EPFL, Switzerland.
My research focuses on generative AI, with an emphasis on multimodal generation, video-to-audio synthesis, and controllable video generation. I am interested in building models that connect perception across modalities and give users precise, interpretable control over generated content.
News
- Joining HKUST as an MPhil student, advised by Prof. Qifeng Chen.
- AC-Foley accepted to ICLR 2026.
- Finished research internship at Everlyn Labs Inc.
- Text-Driven Portrait Image Animation accepted to ICCV 2025 Workshop.
- Started exchange semester at EPFL, Switzerland.
Publications
* Equal contribution · † Corresponding author
Experience
Everlyn Labs Inc. — Research Intern
Feb 2025 – Aug 2025Developed AC-Foley, a reference-audio-guided video-to-audio synthesis system. Built WanFM, a First–Last–Frame-to-Video pipeline based on Wan 2.2 with bidirectional denoising for stronger temporal consistency.
Selected Projects
VideoTuna
Open-source codebase for text-to-video generationUnified codebase integrating multiple AI video generation models across text-to-video, image-to-video, and text-to-image. Provides end-to-end pipelines for pre-training, continual training, post-training alignment, and fine-tuning.
WanFM
First–Last–Frame-to-Video generationBuilt on Wan 2.2 Image-to-Video with last-frame constraints, bidirectional denoising, and prompt-adapted attention for controllable FLF2V generation.
Contact
Feel free to reach out by email at pfangaf@connect.ust.hk for research discussions or collaborations.



