I’m always open to research collaborations and happy to mentor motivated students.
If you’re interested in the following areas, feel free to reach out!
Jinpeng Yu (于劲鹏) is currently the Content Generation Algorithm Lead for the Qwen App at Alibaba Group, where he focuses on multimodal AI agent systems and generative models for image and video creation. Previously, he was a Senior Researcher on the AIGC team at Xiaohongshu, where he worked on controllable image generation and multimodal understanding. He received his bachelor's degree with honors from Harbin Institute of Technology (HIT) in 2021. Subsequently, he earned his master's degree from SIST, ShanghaiTech University in 2024 under the supervision of Prof. Shenghua Gao (高盛华, HKU).
* indicates equal contribution. † indicates project lead.
A streaming framework that jointly replaces visual identity and vocal timbre in talking videos while preserving motion, scene content, and audio-video synchronization.
Real-time, causal human animation from a reference image and streaming pose controls, designed to remain stable over long-form generation.
An agentic video auto-encoding framework that transforms films into structured knowledge graphs and learns agent-native video representations through reconstruction-driven optimization.
Introduces branch-aware Positive-Direction Matching to address negative-branch asymmetry and improve guidance-scale robustness in on-policy diffusion distillation.
A three-stage multimodal framework for recommending useful, diverse, and visually consistent follow-up edits in image-creation conversations.
An agentic visual generation framework that retrieves external visual knowledge for knowledge-intensive prompts and decides when search is necessary.
Encodes selective subject representation from reference images for subject-driven image generation.
Pseudo-plane regularization for high-fidelity SDF-based reconstruction of texture-less indoor scenes.
A tri-plane integrated transformer for accurate and detailed point cloud completion.