I am currently a third-year Ph.D. candidate in Artificial Intelligence at The Hong Kong University of Science and Technology (HKUST), supervised by Prof. Yike Guo and Prof. Wenhan Luo. Before joining HKUST, I received my M.Sc. in Computer Science and Technology and B.Eng. in Communication Engineering from Xidian University. During my master's studies, I was advised by Prof. Licheng Jiao.

I have conducted research internships with the Kling Team at Kuaishou, the Tongyi Wan Team at Alibaba, and AIR, Tsinghua University, where I worked with Prof. Hao Zhao and Prof. Guyue Zhou.

Research Interests

My research lies at the intersection of multimodal understanding and controllable visual generation, with a focus on:

  • Controllable video generation and editing
  • Unified multimodal understanding and generation
  • Post-training and reinforcement learning for multimodal models

I am open to collaborations in any form. I am also actively seeking internship opportunities in industry or academia. If you would like to collaborate or have a suitable opportunity, please feel free to reach out.

News

  • Our work ReBind on multi-reference video editing is now available.
  • Our paper Lumos-Nexus was accepted to ECCV 2026.
  • Our paper HiPrompt was accepted by IJCV in 2026.
  • Our work AesRM on expert-aligned video aesthetics is now available.
  • Our work DreamVideo-Omni on controllable multi-subject video customization is now available.
  • Our paper Turning Internal Gap into Self-Improvement was accepted to ICLR 2026.
  • Our work ReViSE on reason-informed video editing is now available.
  • Our work VFX Creator on controllable animated visual-effect generation is now available.
  • Our paper LiON was accepted to AAAI 2025.
  • Our paper TOD3Cap was accepted to ECCV 2024.
  • I commenced my doctoral journey at HKUST under the supervision of Prof. Yike Guo and Prof. Wenhan Luo.

Publications

Google Scholar

* denotes equal contribution.

ReBind instruction generation and reward design overview

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

Xinyu Liu, Shihao Li, Weihong Lin, Xinlong Chen, Yang Shi, Yujin Han, Yiyang Cai, Yanghao Wang, Ruibin Yuan, Yuanxing Zhang, Pengfei Wan, Wenhan Luo, Yike Guo

arXiv, 2026

Lumos-Nexus video demo

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models

Jiazheng Xing, Hangjie Yuan, Lingling Cai, Xinyu Liu, Yujie Wei, Fei Du, Tao Feng, Hai Ci, Jiasheng Tang, Weihua Chen, Fan Wang, Yong Liu

European Conference on Computer Vision (ECCV), 2026

HiPrompt overview

HiPrompt: Tuning-Free Higher-Resolution Generation with Hierarchical MLLM Prompts

Xinyu Liu, Yingqing He, Lanqing Guo, Xiang Li, Bu Jin, Peng Li, Yan Li, Chi-Min Chan, Qifeng Chen, Wei Xue, Wenhan Luo, Qifeng Liu, Yike Guo

International Journal of Computer Vision (IJCV), 2026

AesRM video aesthetics comparison demo

AesRM: Improving Video Aesthetics with Expert-Level Feedback

Yujin Han, Yujie Wei, Yefei He, Xinyu Liu, Tianle Li, Zichao Yu, Andi Han, Shiwei Zhang, Tingyu Weng, Difan Zou

arXiv, 2026

DreamVideo-Omni multi-subject customization demo

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

Yujie Wei*, Xinyu Liu*, Shiwei Zhang, Hangjie Yuan, Jinbo Xing, Zhekai Chen, Xiang Wang, Haonan Qiu, Rui Zhao, Yutong Feng, Ruihang Chu, Yingya Zhang, Yike Guo, Xihui Liu, Hongming Shan

arXiv, 2026

Generation-understanding unification analysis

Turning Internal Gap into Self-Improvement: Promoting Generation–Understanding Unification in MLLMs

Yujin Han, Hao Chen, Andi Han, Zhiheng Wang, Xinyu Liu, Yingya Zhang, Shiwei Zhang, Difan Zou

International Conference on Learning Representations (ICLR), 2026

ReViSE framework

ReViSE: Towards Reason-Informed Video Editing in Unified Models with Self-Reflective Learning

Xinyu Liu, Hangjie Yuan, Yujie Wei, Jiazheng Xing, Yujin Han, Jiahao Pan, Yanbiao Ma, Chi-Min Chan, Kang Zhao, Shiwei Zhang, Wenhan Luo, Yike Guo

arXiv, 2025

VFX Creator results

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer

Xinyu Liu, Ailing Zeng, Wei Xue, Harry Yang, Wenhan Luo, Qifeng Liu, Yike Guo

arXiv, 2025

LiON LiDAR outlier detection results

LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier Detection Using Diverse Synthetic Data

Shaocong Xu, Pengfei Li, Qianpu Sun, Xinyu Liu, Yang Li, Shihui Guo, Zhen Wang, Bo Jiang, Rui Wang, Kehua Sheng, Bo Zhang, Li Jiang, Hao Zhao, Yilun Chen

AAAI Conference on Artificial Intelligence (AAAI), 2025

TOD3Cap overview

TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes

Bu Jin, Yupeng Zheng, Pengfei Li, Weize Li, Yuhang Zheng, Sujie Hu, Xinyu Liu, Jinwei Zhu, Zhijie Yan, Haiyang Sun, Kun Zhan, Peng Jia, Xiaoxiao Long, Yilun Chen, Hao Zhao

European Conference on Computer Vision (ECCV), 2024

SAZS framework

Delving into Shape-Aware Zero-Shot Semantic Segmentation

Xinyu Liu, Beiwen Tian, Zhen Wang, Rui Wang, Kehua Sheng, Bo Zhang, Hao Zhao, Guyue Zhou

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

ADAPT driving caption example

ADAPT: Action-Aware Driving Caption Transformer

Bu Jin, Xinyu Liu, Yupeng Zheng, Pengfei Li, Hao Zhao, Tong Zhang, Yuhang Zheng, Guyue Zhou, Jingjing Liu

IEEE International Conference on Robotics and Automation (ICRA), 2023

Research Experience

Academic Service & Honors

Academic Service

Reviewer: NeurIPS, ICLR, CVPR, ICML, ECCV, AAAI, IJCAI, and IEEE Transactions on Image Processing.

Teaching Assistant: AMCC5110 Programming for Arts and Creativity; EMIA6500(G) AI Animation and Video Generation.

Honors & Awards

  • National Scholarship for Graduate Students, 2021
  • CETC Leis Scholarship, Top 1%, 2021
  • Mathematical Contest in Modeling, Meritorious Winner