Yisheng He (何益升)

Yisheng He is a research scientist at Tongyi Lab, Alibaba (Alibaba Star Project). Before that, he obtained his Ph.D. at HKUST.

We are hiring research interns. To apply, please email your CV to ethanheysh@gmail.com.

Portrait of Yisheng He

Research

I'm interested in Embodied AI, 3D Computer Vision, AIGC, and Digital Avatar.

* denotes equal contribution; ^ denotes intern student; † denotes project leader; ✉ denotes corresponding author.

Human-to-robot scale alignment followed by causal state reconstruction in EgoAlign.

EgoAlign: Bridging the Human-Humanoid Gap for Long-Range Loco-Manipulation

Yiming Jiang^, Jin Chen, Chongyang Xu, Yilun Chen†, Aimin Hao✉, Yisheng He†✉

Technical Report, 2026

EgoAlign adapts egocentric human demonstrations through controller-aware motion alignment and causal robot-state reconstruction. This enables human-only task training of vision-language-action models and zero-shot deployment for long-range humanoid loco-manipulation.

Attention maps for two facial query points, comparing models with and without regularization.

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head

Yingdong Hu*^, Yisheng He*✉, Yiming Jiang, Zehong Lin, Steven Hoi, Jun Zhang

NeurIPS, 2026

FA-LAM is a focus-aware large avatar model for one-shot animatable 3D Gaussian head and streaming 4D full-head reconstruction, with symmetric semantic attention regularization, a dual-phase training pipeline, and an autoregressive design with visibility-gated fusion.

Three scenes shown as photographs, sparse depth observations, completed depth maps, and colored 3D reconstructions.

LDCM: Large Depth Completion Model from Sparse Observations

Zhu Yu, Zhengyi Zhao, Runmin Zhang, Lingteng Qiu, Kejie Qiu, Yisheng He, Siyu Zhu, Zilong Dong, Si-Yuan Cao, Hui-liang Shen

ICLR, 2026

LDCM is a transformer-based framework for metric-depth completion from sparse observations, with Poisson depth initialization and a point-map head to regress per‑pixel 3D coordinates without camera intrinsics.

Three upper-body avatar examples shown with changing expressions and hand gestures.

ViSA: 3D-Aware Video Shading for Real-Time Upper-Body Avatar Creation

Fan Yang*, Heyuan Li*, Peihao Li, Weihao Yuan, Lingteng Qiu, Chaoyue Song, Cheng Chen, Yisheng He, Shifeng Zhang, Xiaoguang Han, Steven Hoi, Guosheng Lin

Preprint, 2025

ViSA integrates 3D reconstruction priors with a real-time autoregressive video diffusion model to generate photorealistic, temporally coherent upper-body avatars from a single image for gaming and VR.

Token grids illustrate spatial attention followed by temporal attention, above a flattened token sequence.

MoGenTS: Motion Generation based on Spatial-Temporal Joint Modeling

Weihao Yuan*, Yisheng He*, Weichao Shen, Yuan Dong, Xiaodong Gu, Zilong Dong, Liefeng Bo, Qixing Huang

NeurIPS, 2024

We introduce a 2D joint VQ-VAE to quantize each joint instead of all joints into tokens. A spatial-temporal modeling framework with temporal-spatial 2D masking and 2D attention is also proposed for motion generation.

Academic Challenge

Experience

  1. Tongyi Lab, Alibaba Group

    Research Scientist

    – Present

    Alibaba Star Program, 1%

  2. Megvii Technology (Face++)

    Senior Research Intern in Computer Vision and Robotics

    –

    Supervisor: Dr. Jian Sun, Chief Scientist, Megvii Research

    Mentor: Haoqiang Fan, Megvii Research

    Collaborator: Dr. Haibin Huang, Megvii Research

    Research Intern in Computer Vision

    –

    Mentors: Haoqiang Fan and Dr. Yuzhi Wang, Megvii Research

  3. Microsoft

    Software Development Engineer Intern

    Mentors: Raymond Xue and Hao Lin, Microsoft

Services

Conference Reviewer
  • IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
  • IEEE/CVF International Conference on Computer Vision (ICCV)
  • European Conference on Computer Vision (ECCV)
  • Conference on Neural Information Processing Systems (NeurIPS)
  • International Conference on Learning Representations (ICLR)
  • ACM SIGGRAPH Conference on Computer Graphics and Interactive Techniques (SIGGRAPH)
  • ACM SIGGRAPH Conference and Exhibition on Computer Graphics and Interactive Techniques in Asia (SIGGRAPH Asia)
  • AAAI Conference on Artificial Intelligence (AAAI)
  • ACM International Conference on Multimedia (ACM MM)
  • IEEE International Conference on Robotics and Automation (ICRA)
  • IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
Journal Reviewer
  • IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI)
  • International Journal of Computer Vision (IJCV)
  • IEEE Transactions on Visualization and Computer Graphics (TVCG)
  • IEEE Robotics and Automation Letters (RA-L)
  • Neurocomputing
Teaching @ HKUST
  • COMP 4201 (Spring 2019)
  • COMP 1029 (Fall 2020)
  • COMP 4201 (Spring 2021)