Yifei Shao

AI · Machine Learning · Reasoning

Yifei Shao

Building models that see, reason, and act — toward unified multimodal intelligence.

Research Scientist @ Aether AI M.S. CSE @ UC San Diego
5Publications
37Citations
2h-index
2i10-index

5 distinct publications · Citation metrics from Google Scholar · Updated

01

About

I'm an AI researcher focused on unified & multimodal models, large language / vision-language models, and reasoning & planning. My work asks how a single model can perceive across modalities, reason over long chains of thought, and plan actions in the world.

I'm currently a Research Scientist at Aether AI, which I joined in May 2026. I received my M.S. in Computer Science & Engineering from UC San Diego, where I worked with Zhiting Hu and Biwei Huang. I earlier completed my B.S. at the University of Electronic Science and Technology of China (UESTC).

Open to research collaboration, idea exchange, and paper discussions.

Toolbox

  • PyTorch
  • HuggingFace
  • vLLM
  • Diffusion Models
  • Reinforcement Learning
  • ML Systems

Focus

  • Multimodal Reasoning
  • Unified Models
  • Embodied AI
  • World Models
02

Research Directions

🧩

Unified & Multimodal Models

One backbone that jointly understands and generates across vision and language, closing the gap between perception and generation.

🧠

Reasoning & Planning

Eliciting long chain-of-thought and latent reasoning so models can plan, self-correct, and think before they act.

🌍

World Models & Embodied AI

Learning predictive models of the world that let agents imagine, simulate, and act with foresight.

⚙️

Efficient ML Systems

Making large models practical — activation control, memory, and inference systems that scale.

03

Selected Publications

  1. 2026
    CausalWM framework: optical flow, pointmaps, and future video prediction with causal chain-of-thought reasoning arXiv 2026

    CausalWM: Causal Chain-of-Thought Reasoning for Embodied World Model

    Z. Xu, S. Liang, R. Han, Z. Xi, M. Rao, K. Zhou, Z. Zhang, Y. Yan, Y. Wei, J. Huang, Y. Shao, F. Nan, B. Huang

    World ModelsEmbodied AICausal Reasoning arXiv:2609.23184 ↗ Project ↗ Code ↗
  2. 2026
    Modal-Mixed CoT framework figure arXiv 2026

    Learning Modal-Mixed Chain-of-Thought Reasoning with Latent Embeddings

    Y. Shao, K. Zhou, Z. Xu, M. A. Quamar, S. Hao, Z. Wang, Z. Hu, B. Huang

    First authorMultimodalReasoning1 citation arXiv:2602.00574 ↗
  3. 2026
    Decentralized Arena framework figure ACL 2026

    Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

    Y. Yin, K. Zhou, Z. Wang, X. Zhang, Y. Shao, S. Hao, Y. Gu, J. Liu, S. Singla, et al.

    EvaluationLLM2 citations arXiv:2505.12808 ↗
  4. 2025
    Activation Control figure NeurIPS 2025 · Spotlight

    Activation Control for Efficiently Eliciting Long Chain-of-Thought Ability of Language Models

    Z. Zhao, Q. Liu, K. Zhou, Z. Liu, Y. Shao, Z. Hu, B. Huang

    ReasoningLLM10 citations arXiv:2505.17697 ↗
  5. 2025
    Continuous Memory (CoMEM) framework figure NeurIPS 2025

    Towards General Continuous Memory for Vision-Language Models

    W. Wu, Z. Song, K. Zhou, Y. Shao, Z. Hu, B. Huang

    MultimodalMemory22 citations arXiv:2505.17670 ↗
Full list on Google Scholar ↗
04

News

05

Get in touch

Always happy to talk research, collaborations, or new ideas.