I am currently a forth-year PhD student at the Fudan University, NLP Lab, supervised by Prof. Xipeng Qiu. I graduated from Huazhong University of Science and Technology with a bachelor’s degree in software engineering.

My current research focuses on unified multimodal models. Feel free to contact me via email at jzhan24@m.fudan.edu.cn.

On the job market: I expect to graduate in June 2027 and am seeking job opportunities in the United States, China, and Europe. Please feel free to contact me at jzhan24@m.fudan.edu.cn.

CV (English) · 简历(中文)

🔥 News

  • 2026.07:  🎬 We release OmniVAE, an audio-video VAE with cross-modal alignment for joint generation!
  • 2026.06:  💼 I joined Meta FAIR as a research intern!
  • 2025.09:  🎙️ We release VStyle, a benchmark for voice style adaptation in speech language models.
  • 2025.06:  🗣️ We release MOSS-TTSD, a bilingual spoken dialogue synthesis model.
  • 2025.05:  💼 I joined Alibaba’s Future Life Lab as a research intern!
  • 2025.02:  🤖 We release SpeechGPT2.0-preview, a human-like real-time interaction system.
  • 2024.09:  🎓 I became a phd student at FudanNLPLab!
  • 2024.07:  🤖 We released SpeechGPT2, a emotional intelligent end-to-end spoken dialogue LLM.
  • 2024.05:  🎉 One first-author paper accepted to ACL 2024!
  • 2024.03:  🤖 We release the model, data, and code of AnyGPT. Welcome to STAR and FORK!
  • 2023.05:  🤖 We released SpeechGPT, a conversational speech large language model.
  • 2022.09:  🎓 I joined FudanNLPLab as a master student.

📝 Publications

2026
OmniVAE architecture

OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Jun Zhan*, Chen Yang*, Yitian Gong*, Donghua Yu, Kuangwei Chen, Wenbo Zhang, Kexin Huang, Qi Luo, Zhe Xu, Ying Zhu, Jin Wang, Tengyue Zhang, Qi Chen, Cheng Chang, Songlin Wang, Junqi Dai, Jiasheng Ye, Xiaogui Yang, Tianyi Liang, Xiangyu Peng, Zhaoye Fei, Shimin Li, Qinyuan Cheng, Xie Chen, Xinchi Chen, Xipeng Qiu
Project

  • Role: Project Lead
  • OmniVAE is a jointly trained audio-video VAE that learns fine-grained cross-modal alignment, improving generation quality and audio-video synchronization without additional inference cost.
ICASSP 2026
VStyle voice style categories and instructions

VStyle: A Benchmark for Voice Style Adaptation with Spoken Instructions
Jun Zhan*, Mingyang Han*, Yuxuan Xie*, Chen Wang, Dong Zhang, Kexin Huang, Haoxiang Shi, DongXiao Wang, Tengtao Song, Qinyuan Cheng, Shimin Li, Jun Song, Xipeng Qiu, Bo Zheng
Project

  • Role: Project Lead
  • VStyle is a benchmark for voice style adaptation in speech language models, with automated evaluation via LALM.
ACL 2024
AnyGPT model architecture

AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Jun Zhan*, Junqi Dai*, Jiasheng*, Yunhua Zhou*, Dong Zhang, Zhigeng Liu, Xin Zhang, Ruibin Yuan, Ge Zhang, Linyang Li, Hang Yan, Jie Fu, Tao Gui, Tianxiang Sun, Yu-Gang Jiang1, Xipeng Qiu

Project

  • Role: Project Lead
  • AnyGPT is the first any-to-any multimodal LLM based on discrete representations.
  • Academic Impact: This work has received 800+ stars on GitHub, promoted by more than 10 media and forums, such as 机器之心
EMNLP 2023
sym

SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities
Dong Zhang, Shimin Li, Xin Zhang, Jun Zhan, Pengyu Wang, Yaqian Zhou, Xipeng Qiu

Project

  • SpeechGPT is the first LLM with intrinsic cross-modal generation abilities.
  • Academic Impact: This work has received 1.3k+ stars on GitHub, promoted by more than 10 media and forums, such as 机器之心

🛠️ Projects

sym

MOSS: A Conversational Language Model

  • Tianxiang Sun, Xiaotian Zhang, Zhengfu He, Peng Li, Qinyuan Cheng, Hang Yan, Xiangyang Liu, Yunfan Shao, Qiong Tang, Xingjian Zhao, Ke Chen, Yining Zheng, Zhejian Zhou, Ruixiao Li, Jun Zhan, Yunhua Zhou, Linyang Li, Xiaogui Yang, Lingling Wu, Zhangyue Yin, Xuanjing Huang, Xipeng Qiu.
  • 🌟 MOSS is the first ChatGPT-like LLM in China, and fully open sourced with 12k+ stars. GitHub
sym

SpeechGPT 2.0-preview: End-to-End Human-Like Spoken Chatbot

  • Dong Zhang, Qian Tu*, Ruifan Deng*, Jun Zhan*, Zixin Wang*, Xingjian Zhao, Ke Chen, Xin Zhang, Pengyu Wang, Zhaowei Li, Shimin Li, Yaqian Zhou, Xipeng Qiu.
  • 🌟 An end-to-end spoken dialogue LLM that directly models audio and supports expressive conversations across emotions, speaking styles, and voices, together with natural real-time interaction.
  • Project · Code

🎖 Honors and Awards

  • 2024.08: 😄 The Best Poster at the 3rd HIT-SCIR & THUNLP & FudanNLP Academic Symposium.
  • 2024.12: ⚽️ The runner-up of the 2024 Fudan University Graduate Football Tournament.

📖 Educations

  • 2024.09 - (now), Phd student, Fudan University, Shanghai.
  • 2022.09 - 2024.06, Master, Fudan University, Shanghai.
  • 2018.09 - 2022.06, Undergraduate, Huazhong University of Science and Technology, Wuhan.

💻 Internships

  • 2026.06 - Present, Research Intern, Meta FAIR.
  • 2025.09 - 2025.11, Industry-Academia Project Intern, OPPO.
  • 2025.05 - 2025.08, Research Intern, Alibaba Future Life Lab.