I am a faculty member at the Institute of Trustworthy Embodied AI (TEAI) at Fudan University and a full-time supervisor at the Shanghai Innovation Institute (上海创智学院全时导师). I also serve as an Honorary Fellow at the University of Melbourne, Australia. My research focuses on Trustworthy AI and AI Safety, with particular interests in the safety of foundation models, AI agents, and embodied AI systems. Beyond my core research, I am deeply interested in leveraging AI to advance our understanding of both the mind and the universe.
I received my Ph.D. from the University of Melbourne, where I subsequently spent two wonderful years as a postdoctoral research fellow. Before joining Fudan University, I was a lecturer at Deakin University for ~2 years. I received my bachelor's degree from Jilin University and my master's degree from Tsinghua University.
"Everything should be as simple as possible, but not simpler."
Email / Google Scholar / GitHub
We welcome motivated PhD applicants to join our lab through the Shanghai Innovation Institute Summer Camp. 欢迎优秀的同学通过上海创智学院夏令营加入我们。 We also welcome overseas undergraduates to join us through direct Ph.D. programs, especially those with clear interests and strong hands-on skills. 欢迎海外本科生通过直博项目加入我们,尤其是兴趣明确、动手能力强的同学。
Introducing OpenTAI: The Open Hub for Trustworthy AI We are building OpenTAI, an open platform for sharing datasets, benchmarks, models, arenas, and tools to advance trustworthy AI. Contributions are welcome!
Books
- Endogenous Safety in Artificial Intelligence (《人工智能内生安全》)
- Artificial Intelligence: Data and Model Safety (《人工智能:数据与模型安全》)
Surveys
- Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
- Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses
- Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
News
- [7/2026] Four papers have been accepted to ACM MM 2026.
- [5/2026] One paper on fairness has been accepted to KDD 2026.
- [5/2026] Seven papers have been accepted to ICML 2026.
- [4/2026] Three papers have been accepted to ACL 2026.
- [2/2026] Two papers have been accepted to CVPR 2026.
- [1/2026] Two/Two/One/One papers have been accepted to ICLR 2026/TPAMI/FCS/WWW 2026, respectively.
- [11/2025] Two papers have been accepted to AAAI 2026.
- [09/2025] Five papers accepted to NeurIPS 2025. Congrats to all the authors!
- [09/2025] Our survey paper Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety has been published in Foundations and Trends® in Privacy and Security.
- [08/2025] Our work on multi-trigger backdoor attacks has been accepted to TDSC.
- [08/2025] Our paper VeriFi: Towards Verifiable Federated Unlearning has been selected as the Runner-up for the 2024 Best Paper Award by the IEEE TDSC journal. Congratulations to all co-authors!
- [07/2025] Four papers have been accepted to ACM Multimedia 2025.
- [06/2025] Three papers have been accepted to ICCV 2025.
- [05/2025] Our BackdoorLLM Benchmark received First Prize in the SafeBench Competition, organized by the Center for AI Safety. Congratulations to all co-authors!
- [05/2025] Our work on super transferable attacks X-Transfer Attacks has been accepted to ICML 2025.
- [03/2025] The preprint of our long survey paper Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety is available on arXiv. Many thanks to all collaborators!
- [02/2025] Our works on Million-scale Adversarial Robustness Evalution, Test-time Adversarial Prompt Tuning, and AnyAttack have been accepted to CVPR 2025.
- [01/2025] Our works on RL-based jailbreak defense for VLMs and backdoor sample detection in CLIP have been accepted to ICLR 2025.
- [12/2024] I will serve as an Area Chair for ICML 2025.
- [12/2024] Our works on targeted transferable adversarial attack, defense against model extraction attacks, and RL-based LLM auditing have been accepted to AAAI 2025.
- [09/2024] I will serve as an Area Chair for ICLR 2025.
- [09/2024] One paper on unlearnable examples for segmentation models has been accepted to NeurIPS 2024.
- [07/2024] Our works on model lock , detecting query-based adversarial attacks , and multimodal jailbreak attacks on VLMs have been accepted to MM 2024.
- [07/2024] Our work on adversarial prompt tuning has been accepted to accepted by ECCV 2024.
- [04/2024] Our work on intrinsic motivation for RL has been accepted to IJCAI 2024.
- [03/2024] Our work on adversarial policy learning in RL is accepted by DSN 2024.
- [03/2024] Our work on safety alignment of LLMs is accepted by NAACL 2024.
- [03/2024] Our work on federated machine unlearning has been accepted to TDSC.
- [01/2024] Our work on self-supervised learning have been accepted to ICLR 2024.
Research Areas
- Trustworthy AI
- LLM/MLLM safety
- Agentic safety
- Embodied AI safety
- Reinforcement learning
- Privacy, fairness, memorization
Professional Activities
- Program Committee Member
- ICLR (2019-2026), ICML (2019-2026), NeurIPS (2019-2026), CVPR (2020-2026), ICCV (2021-2026), ECCV (2020), AAAI (2020-2022), IJCAI (2020-2021), KDD (2019,2021), ICDM (2021), SDM (2021), AICAI (2021)
- Journal Reviewer
- Nature Communications, Pattern Recognition, TPAMI, TIP, IJCV, JAIR, TNNLS, TKDE, TIFS, TOMM, KAIS