I am an undergraduate student in the School of Artificial Intelligence at Nanjing University, expected to join the LAMDA Lab as a graduate student under the supervision of Prof. Peng Zhao. My research centers on efficient inference for large language models, with a particular focus on speculative decoding, online learning, and diffusion language models. I am also broadly interested in LLM post-training and scalable inference systems. I currently work as an inference optimization intern at StepFun, focusing on efficient distributed training and deployment of speculative-decoding draft models.