Xu Guo | Fudan University

PhD Student · School of Computer Science, Fudan University · Large Language Models and Data-Centric AI

guoxu-chuanxi.jpg

I am Xu Guo (郭旭), a third-year PhD student at the School of Computer Science, Fudan University, working on large language models and data-centric artificial intelligence.

I believe that training data, including synthetic data, is a fundamental driver of modern AI systems. My research focuses on data-centric methods for language model training, spanning pre-training, post-training data synthesis, and agentic training. My recent first-author work explores robust pre-pre-training under noisy data, data and reward design for reinforcement learning with verifiable rewards, and instruction-following post-training.

This official academic homepage collects my publications and research updates. You can also find my work and citation record on Xu Guo’s Google Scholar profile, with additional publication and identity records on DBLP, ORCID, and GitHub. You can reach me at guox24@m.fudan.edu.cn.

news

Jul 03, 2026 Our new preprint, When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling, is now available on arXiv.
Jul 02, 2026 I launched my personal academic homepage.

selected publications

  1. When Does Generating More Help? Disentangling Fixed-Source Synthesis from Source Expansion in Synthetic Data Scaling
    Xu Guo, Jian Tong, Zhihui Lu, and Qipeng Guo
    arXiv preprint arXiv:2607.01727, 2026
  2. Synthetic Pre-Pre-Training Improves Language Model Robustness to Noisy Pre-Training Data
    Xu Guo, Runyu Peng, Jian Tong, Yunhua Zhou, Haijun Lv, Zhihui Lu, and Qipeng Guo
    arXiv preprint arXiv:2605.10129, 2026
  3. Rethinking Multiple-Choice Questions for RLVR: Unlocking Potential via Distractor Design
    Xu Guo, Qiming Ge, Jian Tong, Kedi Chen, Jin Zhang, Xiaogui Yang, Xuan Gao, Haijun Lv, Zhihui Lu, Yicheng Zou, and 1 more author
    In Findings of the Association for Computational Linguistics: ACL 2026, 2026
  4. IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
    Xu Guo, Tianyi Liang, Tong Jian, Xiaogui Yang, Ling-I Wu, Chenhui Li, Zhihui Lu, Qipeng Guo, and Kai Chen
    arXiv preprint arXiv:2508.04632, 2025
  5. Towards transferable adversarial attacks on vision transformers for image classification
    Xu Guo, Peng Chen, Zhihui Lu, Hongfeng Chai, Xin Du, and Xudong Wu
    Journal of Systems Architecture, 2024