🎓 I'm a graduate student majoring in Software Engineering at the University of Science and Technology of China.
- 🔬 I work on LLM post-training — SFT, GRPO/PPO, and RL for multi-turn tool-using agents.
- 🧪 Currently building an agentic RL pipeline on τ2-bench: failure attribution → trajectory distillation → SFT → GRPO, with ablations on reward shaping and KL anchoring.
- ⚙️ Also interested in the systems side: vLLM rollout, FSDP training, weight sync, and memory optimization.
- 💬 Ask me about GRPO/PPO internals, agent evaluation & failure analysis, Python, C++, algorithms.
- 📫 Reach me at fangchengjie@outlook.com.