Job Description:
Post-Training Pipeline Implementation
- Participate in the development and deployment of post-training pipelines such as SFT, DPO, PPO, and Reward Modeling responsible for concrete execution from training, hyperparameter tuning, evaluation to online regression.
- Participate in training optimization in Agent and Tool-use directions, improving metrics such as tool-calling accuracy, multi-turn instruction following, and hallucination control.
- Participate in iteration of core-scenario models such as Router, Query rewriting, Agent tool calling, Web Search decision-making, Memory, and e-commerce search relevance.
Data and Evaluation
- Manage the collection, cleaning, synthesis, and annotation workflows for post-training data participate in annotation guideline development and vendor coordination.
- Build and maintain automated evaluation combining human review and LLM-as-Judge produce reproducible effectiveness reports.
- Keep up to date with frontier methods in the industry and experiment with them based on the team's direction.
Requirements:
- Master's degree in Computer Science, AI, or a related field.
- Prior development experience in at least one LLM post-training area (SFT, DPO, PPO, RLHF, Reward Modeling, etc.) able to independently complete key steps from data through training to evaluation.
- Prior experience with at least one mainstream training framework (e.g., Megatron, veRL, etc.) and understanding of basic principles of distributed training.
- Understanding of Agent / Tool-use / multi-turn dialogue familiar with data construction and basic alignment approaches for Function Calling.
- Strong engineering and troubleshooting capabilities able to make reasonable trade-offs between effectiveness and iteration efficiency.
- Prior End-to-end post-training deployment experience, or participation in 1000-GPU-scale distributed training.
- Prior experience with leading large model teams will be a strong plus.
- A strong understanding or prior hands on experience in Agent RL, Self-play, synthetic data, or inference-time compute will be a strong plus