Job Responsibilities
- Pre-training Strategy & Architecture
- Pre-training Data Engineering
- Large-scale Distributed Training
- Long-context Training
- Training Monitoring & Optimization
- Evaluation & Iterative Optimization
Job Requirements
- Bachelor's degree or above in Computer Science, Artificial Intelligence, NLP, Machine Learning, Distributed Systems, or a related field.
- LLM Pre-training Experience
- Hands-on experience with complete LLM pre-training projects.
- Experience participating in the pre-training of 7B+ parameter models.
- Distributed Training
- Strong expertise in large-scale distributed training frameworks such as Megatron
- LM, DeepSpeed, or FSDP.
- Practical experience with 64+ GPU training environments.
- Pre-training Data Engineering
- Solid understanding of large-scale pre-training data pipelines, including data cleaning, deduplication, quality filtering, tokenization, data mixing, and data quality optimization.
- Training Monitoring & Debugging
- Strong ability to analyze training loss, gradients, convergence, and training stability.
- Experience troubleshooting large-scale distributed training issues.
- Long-context Training
- Familiarity with long-context training and extension techniques, including RoPE scaling, NTK-aware interpolation, and YaRN.
- Preferred QualificationsExperience with 70B+ parameter model pre-training.
- Experience with MoE (Mixture-of-Experts) model pre-training.Publications in top-tier AI/ML conferences such as NeurIPS, ICML, ICLR, ACL, or EMNLP, particularly in LLM pre-training, model architecture, or training optimization.
- Experience optimizing large-scale GPU clusters and training infrastructure.
Pay: $5,000.00 - $10,000.00 per month
Benefits:
- Additional leave
- Food provided
- Health insurance
- Professional development
Work Location: In person