Job Responsibilities
- Develop LLM pre-training strategies and model architectures.
- Build and optimize large-scale pre-training data pipelines.
- Conduct distributed training using Megatron-LM, DeepSpeed, or FSDP.
- Support 64+ GPU training environments.
- Work on long-context training and techniques such as RoPE scaling, NTK, and YaRN.
- Monitor training performance, troubleshoot issues, and optimize model efficiency.
- Evaluate models and drive iterative performance improvements.
Job Requirements
- Bachelor's degree or above in Computer Science, AI, NLP, ML, Distributed Systems, or related fields.
- Hands-on experience with LLM pre-training projects.
- Experience with 7B+ parameter models.
- Strong experience in large-scale distributed training and 64+ GPU environments.
- Knowledge of large-scale data processing, including cleaning, deduplication, filtering, tokenization, and data mixing.
- Strong understanding of training loss, gradients, convergence, and stability.
- Experience troubleshooting distributed training issues.
- Familiarity with long-context training and extension techniques.
Preferred Qualifications
- Experience with 70B+ models or MoE pre-training.
- Experience optimizing GPU clusters and AI training infrastructure.
- Publications in NeurIPS, ICML, ICLR, ACL, EMNLP, or similar conferences.
Pay: $5,000.00 - $10,000.00 per month
Benefits:
- Food provided
- Health insurance
- Professional development
- Work from home
Work Location: In person