We're hiring on behalf of a Haystack partner!
The Role
- Optimize collective operations for AWS Trainium, focusing on AI compute scaling
- Enhance collective algorithms and topologies for optimal training performance
- Work closely with hardware teams to co-optimize software and Trainium silicon
- Develop and optimize C/C++ implementations of collective communication patterns
- Investigate and implement improvements for specific training topologies used by modern LLMs
- Monitor and analyze processor, DMA, firmware, and workload metrics
What You'll Need
- Experience building complex software systems successfully delivered to customers
- Background in architecture and design of new and current systems
- Bachelor's degree in Computer Science or equivalent
- Knowledge of the full software/hardware/networks development life cycle
- Experience in development within the last 3 years, or embedded C/C++ development
- Familiarity with collective communication algorithms or distributed training frameworks (preferred)
What's On Offer
- Opportunity to impact AI training at AWS scale
- Work on cutting-edge AI training hardware
- Culture of inclusion, work-life balance, and career growth
- Comprehensive benefits package including health insurance, 401(k) matching, and paid time off
Apply via Haystack today!