Optomi, in partnership with a large telecommunications company, is seeking a Principal Inference Software Engineer to join their team in the Plano, TX Area. The ideal candidate will have experience building, deploy, and scale production AI inference infrastructure. You’ll work across Kubernetes, vLLM, Python, GPUs, and cloud/on-prem environments to make model serving reliable, scalable, and highly automated.
What The Right Professional Will Enjoy
- Building and operating LLM inference infrastructure at scale.
- Deploying and managing models using vLLM and Kubernetes.
- Developing Python automation for model deployment, scaling, updates, rollbacks, and health monitoring.
- Managing GPU capacity, performance, reliability, and production workloads.
- Building tools that allow engineering teams to deploy and access models without managing underlying infrastructure.
- Supporting both R&D and production inference environments.
- Troubleshooting and optimizing model-serving performance and reliability.
Apply Today If Your Background Includes
- Strong software engineering and Python skills.
- Deep hands-on Kubernetes experience.
- Production experience with LLM inference/model serving, ideally vLLM.
- Experience with GPU infrastructure, containers, and CI/CD.
- Understanding of scaling, automation, observability, and reliability for distributed systems.
- Experience with cloud and/or on-prem infrastructure preferred.