Platform Engineer - AI/ML Infrastructure
About the Role:
We're seeking a skilled Platform Engineer to deploy and manage our AI/ML infrastructure, e.g LlamaIndex Cloud and KDB.AI applications on Kubernetes Platform. You'll be responsible for building reliable, scalable, and secure deployment pipelines using modern GitOps practices.
What You'll Do:
- Deploy and manage LlamaIndex Cloud and KDB.AI applications and similar products supporting AI workloads across Dev/QA/Prod environments
- Implement GitOps workflows using Flux CD for automated deployments
- Administer Kubernetes clusters across multiple environments
- Configure HashiCorp Vault for secrets management and ExternalSecrets integration
- Maintain CI/CD pipelines with GitHub Actions and Artifactory
- Work with enterprise Identity Management team to configure OIDC authentication with Microsoft Entra ID
- Create and maintain Helm charts and Kubernetes manifests
- Monitor application performance and troubleshoot production issues
- Document procedures, runbooks, and infrastructure patterns
Required Skills:
Core Technologies:
- 3+ years managing production Kubernetes clusters
- 2+ years with Flux CD, ArgoCD, or similar GitOps tools
- Advanced Helm chart development and management
- HashiCorp Vault for secrets management
- Artifactory or similar container registries
- CI/CD with GitHub Actions, Jenkins, or similar
Infrastructure & Database:
- PostgreSQL, MongoDB, Redis, RabbitMQ administration
- Database HA/failover configurations (PgBouncer, HAProxy)
- Linux/Unix systems and shell scripting (Bash, PowerShell)
- Kubernetes networking, Ingress, and Gateway API
Nice to Have:
- Experience with LlamaIndex, LangChain, or AI/ML platforms
- Vector databases or KDB.AI knowledge
- Temporal.io workflow orchestration
- Python/Go for automation