- Implement and operate CI/CD pipelines, automated testing and release processes for AI/ML workloads.
- Build and maintain model registry, model serving and AI gateway integrations for LLM APIs and internal applications.
- Configure and maintain observability for model usage, cost, token consumption, latency, reliability and quality signals using tools such as Prometheus, Grafana, logging and alerting platforms.
- Support the transition of workloads from sandbox or PoC environments into production by following defined standards, runbooks and support models.
- Implement reusable technical components for LLM API integration, RAG pipelines, evaluation pipelines and integration with business applications.
- Execute infrastructure-as-code for platform environments across container and cloud infrastructure, including Docker, Kubernetes and Helm-based deployment patterns.
- Maintain runbooks, operating procedures, technical documentation and operational dashboards for platform components.
- Support incident analysis, reliability improvements, cost optimization and lifecycle maintenance for production AI workloads.
- Work with nearshore, system integration or cloud partners on specific implementation tasks as directed by the AI Platform Engineer.
- Collaborate with data engineering, application development, cloud platform and security teams on integration, identity, access and deployment requirements.
Requirements:
- 3-5 years of experience in DevOps, cloud engineering, ML engineering, MLOps or platform engineering.
- Hands-on experience with CI/CD, infrastructure as code and automated deployment in production environments.
- Strong practical Python skills and Git-based development workflows.
- Experience with Docker and Kubernetes; deployment tooling such as Helm is desirable.
- Working experience with Azure or AWS cloud services, including compute, storage and IAM concepts.
- Experience with observability tooling such as Prometheus, Grafana, logging platforms and alerting practices.
- Familiarity with MLOps concepts such as model registries, evaluation pipelines, drift monitoring and model lifecycle management.
- Understanding of network isolation, identity, secrets management and API access control.
- Fluency in English, spoken and written.