Role and Responsibilities
- Design, develop, and implement deployment pipelines for Data Science, ML, AI, and GenAI solutions on AWS cloud.
- Build and maintain CI/CD and CT pipelines using GitHub Actions, Airflow, or similar orchestration tools.
- Support deployment and lifecycle management of ML models, LLM-based applications, prompt workflows, embeddings, vector search, and API-based AI services.
- Collaborate with data scientists, GenAI engineers, and data engineers to understand technical requirements, solution design, and deployment processes; document standards and operating procedures.
- Continuously monitor and maintain ML and GenAI pipelines in production, ensuring performance, reliability, latency, cost efficiency, and model quality.
- Implement observability for AI/GenAI workloads, including model drift, data quality, prompt quality, hallucination indicators, latency, throughput, and cost metrics.
- Optimize pipelines and runtime environments for scalability, security, automation, and cost-effectiveness.
- Troubleshoot and resolve issues related to deployments, integrations, production incidents, and model/service performance.
- Ensure compliance with security, privacy, responsible AI, and data governance standards across all deployment activities.
- Support experimentation and release processes for model versions, prompt versions, feature pipelines, and evaluation workflows.
- Keep up to date with emerging tools, best practices, and trends in AI Ops, MLOps, and LLMOps.
- Provide support, guidance, and knowledge sharing to other team members on deployment, automation, monitoring, and operational best practices.
Education and Competencies
- Bachelor’s or master’s degree in computer science, engineering, informatics, data science, or equivalent qualification.
- Strong hands-on experience in **Data Science, AI, and GenAI operations** with AWS cloud platform services such as ECS, SageMaker, Batch, Lambda, API Gateway, S3, Redshift, CloudWatch, and related managed services.
- Experience with GenAI ecosystem components such as foundation models, prompt orchestration, retrieval-augmented generation, vector databases, model gateways, and evaluation/monitoring frameworks.
- Proficiency in Python and PySpark, with hands-on experience in containers, Airflow, GitHub Actions, SonarQube, and related automation/tooling stacks.
- Strong understanding of CI/CD, deployment automation, infrastructure as code, and production monitoring tools such as Datadog or equivalent observability platforms.
- Experience with MLOps, LLMOps, and AI lifecycle management in enterprise environments.
- Ability to understand architectural and technical dependencies in customer analytics and AI environments.
- Strong ability to translate business and technical requirements into scalable technical implementations.
- Confident communicator who can present effectively internally and with clients.
- Experience working in Agile delivery models.
- Strong team player who can coordinate effectively across distributed, global teams and time zones.