AI Jobs Map

PRI Technology · New York, NY

Sr Cloud AI Platform Engineer

HybridseniorcontractPosted yesterday
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

awsterraformekspythonobservabilitykubernetesopenaianthropicgeminillmopentelemetryprometheusgrafanadatadogsystem-design

Sr Cloud AI Platform Engineer

Contract

Hybrid in NYC

Must be local.

This is not a Full Stack Developer role. We'd like candidates to have strong AWS infra experience (Terraform, VPC, EKS etc) with Python/Go experience

As a Senior Engineer, you will help build and enhance the reliability, scalability, and resilience of enterprise AI platforms. Working closely with engineering teams, you will develop solutions that minimize the impact of provider outages and system dependencies, establish operational standards through observability and service-level objectives (SLOs), and ensure the platform continues to evolve alongside advancements in AI technologies.

You will design and maintain Kubernetes-based platform services, AI gateways, APIs, and self-service capabilities that enable teams to develop and deploy AI applications efficiently. The role includes integrating with multiple AI providers, including OpenAI, Anthropic, Gemini, and AWS Bedrock, while implementing features such as intelligent request routing, failover strategies, retry logic, rate limiting, and cost management.

You will also strengthen the platform's operational health by developing monitoring and observability capabilities using metrics, logs, distributed tracing, dashboards, and alerting. Additionally, you will design and implement networking solutions that support secure, reliable communication between applications running across both cloud and on-premises environments.

SKILLS:

- 6+ years of experience in software engineering, with a focus on backend or platform development.

- Proficiency in Python/Go for building scalable backend applications and APIs.

- Experience working with distributed systems and cloud-based architectures, with an understanding of reliability, scalability, and fault tolerance.

- Working knowledge of AWS services.

- Experience using Infrastructure as Code tools, with Terraform preferred.

- Familiarity with API/LLM gateways

- Experience with observability and monitoring tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or equivalent solutions.

- Experience deploying and managing containerized applications using Kubernetes, with exposure to Amazon EKS, Karpenter, or similar autoscaling technologies preferred.

More jobs at PRI Technology

Similar roles in New York