AI Jobs Map

XpertDirect · Berlin, Germany

ML Observability Engineer

Hybridseniorfull timePosted 6 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

observabilitymlopssrepythonopentelemetryprometheuskubernetesmlflowgrafanapytorchtensorflowawsgcpterraformllmmachine-learning

ML Observability Engineer

Berlin, Germany — Hybrid

AI SaaS | ML Observability | MLOps | ML Reliability | AI Infrastructure

Our client, a growing AI SaaS company based in Berlin, is looking for an ML Observability Engineer to build the systems that give engineering teams visibility into model performance, infrastructure behaviour, latency, drift, and production failures.

You'll work at the intersection of MLOps, Site Reliability Engineering, and ML Engineering, helping teams understand not just whether their services are running—but whether their models are performing as expected.

What You'll Work On

• Build observability systems for production ML workloads

• Monitor model performance, drift, latency, errors, and resource utilisation

• Develop ML monitoring and reliability tooling in Python

• Build telemetry pipelines using OpenTelemetry

• Create metrics, dashboards, and alerting with Prometheus

• Integrate observability across Kubernetes-based ML infrastructure

• Track experiments, deployments, and model versions using MLflow

• Define meaningful reliability indicators for production ML systems

• Detect degradation and anomalous model behaviour before it impacts users

• Improve incident investigation and root-cause analysis across models and infrastructure

• Collaborate with ML, MLOps, and Platform Engineers to improve production reliability

Core Skills

• 4+ years in MLOps, ML Infrastructure, SRE, Platform Engineering, or similar roles

• Python

• MLflow

• OpenTelemetry

• Kubernetes

• Prometheus

• Model monitoring

• Strong understanding of production ML systems and observability

Nice to Have

Grafana

Evidently / Arize / WhyLabs or similar ML observability tooling

Data and concept drift detection

Distributed tracing

SLI / SLO design

PyTorch / TensorFlow

Feature and data-quality monitoring

AWS / GCP

Terraform

Incident management and postmortems

LLM observability and evaluation

More jobs at XpertDirect