AI Jobs Map

XpertDirect · Berlin, Germany

AI Inference Platform Engineer

Hybridseniorfull timePosted 5 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

kubernetesvllmpythonterraformobservabilitymlopslinuxcudapytorchprometheusgrafanaopentelemetryllmawsgcpsystem-designray

AI Inference Platform Engineer

Berlin, Germany — Hybrid

AI Infrastructure | Model Serving | GPU Computing | Inference Engineering | ML Platforms

Our client, a growing AI Infrastructure company based in Berlin, is looking for an AI Inference Platform Engineer to build and optimise the platform used to serve production AI models across GPU-enabled infrastructure.

You'll work at the intersection of AI Infrastructure, Distributed Systems, and Platform Engineering, focusing on inference performance, GPU utilisation, autoscaling, latency, and reliability.

What You'll Work On

• Build and operate Kubernetes infrastructure for production AI inference

• Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton

• Improve GPU utilisation, throughput, and inference latency

• Design autoscaling strategies for dynamic AI workloads

• Build platform tooling and automation in Python

• Provision and manage infrastructure using Terraform

• Develop observability across models, GPUs, Kubernetes, and serving infrastructure

• Profile and troubleshoot inference performance bottlenecks

• Improve batching, concurrency, caching, and resource allocation strategies

• Build reliable deployment workflows for new models and model versions

• Partner with ML Engineers to move models efficiently into production

Core Skills

• 4+ years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering, or similar roles

• Kubernetes

• Python

• NVIDIA GPU infrastructure

• vLLM

• NVIDIA Triton Inference Server

• Terraform

• Observability

• Strong understanding of Linux and distributed production systems

Nice to Have

CUDA

NVIDIA GPU Operator

PyTorch

KServe

Ray Serve

Prometheus / Grafana / OpenTelemetry

LLM inference optimisation

Quantisation techniques

Multi-GPU inference

AWS / GCP GPU infrastructure

Experience operating high-throughput or latency-sensitive inference services

More jobs at XpertDirect

  • XpertDirect · Copenhagen, Capital Region of Denmark, Denmark

    4 days ago

    Cloud Security Policy Engineer

    Hybridseniordevsecopsawskubernetesterraform+4LinkedIn
  • XpertDirect · Berlin, Germany

    6 days ago

    ML Observability Engineer

    Hybridseniorobservabilitymlopssrepython+4LinkedIn
  • XpertDirect · Munich, Bavaria, Germany

    6 days ago

    AI Platform Security Engineer

    Hybridseniorkubernetespythonterraformdevsecops+4LinkedIn
  • XpertDirect · Zurich, Switzerland

    6 days ago

    Data Platform SRE

    Hybridseniorsreobservabilitydata-engineeringapache-kafka+4LinkedIn
  • XpertDirect · Vienna, Austria

    7 days ago

    Kubernetes Reliability Engineer

    Hybridseniorkubernetessreobservabilityprometheus+4LinkedIn