AI Inference Platform Engineer
Berlin, Germany — Hybrid
AI Infrastructure | Model Serving | GPU Computing | Inference Engineering | ML Platforms
Our client, a growing AI Infrastructure company based in Berlin, is looking for an AI Inference Platform Engineer to build and optimise the platform used to serve production AI models across GPU-enabled infrastructure.
You'll work at the intersection of AI Infrastructure, Distributed Systems, and Platform Engineering, focusing on inference performance, GPU utilisation, autoscaling, latency, and reliability.
What You'll Work On
• Build and operate Kubernetes infrastructure for production AI inference
• Deploy and optimise model-serving workloads using vLLM and NVIDIA Triton
• Improve GPU utilisation, throughput, and inference latency
• Design autoscaling strategies for dynamic AI workloads
• Build platform tooling and automation in Python
• Provision and manage infrastructure using Terraform
• Develop observability across models, GPUs, Kubernetes, and serving infrastructure
• Profile and troubleshoot inference performance bottlenecks
• Improve batching, concurrency, caching, and resource allocation strategies
• Build reliable deployment workflows for new models and model versions
• Partner with ML Engineers to move models efficiently into production
Core Skills
• 4+ years in AI Infrastructure, ML Infrastructure, MLOps, Platform Engineering, or similar roles
• Kubernetes
• Python
• NVIDIA GPU infrastructure
• vLLM
• NVIDIA Triton Inference Server
• Terraform
• Observability
• Strong understanding of Linux and distributed production systems
Nice to Have
CUDA
NVIDIA GPU Operator
PyTorch
KServe
Ray Serve
Prometheus / Grafana / OpenTelemetry
LLM inference optimisation
Quantisation techniques
Multi-GPU inference
AWS / GCP GPU infrastructure
Experience operating high-throughput or latency-sensitive inference services