AI Jobs Map

XpertDirect · Stockholm, Stockholm County, Sweden

Streaming Reliability Engineer

Hybridseniorfull timePosted yesterday
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sreapache-kafkaapache-flinkkubernetesprometheusterraformjavascalagrafanaopentelemetryawsgcpargocdsystem-designkafka-streams

Streaming Reliability Engineer

Stockholm, Sweden — Hybrid

Mobility Technology | Streaming Infrastructure | Site Reliability Engineering | Real-Time Data | Distributed Systems

Our client, a growing Mobility Technology company based in Stockholm, is looking for a Streaming Reliability Engineer to own the performance, resilience, and operability of high-throughput streaming infrastructure processing real-time operational data.

You'll work at the intersection of Streaming Engineering, SRE, and Platform Engineering, ensuring Kafka and Flink workloads remain fast, observable, and reliable as data volumes and platform complexity grow.

What You'll Work On

• Operate and improve high-throughput Kafka infrastructure

• Build and optimise real-time processing workloads using Apache Flink

• Improve reliability across Kubernetes-based streaming environments

• Define SLIs, SLOs, and reliability standards for critical streaming services

• Build monitoring and alerting using Prometheus

• Investigate latency, throughput, consumer lag, backpressure, and processing failures

• Automate infrastructure provisioning and configuration using Terraform

• Improve partitioning, scaling, and resource-allocation strategies

• Build reliability tooling in Java and/or Scala

• Automate recovery and reduce manual intervention during incidents

• Perform capacity planning for growing event volumes

• Partner with Data and Platform Engineers to design more resilient streaming systems

Core Skills

• 4+ years in Streaming Engineering, SRE, Data Infrastructure, Platform Engineering, or similar roles

• Apache Kafka

• Apache Flink

• Kubernetes

• Java and/or Scala

• Prometheus

• Terraform

• Strong understanding of distributed systems and production reliability

Nice to Have

Kafka Streams / Kafka Connect

Schema Registry / Avro / Protobuf

Grafana / OpenTelemetry

Exactly-once processing concepts

Event-driven architectures

AWS / GCP

GitOps / Argo CD

JVM performance tuning

Incident management and postmortems

High-throughput or low-latency systems

Experience operating multi-region streaming infrastructure

More jobs at XpertDirect