AI Jobs Map

FPT France · Toulouse, Occitanie, France

Senior Site Reliability Engineer

Hybridseniorfull timePosted 8 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sreawsreactnode.jsnginxspringecss3observabilityeksjavaopentelemetrylinuxbashelasticsearch

1. Role summary

You own end-to-end visibility of a distributed microservice platform, and you are the person who determines why a request is slow — instrumenting services, repairing broken traces, reading thread dumps, and turning production symptoms into root causes the customer's engineering teams can act on. You lead the team technically and you are the customer's first technical interlocutor. The engagement is measured on time-to-first-usable-trace, and you set the specification that everything else is built against.

2. Engagement context

The customer operates a distributed application on AWS: React front-ends, a Node.js authentication proxy, Nginx, and Spring Boot / Tomcat services running on Amazon ECS Fargate, backed by Amazon RDS and Amazon S3. Its current tooling cannot explain the production incidents it experiences, and centralizing observability on a self-managed Elastic Enterprise Stack deployed on Amazon EKS is the response to that.

3. Key responsibilities

Specification before construction

- Produce the instrumentation gap analysis: establish precisely what the customer's current tooling cannot show and why, as the written justification of the programme and as leverage for the customer inside their own organization.

- Produce the measurement specification — what must be measured at each layer and the fields required to measure it — which is the input the platform's data model is built against. Nothing is ingested before this exists.

- Open the customer-side dependencies that block the first diagnosis: the Nginx log format (request time and upstream response time captured separately), the Tomcat MBean registry exposure required for thread pool metrics, and the developer availability needed to instrument the pilot service.

- Select the pilot service with the customer on the criterion of diagnostic difficulty, not ease of instrumentation.

Distributed tracing and instrumentation

- Instrument Java (Spring Boot / Tomcat) and Node.js services using Elastic APM agents or the Elastic Distributions of OpenTelemetry (EDOT), ensuring uninterrupted end-to-end trace continuity.

- Diagnose and repair trace context propagation failures across Nginx and the Node.js authentication proxy, including manual instrumentation wherever auto-instrumentation falls short.

- Define and operate the sampling strategy — head-based versus tail-based — balancing diagnostic value against ingestion cost.

- (Optional, if possible) Implement browser-side Real User Monitoring on the React front-ends, in line with the customer's privacy and consent requirements.

Performance diagnosis

- Investigate latency, saturation and error incidents end to end: JVM and Tomcat thread pools, JDBC connection pools, garbage collection behavior, thread dumps and heap dumps.

- Diagnose operating-system and network constraints — file descriptor limits, ephemeral port exhaustion, socket timeouts, keep-alive and connection reuse at the proxy layer.

- Analyze downstream managed-service behavior (Amazon RDS, Amazon S3) and correlate it with application-level symptoms.

- Separate symptom from cause and produce written root-cause analyses that the customer's engineering teams can act on without further interpretation.

Alerting and observability by design

- Build Kibana dashboards that expose saturation and queueing rather than averages: percentile latency, thread and connection pool utilization, error rates and error budgets.

- Define alerting thresholds across APM metrics, 5xx rates and host-level resource utilization, together with the routing and escalation rules that send each alert to the team able to act on it.

- Treat alert noise reduction as an explicit, measured objective rather than a by-product.

Enablement and knowledge transfer

- Work directly with the customer's senior developers on instrumentation standards, trace context propagation and the interpretation of trace data.

- Run the team's technical decisions and act as its point of arbitration on observability design.

- Maintain runbooks as a condition of completion.

4. Required qualifications

- 8+ years in software engineering and production operations, of which at least 4 in a production-facing SRE, production engineering or performance engineering role.

- Demonstrable JVM expertise: you have used thread dumps and heap dumps to resolve real production incidents, and you can explain the difference between thread pool saturation and downstream latency without hesitation.

- Hands-on distributed tracing with Elastic APM and/or OpenTelemetry in Java and Node.js environments, including manual trace context propagation.

- Strong Linux fundamentals: shell scripting, and a working understanding of the network stack, file descriptors and kernel-level connection limits.

- Practical experience configuring and monitoring reverse proxies (Nginx, HAProxy): timeouts, connection limits, keep-alive behaviour.

- Working knowledge of AWS ECS Fargate, CloudWatch and CloudTrail.

- Professional working proficiency in English, spoken and written.

5. Desirable

- Elasticsearch and Kibana experience as an operator, not only as a consumer of dashboards.

- AWS FireLens / Fluent Bit log routing on ECS.

- Experience in aerospace, defense, or another regulated industry with export-control constraints.

- French proficiency

6. Working arrangements and constraints

- Based on the customer's site in Toulouse, hybrid; the minimum on-site presence is to be confirmed with the customer.

- Business-hours engagement.

More jobs at FPT France

  • FPT France · Lyon, Auvergne-Rhône-Alpes, France

    today

    Senior AI Engineer

    Senior$54,721llmragopenaianthropic+4LinkedIn