AI Jobs Map

Thrive IT Systems · Cracow, Małopolskie, Poland

Site Reliability Engineer

Hybridseniorfull timePosted today
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sredevopssystem-designobservabilityansiblejenkinsprometheusgrafanajavapythonnode.jssql

Role: Site Reliability Engineer

Location: Poland, Krakow

Mode: Hybrid - 3 day a week

Description

- Serve as a Site Reliability Engineer within a global DevOps team supporting highly available 24x7 production services

- Implement solutions using SRE best practices to improve service availability performance security and transparency

- Resolve incidents conduct root cause analysis and facilitate postincident reviews

- Participating in software architecture design

- Strong experience with the Software Development Life Cycle SDLC including requirements gathering design development testing deployment and maintenance

- Define application SLIs and SLOs build maintain observability to continuously monitor the operational performance create execute action plans to address failures to meet the desired metrics

- Plan and execute application infrastructure migration disaster recovery exercise and product upgrade

- Enhance automation and develop self-service capability to improve user experience and reduce manual effort

- Provide oncall support as part of a rotation to ensure rapid response to critical incidents

- Participate in scheduled maintenance activities including those that may occur during weekends to ensure system reliability and minimal disruption to users

Key Skills and Qualifications for this role

- Hands on years of professional experience in Production Application Support or Site Reliability Engineering demonstrating strong troubleshooting resolution and issue prevention skills in high-pressure environments

- Proficiency with automation build and monitoring tools such as Ansible Jenkins Prometheus and Grafana

- Strong analytical and troubleshooting skills

- Strong engineering skills with some of the following Java Python NodeJS plus SQL

- Knowledge to the Software Development Life Cycle SDLC and practice its principles

- Good communication skills with the ability to collaborate effectively with globally dispersed cross functional teams and vendors

- Prior experience supporting large Atlassian Jira and Confluence Data Centre instances is preferred but not required Candidates who demonstrate a strong ability to learn quickly and adapt to new technologies will also be considered

- Solid knowledge on observability and monitoring tools such as Grafana and Prometheus is an advantage

More jobs at Thrive IT Systems

Site Reliability Engineer at Thrive IT Systems (Cracow, Małopolskie, Poland) | AI Jobs Map