AI Jobs Map

BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD. · Outram

SRE Engineer

Posted 21 days ago
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sredevsecopsobservabilityagentic-aigrafanasplunkdatadogci/cddevopslinuxbashgitdockerkubernetesprometheusopentelemetrypythonawsazuregcp

Job Description & Requirements

About Bounteous

Bounteous is a leading global

digital engineering and technology solutions provider

, trusted by top-tier organizations across Banking, Financial Services & Insurance (BFSI). With a combined strength of more than

5,000 engineers

worldwide, we specialize in delivering high-impact solutions across Capital Markets, Core Banking, Payments, Digital Transformation, Data & AI, Cloud Engineering, and modern enterprise platforms.

In Singapore, we partner closely with major regional and global financial institutions, including banks, asset managers, and market infrastructure providers, to drive advanced engineering, modernization, and AI-led transformation. Our teams bring deep expertise across

Site Reliability Engineering (SRE), Cloud-Native Platforms, DevSecOps, Infrastructure Automation, Observability, AI Engineering, and Enterprise Operations

, supported by a strong delivery presence across APAC, India, Europe, and North America.

About the Role

We are seeking an experienced

Site Reliability Engineer (SRE)

to support the implementation, operation, and continuous improvement of an enterprise-grade

Agentic AI Security Control Platform (MSAgentShield)

.

The successful candidate will play a key role in ensuring platform reliability, availability, observability, monitoring, automation, incident management, and operational excellence across a high-availability production environment. This role requires strong expertise in monitoring, alerting, log analytics, production support, and operational engineering practices.

Key Responsibilities

Design, implement, and maintain monitoring, observability, and alerting frameworks for mission-critical production systems.

Develop and optimize dashboards, metrics, logging, and alerting solutions using tools such as Grafana, Splunk, Datadog, or equivalent platforms.

Monitor system health, platform availability, and performance to proactively identify and resolve issues.

Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and operational KPIs.

Participate in production incident management, root cause analysis (RCA), and post-incident reviews.

Support and enhance enterprise log analytics and monitoring solutions.

Develop automation scripts and tooling to reduce operational overhead and improve platform reliability.

Maintain and improve CI/CD pipelines and deployment processes.

Collaborate with engineering, security, infrastructure, and cloud teams to ensure operational readiness.

Create and maintain runbooks, standard operating procedures, troubleshooting guides, and escalation processes.

Support platform upgrades, maintenance activities, disaster recovery testing, and operational resilience initiatives.

Contribute to continuous improvement initiatives focused on reliability, scalability, and service quality.

Required Skills & Experience

Strong experience in

Site Reliability Engineering (SRE), DevOps, Platform Engineering, Infrastructure Engineering, or Production Support Engineering

.

Solid hands-on experience with

Linux administration, shell scripting, Git, and CI/CD pipelines

.

Working knowledge of

Docker and Kubernetes

in production environments.

Strong expertise in

observability, monitoring, dashboarding, and alerting

.

Hands-on experience with

Grafana

and enterprise monitoring platforms.

Experience with

Splunk, Datadog, ELK Stack, Prometheus, OpenTelemetry, or similar observability solutions

.

Experience defining and managing SLOs, SLIs, alert thresholds, and operational metrics.

Strong incident response, troubleshooting, and root cause analysis experience.

Basic to intermediate Python scripting for automation and operational tooling.

Experience supporting high-availability production systems.

Strong analytical, problem-solving, and communication skills.

Experience working in Agile delivery environments.

Preferred Qualifications

Bachelor's Degree in Computer Science, Engineering, Information Technology, or a related discipline.

Experience supporting large-scale cloud-native environments on AWS, Azure, or GCP.

Experience with log analytics and enterprise monitoring ecosystems.

Exposure to AI/ML platforms, agentic AI solutions, or security platforms.

Experience implementing Infrastructure as Code (Terraform, Ansible, etc.).

Experience with ITIL-based operational processes.

Familiarity with security monitoring, SIEM platforms, and DevSecOps practices.

Experience working in global enterprise environments.

What's on Offer

You will join a high-caliber engineering environment where reliability, automation, observability, and operational excellence are critical to success. This role offers the opportunity to work on a next-generation

Agentic AI Security Platform

supporting enterprise-scale production environments.

More jobs at BOUNTEOUSXACCOLITE SINGAPORE PTE. LTD.