Job Description
We are seeking an experienced
Platform Engineer / Site Reliability Engineer (SRE)
to support the deployment, operations, and reliability of enterprise platform services. The ideal candidate will have strong hands-on experience with
Kubernetes
,
observability platforms
,
CI/CD pipelines
, and
production support
in large-scale environments.
Key Responsibilities
Manage and support Kubernetes-based applications and platform services in production environments.
Troubleshoot and resolve issues related to pods, deployments, services, ingress, scaling, and platform performance.
Design, implement, and maintain monitoring, alerting, and observability solutions using Grafana, Prometheus, Splunk, or similar tools.
Support CI/CD pipelines and deployment processes across development, testing, and production environments.
Perform production incident management, root cause analysis (RCA), and service reliability improvements.
Work closely with development, infrastructure, and security teams to ensure platform stability and availability.
Automate operational tasks using scripting and infrastructure-as-code tools where applicable.
Contribute to system capacity planning, performance optimization, and operational readiness.
Required Skills & Experience
5+ years of experience in Platform Engineering, SRE, DevOps, Infrastructure Engineering, or related roles.
Strong hands-on experience with
Kubernetes
and containerized workloads.
Experience with
Helm
, Kubernetes deployments, services, ingress, scaling, and troubleshooting.
Good experience with
Grafana
,
Prometheus
,
Datadog
,
Splunk
, or equivalent monitoring and observability tools.
Experience working with
CI/CD tools
such as Jenkins, GitLab CI/CD, GitHub Actions, ArgoCD, or similar.
Strong Linux administration and troubleshooting skills.
Experience with production support, incident management, and RCA.
Exposure to cloud platforms such as AWS, Azure, or GCP.
Working knowledge of scripting and automation using Python, Shell, or similar tools.
Preferred Skills
Experience in banking, financial services, or other large-scale enterprise environments.
Exposure to Infrastructure as Code tools such as Terraform or Ansible.
Knowledge of SRE concepts including SLI, SLO, SLA, and Error Budgets.
Experience with high-availability systems, observability platforms, and distributed systems.
Employment Type
Full-time / Contract (as applicable)
Location
Singapore
Experience
5-12 years preferred
This version is concise enough for MCF posting while still attracting the exact profile type that the client has been responding positively to:
Kubernetes + SRE + Monitoring + Production Support engineers rather than pure software developers.