The Opportunity
We are seeking an accomplished and highly skilled Senior DevOps Engineer to join our growing engineering team.
In this pivotal role, you will help design, implement, and operate the cloud infrastructure, automation tooling, CI/CD pipelines, and platform services that support our External Marketplace platform and API ecosystem. You will work closely with engineering teams, architects, security specialists, and product stakeholders to build reliable and secure cloud-native solutions.
You will play a key role in establishing infrastructure standards, driving automation initiatives, improving operational resilience, and enabling engineering teams to deliver software quickly and safely. This is an excellent opportunity for an experienced DevOps professional who thrives in complex enterprise environments and enjoys solving large-scale engineering challenges.
What You'll Do
Infrastructure Automation & Platform Engineering
- Design, implement, and manage scalable, secure, and highly available cloud infrastructure primarily on Google Cloud Platform (GCP).
- Develop and maintain Infrastructure as Code solutions using Terraform and reusable infrastructure modules.
- Build self-service platform capabilities that enable engineering teams to provision and manage infrastructure efficiently.
- Implement infrastructure standards, governance controls, and automation frameworks across multiple environments.
CI/CD & Release Engineering
Design, build, and optimize enterprise-grade CI/CD pipelines using tools such as:
- GitHub Actions
- Jenkins
- GitLab CI/CD
- Harness
Drive deployment automation, release management, environment provisioning, and configuration management practices.
Improve deployment reliability through progressive delivery approaches, automated testing, and deployment guardrails.
Champion GitOps and Infrastructure-as-Code practices across the engineering landscape.
Cloud Architecture & Operations
Provide technical leadership and guidance on cloud architecture, best practices, and operational excellence.
Support cloud-native services including:
- GKE (Google Kubernetes Engine)
- Cloud Run
- Cloud Functions
- Pub/Sub
- Cloud SQL
- Apigee
- VPC Networking
- IAM
- Cloud Monitoring
- Design resilient and observable systems aligned to enterprise security and reliability standards.
- API Platform & Marketplace Enablement
- Build and maintain infrastructure supporting enterprise API platforms and external marketplace services.
- Enable secure onboarding of internal and external API consumers.
- Establish standardized deployment and operational processes for API products and services.
- Support API gateway technologies such as Apigee and enterprise integration platforms.
- Drive incident management, root cause analysis, and post-incident reviews.
- Improve platform availability, scalability, and performance through automation and engineering improvements.
- Monitoring & Observability
- Implement enterprise monitoring, alerting, logging, and observability solutions.
Utilize tools including:
- Prometheus
- Grafana
- Google Cloud Monitoring
- ELK/Elastic Stack
- OpenTelemetry
- Proactively identify issues before they impact customers and engineering teams.
- Develop dashboards and operational reporting for platform health and reliability.
- DevSecOps & Security
- Embed security best practices throughout the software delivery lifecycle.
- Integrate automated vulnerability scanning, code analysis, policy enforcement, and compliance controls into CI/CD pipelines.
- Partner with Cyber Security teams to ensure compliance with internal and regulatory requirements.
- Support implementation of secure identity and access management controls across cloud platforms.
Operational Excellence:
- Lead resolution of complex production incidents and service disruptions.
- Continuously improve operational procedures, runbooks, and platform support processes.
- Drive automation initiatives that reduce manual effort and improve service reliability.
- Identify opportunities for cloud cost optimization while maintaining performance and resilience.
- Mentoring & Leadership
- Mentor junior engineers and promote DevOps best practices across the engineering community.
- Collaborate with architects, developers, testers, platform teams, and product owners.
- Drive engineering excellence through knowledge sharing, technical leadership, and continuous learning.
What You'll Need
- Essential Skills & Experience
- DevOps & Cloud Engineering
- Proven experience (10+ years) in DevOps, Platform Engineering, Site Reliability Engineering, or Cloud Infrastructure roles.
- Significant experience operating highly available production systems at enterprise scale.
- GCP Expertise (Essential)
Strong hands-on experience with:
- GKE (Google Kubernetes Engine)
- Cloud Run
- Cloud Functions
- Pub/Sub
- Cloud SQL
- VPC Networking
- IAM
- Cloud Monitoring
- Secret Manager
- Cloud Storage
- Infrastructure as Code (IaC)
- Expert-level knowledge of Terraform.
- Experience building reusable infrastructure modules and automation frameworks.
- Experience with policy-as-code and infrastructure governance.
Kubernetes & Containerization
- Strong experience with:
- Kubernetes
- Docker
- Container security
- Service mesh technologies
- Experience operating production Kubernetes environments.
- CI/CD & Release Management
Experience with:
- GitHub Actions
- Jenkins
- GitLab CI/CD
- Harness
- ArgoCD (desirable)
Strong understanding of:
- GitOps
- Deployment strategies
- Release automation
- Configuration management
- API & Integration Technologies
Experience working with API Gateway technologies such as:
- Apigee
- Azure API Management
- Kong (desirable)
- Strong understanding of:
- REST APIs
- OpenAPI Specifications
- OAuth2
- JWT
- API Security Standards
- Observability & Monitoring
Experience with:
- Prometheus
- Grafana
- ELK Stack
- OpenTelemetry
- Google Cloud Monitoring
- Scripting & Automation
Proficiency in one or more of:
- Python [optional]
- Bash
- Go
- PowerShell
- Networking & Security
You'll thrive in this role if you:
- Have a passion for automation and engineering excellence.
- Are comfortable working in complex enterprise-scale environments.
- Enjoy troubleshooting challenging technical problems.
- Focus on continuous improvement and operational resilience.
- Can communicate effectively with both technical and non-technical stakeholders.
- Take ownership of outcomes and drive delivery through collaboration and influence.
Why Join Us?
- Be part of a strategic platform powering our Customer’s digital transformation.
- Work on cutting-edge cloud-native technologies and large-scale API ecosystems.
- Influence the future direction of our DevOps and Platform Engineering practices.
- Collaborate with talented engineers, architects, and technology leaders.
- Access industry-leading learning, certification, and development opportunities.
- Enjoy a comprehensive benefits package and strong commitment to work-life balance.
- Make a real impact on services used by millions of customers across the UK.