AI Jobs Map

Keysight Technologies · Gurugram, Haryana

Senior Web Operations Engineer (SRE)

Remoteseniorfull timePosted today
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

sreincident-responseobservabilityprometheusgrafanaopentelemetrydatadogmicroservicesanomaly-detectionllmtime-seriesvulnerability-managementci/cdawsterraformpulumikubernetescloudflarejavajenkins

Overview:

Keysight is at the forefront of technology innovation, delivering breakthroughs and trusted insights in electronic design, simulation, prototyping, test, manufacturing, and optimization. Our ~16,800 employees create world-class solutions in communications, 5G, automotive, energy, quantum, aerospace, defense, and semiconductor markets for customers in over 100 countries. Learn more about what we do.

Our award-winning culture embraces a bold vision of where technology can take us and a passion for tackling challenging problems with industry-first solutions. We believe that when people feel a sense of belonging, they can be more creative, innovative, and thrive at all points in their careers.

Responsibilities:

Key Responsibilities

Reliability & Availability

-
Lead incident response, conduct RCAs and ensure action items are tracked to closure

-
Build and maintain runbooks, playbooks and escalation frameworks for proactive and reactive response

-
Drive toil reduction by identifying repetitive operational work and engineering it away

Observability

-
Design and own the full observability stack — metrics, logs, traces and events — using tools like Prometheus, Grafana, OpenTelemetry, Datadog, Dynatrace or similar

-
Build intelligent alerting that reduces noise, eliminates alert fatigue and surfaces actionable signals

-
Implement distributed tracing and dependency mapping to provide end-to-end visibility across microservices

-
Drive adoption of continuous profiling and real user monitoring (RUM) for proactive performance management

AI Adoption in SRE

-
Leverage AIOps platforms to enable anomaly detection, predictive alerting and automated root cause analysis

-
Implement AI-assisted incident triage — using LLM-powered tools to summarise incidents, suggest fixes and accelerate MTTR

-
Build and maintain ML-powered capacity forecasting models to optimise infrastructure spend and prevent resource saturation

Security & Vulnerability Management

-
Embed security-as-reliability principles — treating security incidents with the same urgency as availability incidents

-
Own issue remediations across infrastructure (OS, containers, dependencies)

-
Integrate SAST, DAST and SCA tools into CI/CD pipelines to shift security left

Infrastructure & Platform Engineering

-
Design, build and maintain cloud-native infrastructure on AWS using Infrastructure as Code (Terraform, Pulumi) & drive rightsizing, reserved capacity planning and cost anomaly detection

-
Own Kubernetes cluster operations — autoscaling, resource management, networking and upgrade strategy

Leadership & Culture

-
Mentor and guide junior and mid-level SREs — conducting technical reviews and pair debugging sessions

-
Define and evolve SRE team standards, best practices and engineering principles

-
Collaborate closely with product, development and security teams as an embedded reliability partner

-
Contribute to on-call rotation and drive continuous improvement of on-call experience

-
Represent SRE in architecture reviews, sprint planning and cross-functional forums

The Competitive Edge

AEM Administration

-
Own end-to-end reliability and availability of AEM environments — Author, Publish, Dispatcher and AEM as a Cloud Service (AEMaaCS) — across dev, staging and production

-
Monitor and manage AEM instance health, optimise Dispatchers, Manage DAM, OSGi Configurations, replication queues.

Exposure to CDN

-
Experience with Cloudflare - CDN, Workers,

Qualifications:

B.E/B.Tech/M.Tech/MCA in Computers, IT, EC, telecommunications or other related streams with minimum 7+ years of working experience with following skills:

-
Strong understanding of cloud-based web architecture and enterprise-level systems, including Apache configuration, rewrite rules, redirects, regular expressions, reverse proxies, restrictive forward proxies, and Apache Mod Proxy configuration, tuning, and optimization.

- Hands-on experience with SSL certificate configuration, TLS, cipher suites, security hardening, and best practices.

- Experience with domain management, network configuration, DNS routing, subnets, zone management, and troubleshooting.

- Experience in vulnerability management and remediation, including the ability to analyze penetration test and vulnerability assessment reports and implement appropriate mitigation strategies.

- Experience with Global Traffic Management (GTM), Content Delivery Networks (CDNs), and network services such as Akamai and Cloudflare.

- Experience managing technology vendors, software licenses, support contracts, and third-party service providers.

- 5–7 years of experience working with Java/J2EE frameworks.

- Experience with Continuous Integration and Continuous Deployment (CI/CD) tools such as Jenkins.

- Experience with cloud-native applications and services, including AWS, Lambda, RDS, and Microsoft Azure.

- Experience with Adobe Experience Manager (AEM) or other enterprise content management systems; AEM experience is preferred.

- Good understanding of designing and developing RESTful APIs and SOAP-based web services.

- Experience with observability and monitoring tools such as New Relic, Grafana, and Datadog; knowledge of Site Reliability Engineering (SRE) concepts is a plus.

- Experience in scripting and automation using Python, Shell scripting, or similar technologies to improve operational efficiency and reduce manual effort.

Careers Privacy Statement***Keysight is an Equal Opportunity Employer.***

More jobs at Keysight Technologies