Sapiens is on the lookout for a Lead Observability Engineer/ APM Engineer to become a key player in our Bangalore team, as we continue to scale observability across a large, multi-customer managed service environment. If you're an expert in Datadog APM who thrives inside traces, service maps, and flame graphs — and is equally comfortable in design forums shaping enterprise monitoring strategy — this role could be the perfect fit, offering high autonomy, deep technical ownership, and a seat at the table in defining how every application we support gets observed, understood, and kept healthy.
About Us
Sapiens International Corporation N.V. is a global leader in intelligent, SaaS-based software solutions. With Sapiens’ robust platform, customer-driven partnerships, and rich ecosystem, insurers are empowered to future-proof their organizations with operational excellence in a rapidly changing marketplace.
Our solutions help insurers harness the power of AI and advanced automation to support core solutions for property and casualty, workers’ compensation, and life insurance, including reinsurance, financial & compliance, data & analytics, digital, and decision management.
Sapiens boasts a longtime global presence, serving over 600 customers in more than 30 countries with our innovative offerings. Recognized by industry experts and selected for the Microsoft Top 100 Partner program, Sapiens is committed to partnering with our customers for their entire transformation journey and is continuously innovating to ensure their success.
What You’ll Do
We are looking for an expert Datadog observability engineer whose primary depth is in Application Performance Monitoring. You will own the APM practice across a large, multi-customer managed service environment — defining the architecture, instrumenting applications, leading deep-dive performance investigations, and setting the standards that every application follows when it is onboarded.
This is a hands-on technical role. You will spend your time inside traces, service maps and flame graphs, and equally in design forums and steering committees where APM strategy, coverage and cost are decided. Alongside expert APM skills, we expect strong all-round Datadog observability capability spanning infrastructure, logs, database monitoring, network, Kubernetes, RUM, synthetics, ServiceNow integration and incident management.
Purpose of the Role
- Improve service reliability and customer experience across all supported applications.
- Reduce mean time to detect and mean time to resolve through trace-led root cause analysis.
- Raise alert quality by replacing symptom-based alerting with dependency-aware, service-level detection.
- Establish enterprise APM standards that make onboarding repeatable, governed and cost-aware.
Key Responsibilities — Datadog APM APM Architecture and Implementation
- Design and implement Datadog APM across enterprise applications and microservices.
- Lead APM onboarding for new and existing applications, end to end.
- Define the APM architecture, instrumentation standards and deployment patterns.
- Configure Datadog Agents and APM components across all application environments.
Distributed Tracing
- Implement and troubleshoot distributed tracing across multi-tier applications.
- Trace end-to-end transactions across microservices and APIs, including asynchronous paths.
- Pinpoint performance bottlenecks along the request path.
- Analyse trace latency, errors, throughput and resource utilisation.
- Use trace analytics to detect performance degradation and transaction bottlenecks early.
Application Performance Troubleshooting
Perform deep-dive analysis of application performance issues using Datadog APM, covering:
- Response time, latency and throughput.
- Error rates, error tracking and request volume.
- Service dependencies, database calls and external API calls.
- Resource consumption at process, container and host level.
- Correlate APM traces with infrastructure metrics and logs.
- Support complex production incidents and major incident investigations as a senior escalation point.
APM Instrumentation
Hands-on experience instrumenting applications with Datadog APM across enterprise technologies, including:
- Java
- .NET and C#
- Node.js
- Python
- Other enterprise application technologies and runtimes
Practical experience with common frameworks, application servers and microservices architectures is highly desirable, as is familiarity with automatic versus manual instrumentation, custom spans and library-level configuration.
APM Configuration and Optimisation
- Configure appropriate trace collection and sampling strategies per service and environment.
- Configure error tracking and analytics.
- Optimise APM telemetry volume and cost by tuning sampling and removing unnecessary trace collection, without losing diagnostic value.
Database and APM Correlation
- Correlate application traces with database activity using Datadog Database Monitoring.
- Identify slow SQL and database calls that contribute to application latency.
- Analyse database query performance from the application perspective.
- Troubleshoot connection pool exhaustion and database dependency issues.
- Work closely with DBA teams across SQL Server, Oracle, PostgreSQL and other database platforms.
Service Mapping and Dependencies
- Build application dependency maps using Datadog APM and Universal Service Monitoring.
- Identify upstream and downstream dependencies and critical transaction paths.
- Establish and maintain service ownership through the Datadog Service Catalogue.
- Support service-level observability and application health monitoring.
APM on Kubernetes and Cloud Platforms
- Implement and support APM for applications running on Azure, AWS and Kubernetes / AKS.
- Correlate Kubernetes and container metrics with application traces.
- Isolate whether a performance issue originates in the application, the container, the Kubernetes node, the database, the network or an external dependency.
APM and ServiceNow Integration
- Integrate Datadog APM events with ServiceNow ITOM and Event Management.
- Define how APM events are converted into actionable events and incidents.
- Ensure appropriate event aggregation, deduplication and suppression during maintenance windows.
- Support CI and service mapping between Datadog and the ServiceNow CMDB.
- Ensure APM-driven incidents are routed to the correct resolver groups.
APM Dashboards and Service Level Objectives
- Develop application-level dashboards covering availability, latency, error rate, throughput, dependency health and database performance, that answer: is the service healthy, what changed, what is the impact, where is the bottleneck, and what should the operator do next.
- Define application health indicators and establish SLIs, SLOs and error budgets for critical services.
APM Governance and Standards
- Establish enterprise APM standards and onboarding templates.
- Define tagging, naming and ownership conventions across the estate.
- Define monitoring requirements based on application architecture and business criticality.
- Review APM coverage regularly and identify gaps.
- Define and drive the APM maturity roadmap.
Broader Datadog Observability Responsibilities
Alongside APM depth, the role requires solid breadth across the Datadog platform and the wider observability discipline.
- Platform breadth — infrastructure monitoring, log management, database monitoring, network performance monitoring, Kubernetes monitoring, Universal Service Monitoring, RUM, synthetic tests, continuous profiler, error tracking and Service Catalogue.
- Monitoring design — monitor services rather than individual resources, covering availability, performance, capacity, errors, security, dependencies, user experience and recovery. CPU, memory and disk alone are not monitoring.
- Alert engineering — actionable alerts only, dynamic thresholds where appropriate, composite and dependency-aware monitors, correlation before escalation, deduplication, maintenance suppression and auto-remediation where feasible.
- Incident correlation — topology-aware event correlation, alert grouping, noise reduction and major incident detection; consolidate multiple symptom alerts into a single actionable incident.
- Automation — Datadog Workflow Automation, incident enrichment, runbook automation, alert routing and ticket automation.
- Integrations — Azure Monitor, SolarWinds, ServiceNow, Prometheus, Grafana, Azure Event Hub, Event Grid, Microsoft Sentinel, Logic Apps and Azure Functions, with Datadog as the central observability platform.
- Governance — standard monitoring templates, dashboard standards, ownership models, monitoring lifecycle and consistent customer onboarding across a multi-customer estate.
Essential Skills and Experience
- Extensive hands-on Datadog experience gained in IT operations, application support or platform engineering.
- Expert-level Datadog APM capability, including instrumentation, distributed tracing, sampling strategies and trace analytics.
- Demonstrable experience instrumenting production applications in at least three of Java, .NET / C#, Node.js and Python.
- Strong working knowledge of microservices, containers, Kubernetes / AKS and cloud-native application architectures.
- Practical experience correlating application traces with database performance across SQL Server, Oracle or PostgreSQL.
- Proven track record of leading root cause analysis on complex production and major incidents.
- Experience defining SLIs, SLOs and error budgets for business-critical services.
- Experience integrating monitoring platforms with ServiceNow ITOM / Event Management.
- Experience working in a large enterprise or managed service environment supporting multiple customers.
- Sound understanding of observability cost management and telemetry optimisation.
Desirable Skills and Experience
- Experience with Datadog Cloud Security, Cloud SIEM, Incident Management, Notebooks and Bits AI.
- Centralised logging experience with platforms such as Elastic / ELK, Splunk, Grafana Loki or Azure Log Analytics, including log shipping, parsing and retention design.
- Practical log analysis skills — building structured log pipelines, correlating logs with traces and metrics, and tuning log volume and indexing for cost.
- Experience migrating from or coexisting with other APM tools such as Dynatrace, AppDynamics, New Relic or Elastic APM.
- Working knowledge of OpenTelemetry and its interoperability with Datadog.
- Scripting or automation skills in Python, PowerShell, Go or Terraform, and experience with infrastructure as code for monitoring configuration.
- Experience with CI/CD pipelines and embedding observability into the delivery lifecycle.
- Exposure to AIOps practices, anomaly detection and event correlation at scale.
- Insurance, financial services or other regulated industry experience.
Competency
This role carries influence well beyond the tooling. We are looking for the following additional skills and attributes.
- Problem solving — structured, evidence-led analysis under pressure; comfortable working from an ambiguous symptom to a proven root cause, and honest about what the data does and does not show.
- Communication — explains complex performance behaviour clearly to engineers, application owners and executives; writes concise incident narratives and RCA documents that stand up to customer scrutiny.
- Stakeholder management — builds credibility with application teams, DBAs, cloud engineers, service delivery managers and customers across distributed, multi-vendor teams; negotiates standards rather than imposing them.
- Ownership and accountability — takes end-to-end responsibility for APM coverage and quality, and follows issues through to preventive action.
- Continuous improvement — actively looks for opportunities to extend coverage, increase automation and cut manual effort, and measures the result.
- Executive readiness — presents observability posture, risks, quick wins and roadmap in a form suitable for steering committee discussion.
Certifications
- Essential — Datadog certification, including at least one specialist certification relevant to APM.
- Desirable — cloud certification such as Microsoft Azure Administrator / Solutions Architect or AWS Solutions Architect.
- Desirable — Certified Kubernetes Administrator (CKA) or equivalent container platform certification.
- Desirable — ITIL Foundation v4, and familiarity with SRE practices and frameworks.
- Desirable — ServiceNow ITOM or Event Management accreditation.
Sapiens is an equal opportunity employer. We value diversity and strive to create an inclusive work environment that embraces individuals from diverse backgrounds.
Disclaimer: Sapiens India does not authorise any third parties to release employment offers or conduct recruitment drives via a third party. Hence, beware of inauthentic and fraudulent job offers or recruitment drives from any individuals or websites purporting to represent Sapiens. Further, Sapiens does not charge any fee or other emoluments for any reason (including without limitation, visa fees) or seek compensation from educational institutions to participate in recruitment events. Accordingly, please check the authenticity of any such offers before acting on them and were acted upon, you do so at your own risk. Sapiens shall neither be responsible for honouring or making good the promises made by fraudulent third parties, nor for any monetary or any other loss incurred by the aggrieved individual or educational institution.
If you come across any fraudulent activities in the name of Sapiens, please feel free report the incident at sapiens to [email protected]