Position Overview
We are looking for an experienced Observability & Reliability Architect to define and advance enterprise monitoring and reliability capabilities across cloud and on-premises environments. This position will focus on observability architecture, platform transformation, production readiness, and improving the reliability and supportability of critical technology services.
The architect will work closely with engineering, infrastructure, operations, and production support teams to establish scalable observability practices and ensure that applications and platforms are fully monitored and operationally ready.
Key Areas of Responsibility
- Establish enterprise-wide observability architecture, standards, and engineering practices.
- Provide technical direction and guidance to infrastructure, operations, and production support teams.
- Define monitoring strategies spanning infrastructure, applications, services, and end-user experience.
- Establish appropriate SLIs, SLOs, service health indicators, and reliability objectives.
- Identify opportunities to improve system availability, reliability, incident detection, and resolution times.
- Review technology and platform changes to confirm adequate monitoring and operational support requirements are in place before implementation.
- Collaborate with application and infrastructure teams to drive adoption of standardized observability patterns.
Observability Transformation
- Lead modernization of enterprise monitoring capabilities and support large-scale migration from legacy monitoring platforms to Grafana Cloud.
- Develop scalable approaches for metrics, logs, traces, dashboards, alerting, and service health monitoring.
- Establish observability patterns for AWS and hybrid infrastructure environments.
- Standardize telemetry pipelines using Grafana Alloy and OpenTelemetry.
- Improve signal quality and reduce unnecessary or low-value alerts through intelligent alerting and anomaly-detection approaches.
- Establish frameworks for predictive monitoring and proactive identification of service degradation.
Grafana Ecosystem & Telemetry Architecture
- Design and implement enterprise observability solutions using the Grafana ecosystem.
- Provide architecture guidance for Grafana Mimir, Loki, Tempo, Grafana Cloud, and Incident Response Management (IRM) capabilities.
- Define reusable instrumentation approaches for infrastructure, applications, APIs, and distributed systems.
- Establish dashboard and visualization standards that provide actionable service and platform health information.
- Design alerting strategies aligned with service objectives, SLIs, SLOs, and error budgets.
- Ensure telemetry pipelines remain scalable, reliable, governed, and fit for enterprise adoption.
- Establish best practices around telemetry consistency, data quality, retention, and accessibility.
Reliability & Production Operations
- Partner with production support organizations to improve operational processes and service reliability.
- Support incident triage, technical escalation, service restoration, and cross-functional communications.
- Ensure new or modified services have appropriate monitoring, logging, alerting, and support procedures before production rollout.
- Participate in major incident investigations and root-cause analysis.
- Identify recurring operational issues and implement improvements to reduce MTTD and MTTR.
- Promote proactive monitoring and operational readiness rather than reactive issue management.
Operational Documentation & Enablement
- Develop and maintain comprehensive Tier 2 and Tier 3 operational runbooks.
- Create standardized onboarding materials, monitoring templates, implementation guides, and support documentation.
- Establish repeatable processes for bringing new applications and services into the enterprise observability platform.
- Work with engineering and support teams to improve observability adoption and operational consistency.
- Maintain architecture and operational standards as platforms, technologies, and business requirements evolve.
Required Technical Experience
- Strong background in enterprise observability, monitoring, reliability engineering, or platform architecture.
- Hands-on experience with Grafana Cloud and the broader Grafana ecosystem.
- Strong knowledge of OpenTelemetry and modern telemetry collection architectures.
- Experience with Grafana Alloy, Mimir, Loki, and Tempo.
- Experience designing monitoring solutions across AWS and on-premises/hybrid environments.
- Familiarity with SLO/SLI frameworks, error budgets, anomaly detection, and intelligent alerting.
- Experience with production support, incident management, root-cause analysis, and operational readiness.
- Working knowledge of Splunk, New Relic, SQL, and Power BI.
- Ability to translate complex technical requirements into enterprise architecture standards and reusable implementation patterns.
- Strong communication and collaboration skills with engineering, infrastructure, operations, and business stakeholders.
Preferred Qualifications
- Experience leading enterprise-scale observability transformation initiatives.
- Experience migrating monitoring platforms from New Relic or similar legacy solutions to Grafana Cloud.
- Knowledge of distributed systems, microservices, cloud-native architectures, and modern application instrumentation.
- Experience establishing observability governance and adoption frameworks across large organizations.
- Familiarity with reliability engineering principles and continuous service improvement practices.