Role: Senior DevOps / Site Reliability Engineer
Function: DevOps / Site Reliability Engineering / Infrastructure
Location: Kolkata or Gurugram, India
Type: Full-time
Industry: AI, Medical SaaS, Dental Technology
About Company
The company builds AI-powered medical SaaS for the U.S. dental market.
Its all-in-one platform unifies analytics, scheduling, communications, payments, and marketing automation. Thousands of dental practices across North America rely on it daily.
The company was founded in 2015 by a practicing dentist, a data strategist, and a PhD technologist. It is mission-driven and execution-heavy, with a team of ~88 people.
Engineers get high ownership, cross-functional collaboration, and clear, measurable impact.
Position Overview
This is a builder's role with full production ownership across the company's AWS infrastructure — not a monitoring seat or ticket queue. You will own Terraform modules, build and maintain GitHub Actions pipelines, and drive the reliability of a platform used by thousands of dental practices. The role offers flexibility between a normal and night shift, with the night window providing uninterrupted time for high-impact work: database migrations, pipeline rewrites, and staged cutovers.
Role & Responsibilities
- Own and extend Terraform modules across AWS — ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, Route 53, CloudWatch — managing state safely and eliminating manually created resources via imports and refactors.
- Build and maintain GitHub Actions pipelines for build, test, containerization, and deployment; migrate legacy pipelines (Jenkins, CircleCI, GitLab CI) onto a single maintainable platform with staged rollouts, health-gated releases, and fast rollback.
- Own the production change window on your shift: deploys, schema migrations against live systems, cutovers, and maintenance with minimal disruption.
- Run and improve observability — dashboards, SLOs, and alerting calibrated to real user impact — and debug containerized applications in production across logs, metrics, traces, and resource limits.
- Operate PostgreSQL in production: read query plans, identify slow queries, manage connections, locks, and replication; operate supporting services including Redis/ElastiCache and message brokers.
- Write Python and shell automation to remove repetitive operational work; build self-service tooling, runbooks, and guardrails that make the wider engineering team faster.
- Write clear, complete shift handover notes and maintain infrastructure documentation; participate in on-call rotation and drive blameless post-incident reviews.
Must Have Criteria
- 6+ years in DevOps, SRE, Platform, or Infrastructure engineering with genuine production ownership.
- Hands-on AWS experience across ECS/Fargate or EKS, RDS, VPC networking, IAM, ALB/NLB, S3, and CloudWatch.
- Hands-on Terraform experience: writing modules, managing state, and reviewing plans in production environments.
- Practical CI/CD experience with GitHub Actions (or GitLab CI, CircleCI, or Jenkins with demonstrated ability to migrate).
- Docker proficiency with demonstrated ability to debug containerized applications in production.
- Working PostgreSQL knowledge: query plans, slow query analysis, connections, locks, and replication.
- Solid Linux fundamentals, shell scripting, and Python for operational automation.
Nice to Have
- Experience with Celery, RabbitMQ, NATS, or Kafka in production.
- Secrets management using AWS Secrets Manager, SOPS, or Vault.
- Exposure to sharded or multi-tenant database architectures.
- Experience in compliance-sensitive environments (HIPAA or SOC 2).
- Kubernetes production experience.
- Prior experience as the sole engineer on shift or in a night/off-hours rotation.
- Telephony or VoIP infrastructure exposure.
What We Offer
- Full infrastructure ownership with a protected change window for high-impact work — no ticket shuffling.
- Shift flexibility: normal or night shift based on candidate preference and business coverage needs.
- Modern stack: AWS, Terraform, GitHub Actions, Docker, PostgreSQL, and Kafka-compatible messaging.
- Direct, measurable impact on the reliability of a platform serving thousands of healthcare practices.
- Autonomy, cross-functional collaboration, and support for continuous learning and growth.