About the Role
We're looking for a Principal Engineer to own the architecture of our data-intensive systems end-to-end, with a particular focus on entity resolution and record linkage at scale. You'll lead a team of data engineers, set technical direction, and work directly with government stakeholders to turn ambiguous, messy source data into concrete engineering decisions. This is a hands-on leadership role — you'll be as comfortable in the codebase as you are in the design review.
What You'll Do
- Own end-to-end architecture for large-scale, data-intensive systems, from ingestion through to production-facing outputs.
- Lead and mentor a team of data engineers, setting technical standards, reviewing designs, and guiding delivery.
- Design and evolve entity resolution / record linkage systems and probabilistic matching approaches capable of operating on millions of records.
- Architect data pipelines for scale, reliability, and maintainability, with close attention to performance and memory-bound workloads.
- Work with columnar data formats (Parquet), SQL, and large structured/semi-structured datasets.
- Build and maintain systems in both Python and Golang, choosing the right tool for each part of the stack.
- Interface directly with government stakeholders to understand requirements, clarify ambiguous or inconsistent source data, and translate that ambiguity into clear engineering decisions and system design.
- Drive technical roadmap decisions, balancing near-term delivery with long-term system health.
Must-Have Qualifications
- 10+ years building data-intensive systems, with demonstrated ability to own architecture end-to-end and lead a team of data engineers.
- Hands-on experience with entity resolution / record linkage, probabilistic matching, and data-pipeline design at scale (millions of records).
- Strong Python and comfortable working in Golang.
- Experience with columnar data formats (Parquet), SQL, and performance/memory-bound workloads.
- Proven ability to interface directly with government stakeholders, and to translate ambiguous, messy source data into concrete engineering decisions.
Nice to Have
- Prior experience delivering systems in government, public-sector, or highly regulated environments.
- Familiarity with statistical record-linkage techniques (e.g., Fellegi–Sunter-style models) or fuzzy/phonetic matching.
- Experience with distributed data processing frameworks (Spark, Dask, or similar) at large scale.
Pay: ₹1,178,328.92 - ₹2,596,627.30 per year
Work Location: Hybrid remote in Ahmedabad, Gujarat 380054