About the Role
We're looking for a hands-on Machine Learning Engineer to own the design and implementation of our entity resolution and record-linkage pipeline. This role sits at the intersection of applied data science and large-scale data engineering: you'll be building the models and heuristics that decide when two records from different, messy source systems represent the same real-world entity — and the ETL infrastructure that feeds them.
This is not a "pipeline plumbing" role. We need someone who can genuinely apply data-science and statistical techniques to a hard matching problem, evaluate models rigorously, and iterate based on measured outcomes.
What You'll Do
- Design, build, and evaluate machine learning models for entity matching, including feature engineering on structured and semi-structured data.
- Implement and tune statistical record-linkage methods, including Fellegi–Sunter-style models — estimating field-level match/non-match weights from labelled data, computing likelihood-ratio based match scores, and setting evidence-driven decision thresholds.
- Build and maintain large-scale ETL processes to ingest heterogeneous databases, map disparate schemas to a common data model, and profile and repair data quality issues (placeholders, junk values, inconsistent encodings, malformed fields).
- Develop fuzzy and phonetic string-matching logic, name normalization routines, and blocking/candidate-generation strategies to make matching tractable at scale.
- Calibrate scoring thresholds against labelled data and business requirements, balancing precision and recall.
- Group matched records using graph-style techniques (clustering / clique formation) to resolve multi-record entities.
- Run iterative rerun-and-tune cycles: evaluate model/pipeline output, diagnose failure modes, and refine features, weights, and thresholds accordingly.
- Write clean, production-grade Python and SQL, and maintain Parquet-based data pipelines.
Must-Have Qualifications
- Genuine ML experience — feature engineering, model evaluation methodology, and hands-on application of data-science techniques to real problems (not just orchestrating pipelines).
- Working knowledge of statistical record-linkage methods such as the Fellegi–Sunter model, including:
- Estimating field-level match/non-match weights from labelled data
- Likelihood-ratio based scoring
- Setting evidence-driven decision thresholds
- Proven experience with large-scale ETL:
- Ingesting heterogeneous databases
- Mapping them to a common data model
- Profiling and repairing data quality issues (placeholders, junk values, inconsistent encodings)
- Strong Python (Pandas / PySpark or similar), SQL / Postgres, and Parquet-based pipelines.
- Practical experience with fuzzy and phonetic string matching, name normalization, blocking/candidate-generation strategies, and scoring/threshold calibration.
Nice to Have
- Familiarity with graph-style grouping (clustering / clique formation) for multi-record entity resolution.
- Experience running iterative rerun-and-tune workflows on production matching pipelines.
- Exposure to distributed data processing at scale (Spark, Dask, or similar).
Pay: ₹539,315.28 - ₹1,867,153.48 per year
Work Location: Hybrid remote in Ahmedabad, Gujarat 380054