AI Jobs Map

Basicana · Ahmedabad, Gujarat

Machine Learning Engineer — Data Matching & Record Linkage

Hybridfull timePosted today
Apply on IndeedOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

machine-learningdata-sciencedata-engineeringetldata-governancepythonsqlpandasapache-sparkpostgresql

About the Role

We're looking for a hands-on Machine Learning Engineer to own the design and implementation of our entity resolution and record-linkage pipeline. This role sits at the intersection of applied data science and large-scale data engineering: you'll be building the models and heuristics that decide when two records from different, messy source systems represent the same real-world entity — and the ETL infrastructure that feeds them.

This is not a "pipeline plumbing" role. We need someone who can genuinely apply data-science and statistical techniques to a hard matching problem, evaluate models rigorously, and iterate based on measured outcomes.

What You'll Do

- Design, build, and evaluate machine learning models for entity matching, including feature engineering on structured and semi-structured data.

- Implement and tune statistical record-linkage methods, including Fellegi–Sunter-style models — estimating field-level match/non-match weights from labelled data, computing likelihood-ratio based match scores, and setting evidence-driven decision thresholds.

- Build and maintain large-scale ETL processes to ingest heterogeneous databases, map disparate schemas to a common data model, and profile and repair data quality issues (placeholders, junk values, inconsistent encodings, malformed fields).

- Develop fuzzy and phonetic string-matching logic, name normalization routines, and blocking/candidate-generation strategies to make matching tractable at scale.

- Calibrate scoring thresholds against labelled data and business requirements, balancing precision and recall.

- Group matched records using graph-style techniques (clustering / clique formation) to resolve multi-record entities.

- Run iterative rerun-and-tune cycles: evaluate model/pipeline output, diagnose failure modes, and refine features, weights, and thresholds accordingly.

- Write clean, production-grade Python and SQL, and maintain Parquet-based data pipelines.

Must-Have Qualifications

- Genuine ML experience — feature engineering, model evaluation methodology, and hands-on application of data-science techniques to real problems (not just orchestrating pipelines).

- Working knowledge of statistical record-linkage methods such as the Fellegi–Sunter model, including:

- Estimating field-level match/non-match weights from labelled data

- Likelihood-ratio based scoring

- Setting evidence-driven decision thresholds

- Proven experience with large-scale ETL:

- Ingesting heterogeneous databases

- Mapping them to a common data model

- Profiling and repairing data quality issues (placeholders, junk values, inconsistent encodings)

- Strong Python (Pandas / PySpark or similar), SQL / Postgres, and Parquet-based pipelines.

- Practical experience with fuzzy and phonetic string matching, name normalization, blocking/candidate-generation strategies, and scoring/threshold calibration.

Nice to Have

- Familiarity with graph-style grouping (clustering / clique formation) for multi-record entity resolution.

- Experience running iterative rerun-and-tune workflows on production matching pipelines.

- Exposure to distributed data processing at scale (Spark, Dask, or similar).

Pay: ₹539,315.28 - ₹1,867,153.48 per year

Work Location: Hybrid remote in Ahmedabad, Gujarat 380054

More jobs at Basicana