Data Scientist, Identity & Graph Analytics
About the Role
A high-growth data and analytics organization is seeking a Data Scientist to join its Big Data R&D team. This team is responsible for developing large-scale identity graph and entity resolution capabilities that support compliance, risk, and fraud-related products.
In this role, you will work with massive datasets to develop graph-based algorithms, build scalable data pipelines, support machine learning initiatives, and evaluate new data sources that enhance identity intelligence solutions. You will collaborate closely with experienced data scientists and engineers while expanding your expertise in large-scale machine learning, graph analytics, distributed computing, and entity resolution.
Key Responsibilities
- Contribute to the development of machine learning, statistical, data mining, and graph-based algorithms for identity resolution, anomaly detection, and large-scale data analysis.
- Analyze complex datasets to improve identity matching, record linkage, and entity resolution capabilities.
- Build and maintain scalable data processing pipelines, including ETL workflows, feature engineering, data normalization, and quality checks.
- Support model development through feature engineering, exploratory analysis, error investigation, and experimentation.
- Evaluate new internal and third-party data sources by assessing quality, coverage, and impact on analytical models.
- Develop and maintain SQL, Python, and related data workflows for extraction, transformation, and validation processes.
- Provide analytical support for compliance, risk, and operational teams through investigations, reporting, dashboards, and deep-dive analyses.
- Present findings and recommendations clearly to technical and non-technical stakeholders.
- Operate effectively in a fast-paced environment while owning projects and delivering results independently.
Qualifications
- Master's degree with 2+ years of relevant experience, PhD with 1+ years of experience, or equivalent industry experience in data science, analytics, or machine learning.
- Proficiency in Python, Scala, or another data-focused programming language.
- Strong SQL skills and experience working with large-scale data warehouse or data lake environments.
- Experience with Spark or PySpark and common machine learning libraries such as scikit-learn, XGBoost, TensorFlow, or PyTorch.
- Familiarity with Linux/Unix environments and cloud platforms, particularly AWS services.
- Understanding of supervised and unsupervised machine learning techniques, clustering methods, similarity metrics, and model evaluation approaches.
- Experience with distributed data processing and large-scale analytical workflows.
- Ability to frame ambiguous problems, identify key questions, and iterate quickly based on feedback.
Preferred Qualifications
- Experience working with graph databases, graph analytics, or network analysis technologies.
- Familiarity with identity resolution, entity matching, compliance, fraud, risk, or trust and safety use cases.
- Experience with search technologies, NoSQL databases, or workflow orchestration tools.
- Exposure to automated data pipelines, experimentation frameworks, and production analytics environments.
- Experience evaluating data quality and integrating new data sources into analytical products.