We are seeking a Lead Data Science Engineer to design and deliver reusable data sharing adapters and governed access patterns across a cloud lakehouse and external data platforms. You will guide high-impact architecture decisions, ensure reliable data and AI workflows, and help the team ship secure integrations.
Responsibilities
-
Design a reusable lakehouse write layer with dual-format metadata to support multiple consumers
-
Build and validate ingestion pipeline patterns from object storage to an analytical warehouse for structured operational data
-
Implement change-data-capture patterns for real-time and near-real-time data movement into the lakehouse
-
Develop dependency-aware bookkeeping and data lineage tracking patterns across pipelines
-
Ensure adapter code is modular, version-controlled, tested, and reusable across new integrations
-
Configure external table definitions for shared lakehouse data products in Snowflake catalogs
-
Validate zero-copy read access from Snowflake to shared Iceberg and Delta tables without data movement
-
Implement tenant-scoped access controls aligned with external catalog governance requirements
-
Implement and certify a Delta Sharing adapter for live data sharing to Databricks consumers
-
Configure Delta Sharing endpoints and manage sharing agreements for multiple tenants
-
Validate consumer access via supported clients while meeting freshness and latency expectations
-
Register connector types in a governed connector registry and enforce auditable RBAC on data-out paths
-
Implement metering hooks compatible with governed billing requirements for external data flows
Requirements
-
5+ years of data science or ML engineering experience with production Python and pandas
-
Experience building RAG applications using embeddings and retrieval pipelines
-
Experience writing SQL for analytical data workflows and validation
-
Strong technical leadership skills to drive architecture decisions and mentor peers
-
Proven project delivery skills across multi-system data integration workstreams
-
Solid software engineering skills in modular design, testing, debugging, and Git workflows
-
Hands-on cloud platform skills with GCP, AWS, or Azure fundamentals
-
Strong LLM fundamentals knowledge including tokenization, attention, context windows, and sampling
-
Practical prompt engineering skills with structured outputs and few-shot techniques
-
Robust evaluation and monitoring skills including drift, performance, and hallucination detection
-
Strong communication and collaboration skills across engineering and data stakeholders
-
Upper-Intermediate English proficiency (B2, Upper-Intermediate)
-
Active experience using AI-assisted development tools such as Claude Code, GitHub Copilot, or Cursor
Nice to have
-
Google Cloud Platform experience with BigQuery and object storage patterns
-
Large Language Models (LLM) API integration experience including streaming, rate limits, and cost controls
-
Vector database experience with indexing, chunking strategies, and retrieval tuning
-
Experience optimizing LLM latency and cost using caching, batching, and model routing
We offer
-
International projects with top brands
-
Work with global teams of highly skilled, diverse peers
-
Healthcare benefits
-
Employee financial programs
-
Paid time off and sick leave
-
Upskilling, reskilling and certification courses
-
Unlimited access to the LinkedIn Learning library and 22,000+ courses
-
Global career opportunities
-
Volunteer and community involvement opportunities
-
EPAM Employee Groups
-
Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.