š Join us in Luxoft!
š¦Flexible working hours
š©ŗPrivate Medical & Dental care & Life Insurance
š° Paid Referrals
šš½ āļø MyBenefit program (sports card, well-being program etc.)
š Internal Mobility program - possibility of rotation between projects, locations, accounts
š LuxTalent platform (webinars, training, courses)
š Project Description
Senior data engineer developing and operating Databricks / Spark pipelines and Delta Lake lakehouse layers on Azure for the Client. Accountable for pipeline reliability, data freshness and dataset quality against agreed SLAs.
Key tasks
⢠Develop and operate Databricks / Spark pipelines (PySpark, SQL, Delta Live Tables or Workflows).
⢠Design Delta Lake / lakehouse layers (bronzeāsilverāgold), partitioning and Unity Catalog governance.
⢠Build ETL/ELT jobs and orchestration with Azure Data Factory and/or Airflow; manage dependencies and retries.
⢠Implement data-quality checks and validation (expectations, reconciliation, anomaly alerts).
⢠Tune Spark job performance and cluster cost (autoscaling, Photon, job clusters, spot).
⢠Manage schema evolution and change control; document lineage and transformations.
š Responsibilities
⢠Pipeline reliability and data freshness against agreed SLAs.
⢠Accuracy and completeness of curated datasets.
⢠Schema change management and backward compatibility for downstream consumers.
⢠Documentation of data lineage and transformations (Unity Catalog, data catalogue).
š Skills
What is relevant to have
⢠Bachelor's degree in Computer Science, Engineering, Information Systems or a related field, or equivalent practical experience.
⢠7+ years in data engineering, of which 3+ on Databricks / Apache Spark and 2+ on Azure data services (ADF, ADLS, Delta Lake).
⢠Databricks (Workflows, Delta Live Tables, Unity Catalog), Apache Spark (PySpark, Spark SQL), Delta Lake.
⢠Azure: Data Factory, Data Lake Storage Gen2, Key Vault, Event Hubs, Synapse or SQL DB; Azure DevOps CI/CD for notebooks and jobs.
⢠Python and SQL at expert level; data modelling (dimensional, data vault) and ELT design.
⢠Orchestration (ADF, Airflow), data-quality frameworks (Great Expectations, DLT expectations), monitoring and alerting.
⢠Performance and cost tuning of Spark workloads; Git-based development and testing of pipelines.
What is nice to have
⢠Databricks Certified Data Engineer Professional; Azure DP-203.
⢠Streaming (Structured Streaming, Kafka / Event Hubs).
⢠dbt, Power BI semantic models, MLflow.
⢠Experience in financial services, sovereign wealth / investment holding or other regulated enterprise environments.
⢠Experience working with distributed teams (onsite UAE with nearshore India / offshore Poland squads).
š Languages
English: C1 Advanced
š Seniority
Senior
š Work mode
remote in Poland