Data Engineer – Databricks / Lakehouse
Location:
Alexandra Building
Contract Duration:
12 Months
Experience:
Minimum 6 years of relevant experience
About the Role
We are looking for an experienced
Data Engineer
to design, develop and operationalise enterprise-scale data platforms, Lakehouse solutions and data products.
The successful candidate will have strong hands-on experience in
Data Engineering, Databricks, Apache Spark, PySpark, Python and SQL
, with exposure to modern Data Lake/Lakehouse architectures, cloud platforms and large-scale data ingestion and processing.
The role involves building scalable batch and streaming pipelines, developing reusable data products, implementing data quality frameworks, and supporting modern analytics, AI and GenAI use cases.
Key Responsibilities
Design, develop and operationalise enterprise
Data Lake, Lakehouse and data engineering platforms
.
Develop scalable
batch, streaming, CDC and API-based data ingestion pipelines
.
Build and maintain data pipelines using
Spark, PySpark, Python, SQL and related Big Data technologies
.
Develop reusable foundation and business
data products
, including appropriate data contracts, SLAs and data quality controls.
Implement data ingestion, transformation, reconciliation and data quality frameworks.
Work with modern Lakehouse and open table technologies such as
Delta Lake, Apache Iceberg and Apache Hudi
.
Develop solutions for structured, semi-structured and unstructured data, including extraction and processing of content from different file formats.
Work with streaming and distributed data technologies such as
Kafka, Spark Streaming, Flink, Airflow, Hive, Trino or Dremio
.
Design data architectures supporting
AI, NLP, RAG, vector search and GenAI/agentic applications
.
Support ingestion, curation, governance and consumption of unstructured data for AI-driven analytics.
Expose data through APIs, event streams, dashboards and other enterprise consumption channels.
Perform performance tuning, troubleshooting, production support and root cause analysis.
Implement automated deployment and engineering practices using
Docker, Kubernetes/OpenShift and CI/CD pipelines
.
Work closely with architecture, engineering, analytics and business teams across multiple projects.
Prepare technical documentation, deployment guides and operational runbooks.
Ensure solutions comply with engineering standards, security requirements, DevSecOps controls and software delivery practices.
Requirements
Bachelor's degree in
Computer Science, Engineering, Information Technology
or a related discipline.
Minimum
6 years of relevant experience
in Data Engineering, Big Data, Data Lake, Data Warehouse or Lakehouse implementations.
Strong hands-on experience with
Apache Spark, PySpark, Python and SQL
.
Experience with one or more enterprise data platforms such as
Databricks, Snowflake, Cloudera, Azure, AWS or GCP
.
Strong experience developing
ETL/ELT, data ingestion, transformation and data processing pipelines
.
Experience working with Data Lake/Lakehouse architectures and technologies such as
Delta Lake, Iceberg or Hudi
.
Experience with distributed and streaming technologies such as
Kafka, Spark Streaming, Flink, Hive or Airflow
.
Good understanding of
data modelling, metadata management, data lineage, governance and data quality
.
Experience with containerisation and DevOps technologies such as
Kubernetes, OpenShift, Docker, Jenkins, Git and CI/CD
.
Strong troubleshooting, performance optimisation and root cause analysis skills.
Experience working in Agile environments and delivering enterprise-scale technology solutions.
Strong communication, collaboration and stakeholder management skills.
Good to Have
Experience developing
data products or enterprise data marketplace solutions
.
Experience supporting
RAG, vector search, GenAI, NLP or AI/ML data pipelines
.
Experience processing multimodal or unstructured data such as documents, images, audio and video.
Knowledge of
Scala or Java
for data engineering.
Experience with
MLflow, Spark MLlib, scikit-learn or XGBoost
.
Experience developing internal engineering tools using
Python, shell scripting, Flask or React
.
Exposure to
Teradata, Netezza, Greenplum or other MPP migration programmes
.
Relevant certifications such as
Databricks Certified Data Engineer, Azure Data Engineer Associate, Google Professional Data Engineer, SnowPro or DAMA CDMP
.
Interested candidates are kindly requested to email their CV with their experience to
We look forward to your application!
NTT Singapore Pte Ltd (NTTS)
is the regional headquarters of NTT Communications Corporation (NTT Com) for Asia Pacific Region.
Established in 1997, NTT Singapore has more than 10 years of expertise in providing information and communications technology (ICT) solutions worldwide.
NTT Singapore offers diverse high-quality connectivity, data centre solutions, security services, IT management services, voice and conferencing solutions and solution integration services to its enterprise customers.
NTT Communications is a wholly owned subsidiary of Nippon Telegraph and Telephone Corporation (NTT Corp.), one of the world’s largest providers of telecommunications services.
In 2013, NTT Corp. is ranked no.1 in telecom industry in the Fortune Global 500* list with operating revenues of more than $133,077 million. It is positioned 32nd among the top 500 corporations worldwide.
NTT Com's extensive global infrastructure includes Arcstar secure private networks, which cover 196 countries/regions and a tier-1 IP backbone network connected with major ISPs worldwide, as well as secure data centers at over 150 locations worldwide.