Position: Data Engineer – Generative AI & Data Platforms
Experience: 5+ Years
Location: Bengaluru / Chennai (Hybrid/Remote)
Employment Type: Full-Time
About Latinum:
Latinum is hiring for a client and is looking for an experienced Data Engineer with 5+ years of experience in designing, developing, and maintaining scalable data platforms and ETL/ELT pipelines. The ideal candidate will have strong expertise in Python, SQL, PySpark, cloud data platforms, and modern data engineering frameworks, along with hands-on exposure to Generative AI and Large Language Models (LLMs).
You will play a key role in building reliable, high-performance data solutions that support analytics, reporting, and AI/ML initiatives.
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines and data workflows.
- Build data ingestion processes from structured and unstructured sources.
- Develop and optimize data models, data warehouses, and data lakes.
- Write and optimize complex SQL queries for high-performance processing.
- Implement data cleansing, validation, transformation, and quality checks.
- Develop batch and real-time data processing pipelines.
- Integrate data solutions with AWS, Azure, or GCP.
- Work with large-scale distributed datasets using Apache Spark/PySpark.
- Support data engineering requirements for Generative AI and LLM-based solutions.
- Collaborate with data scientists, analysts, and application teams.
- Follow engineering best practices including Git, CI/CD, testing, monitoring, and documentation.
Required Skills
- 5+ years of hands-on experience in Data Engineering.
- Strong programming skills in Python.
- Hands-on experience with PySpark and Apache Spark.
- Advanced SQL, including query optimization and performance tuning.
- Experience with AWS, Azure, or GCP.
- Strong experience building scalable ETL/ELT pipelines.
- Hands-on experience with Generative AI and Large Language Models (LLMs).
- Strong understanding of data warehousing and dimensional modeling.
- Experience working with large-scale distributed datasets.
- Familiarity with Git and CI/CD pipelines.
- Understanding of Agile/Scrum methodologies.
Preferred Skills
- Experience with Apache Airflow or similar orchestration tools.
- Knowledge of Kafka and real-time data streaming.
- Exposure to Databricks and modern data lake technologies such as Delta Lake, Apache Iceberg, or Apache Hudi.
- Familiarity with Docker, Kubernetes, and Terraform/IaC.
- Understanding of data governance and metadata management.
- Strong analytical, problem-solving, and communication skills.
Qualification
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline.
Interested candidates can share their resume at [email protected] & [email protected]