Our client is looking for an experienced Data Engineer with 3–5 years of hands-on Data Engineering experience.
The ideal candidate should have strong experience in Python, PySpark, and SQL, along with exposure to modern cloud and data platforms such as AWS, Azure, GCP, Snowflake, or Databricks.
Candidates do not need to have experience across all cloud platforms or technologies. Strong hands-on Data Engineering experience with Python, PySpark, SQL, ETL/ELT, and scalable data pipelines is the primary requirement.
Exposure to Generative AI, Large Language Models (LLMs), RAG, AI Agents, or AI-enabled data solutions is preferred but not mandatory for candidates with a strong Data Engineering background.
Key Responsibilities
- Design, develop, and maintain scalable and reliable data pipelines using Python, PySpark, SQL, and modern data engineering platforms.
- Develop and optimize ETL/ELT workflows for batch and/or real-time data processing.
- Write and optimize complex SQL queries for data transformation, integration, and analytics.
- Build data solutions using one or more cloud platforms such as AWS, Azure, or GCP.
- Work with modern data platforms such as Snowflake, Databricks, Spark, or similar technologies.
- Process and transform large-scale structured and unstructured datasets.
- Troubleshoot and optimize pipelines for performance, scalability, reliability, and cost efficiency.
- Integrate data from multiple sources, APIs, databases, files, and cloud storage platforms.
- Support data quality, monitoring, production troubleshooting, and data validation.
- Where applicable, contribute to GenAI/LLM-enabled data engineering solutions.
- Collaborate with engineering, analytics, AI/ML, product, and business teams in Agile and/or Waterfall environments.
Required Skills
- 3–5 years of hands-on Data Engineering experience.
- Strong hands-on experience with Python.
- Strong hands-on experience with PySpark / Apache Spark.
- Strong practical knowledge of SQL, including complex queries, joins, transformations, window functions, and query optimization.
- Good understanding of ETL/ELT, data pipelines, data processing, and data integration concepts.
- Experience working with large datasets and distributed data processing.
- Experience with at least one of the following:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
- Snowflake
- Databricks
- Good understanding of data engineering fundamentals, including data quality, scalability, reliability, and performance optimization.
- Strong problem-solving, communication, and collaboration skills.
Cloud / Platform Experience
Candidates with experience in any one or more of the following environments are encouraged to apply:
- AWS – S3, Glue, EMR, Redshift, Lambda, Athena, Kinesis, Step Functions, or related services.
- Azure – Azure Data Factory, Azure Databricks, ADLS, Synapse Analytics, Event Hubs, Functions, or related services.
- GCP – BigQuery, Dataflow, Dataproc, Cloud Storage, Pub/Sub, Composer, or related services.
- Snowflake – Snowflake SQL, Snowpipe, Streams, Tasks, performance optimization, or data modeling.
- Databricks – Apache Spark, PySpark, Delta Lake, Unity Catalog, Databricks Workflows, or Lakehouse architecture.
Candidates are not expected to have experience with every platform listed above.
GenAI / AI Exposure – Preferred
Exposure to Generative AI or LLM-based solutions is an advantage, but deep AI/ML expertise is not required.
Relevant exposure may include:
- Generative AI / Large Language Models
- Retrieval-Augmented Generation (RAG)
- AI Agents / Agentic AI
- Prompt Engineering
- Vector Databases
- Embeddings and semantic search
- LLM APIs
- AI-powered data processing
- Integration of GenAI solutions with enterprise data pipelines
- AWS Bedrock, Azure OpenAI, Vertex AI, or similar AI platforms
Nice-to-Have Skills
- Experience with Snowflake and/or Databricks.
- Experience with Airflow, Kafka, Spark Streaming, or similar technologies.
- Experience with data lakes, data warehouses, or Lakehouse architectures.
- Experience with dimensional modeling and analytical data platforms.
- Exposure to CI/CD, Git, and production deployment practices.
- Experience with real-time or near-real-time pipelines.
- Exposure to GenAI/LLM integration with data platforms.
- Knowledge of data governance, security, lineage, or observability.
Candidates with strong Python, PySpark, SQL, ETL/ELT, and core Data Engineering skills should be considered even if they have limited AWS, Snowflake, Databricks, or GenAI exposure.
About YMinds.AI
YMinds.AI is a technology and talent solutions company focused on connecting businesses with high-quality technology professionals. We help organizations build strong engineering and AI teams by identifying, evaluating, and delivering skilled talent across emerging and established technology domains.
Keywords
Data Engineer, Cloud Data Engineer, AWS Data Engineer, Azure Data Engineer, GCP Data Engineer, Snowflake Data Engineer, Databricks Data Engineer, Python, PySpark, Apache Spark, SQL, ETL, ELT, Data Pipelines, Data Engineering, Snowflake, Databricks, AWS, Azure, GCP, Big Data, Data Lake, Data Warehouse, Lakehouse, Airflow, Kafka, Generative AI, GenAI, LLM, RAG
Hashtags
#DataEngineer #DataEngineering #CloudDataEngineer #Python #PySpark #SQL #ApacheSpark #Snowflake #Databricks #AWS #Azure #GCP #BigData #ETL #ELT #DataPipelines #GenAI #LLM #AIJobs #DataJobs #CloudJobs #TechJobs #Hiring #HiringNow #YMindsAI