We're hiring: Data Engineer
We're building an open data platform designed to structure, host, and share large-scale datasets, with a particular focus on supporting a low-resource language. This role is centered on designing the data pipelines that power the whole project: collecting, cleaning, structuring, and making large-scale corpora available.
What you'll do
- Design and build pipelines to collect, clean, and structure large-scale text data
- Set up ingestion processes for heterogeneous corpora (web, documents, community sources) in a low-resource language
- Develop workflows for normalization, deduplication, and data quality validation
- Design data schemas and storage formats suited to exploring and training NLP models
- Build and maintain batch and/or streaming pipelines (orchestration, scheduling, monitoring)
- Work closely with the AI/NLP team to deliver ready-to-use datasets for model training and evaluation
- Optimize the performance and cost of data processing on the cloud
- Document pipelines, schemas, and operational procedures
Tech skills we're looking for
- Python (required), strong experience with large-scale data processing
- Experience with data processing frameworks (Pandas, PySpark, or equivalent)
- Familiarity with pipeline orchestration tools (Airflow, Dagster, Prefect, or similar)
- Familiarity with relational and/or NoSQL databases, and columnar storage formats (Parquet, etc.)
- Experience with GCP, particularly data services (BigQuery, Cloud Storage, Dataflow, or equivalents)
- Basic containerization knowledge (Docker) is a plus
- Git / version control
- An interest in NLP or low-resource languages is a plus