AI Jobs Map

Infoser Technological solutions, Hardware, Big Data Experts ยท Madrid, Community of Madrid, Spain

Data Engineer - Python & ML

seniorfull timePosted 3 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

data-engineeringpythondata-modelingexcelpandaspostgresqldjangosqldata-structuresawsetldata-governancemachine-learning

We are looking for an experienced Data Engineer to join an international technology project focused on building and evolving a data-processing platform that handles structured and unstructured information from multiple sources.

You will work on multi-format data ingestion, data normalization, entity reconciliation, relational data modeling and classification workflows, helping transform heterogeneous information into reliable and standardized datasets.

๐Ÿ”Ž What will you work on?

- Build and maintain Python pipelines for processing Excel, PDF and XML sources.

- Extract structured records and specification lists from Excel using pandas and openpyxl.

- Process unstructured requirements from PDF documents using pypdf.

- Parse XML documents securely using defusedxml, including protection against XXE vulnerabilities.

- Map raw source fields into a canonical data model.

- Normalize categorical and hierarchical attributes such as product line, variant, region and language.

- Reconcile inconsistent terminology and entity names using fuzzy and approximate matching.

- Maintain relational data models using PostgreSQL and SQLAlchemy ORM.

- Support and improve existing scikit-learn classification workflows.

- Retrain and evaluate models and adjust decision thresholds according to precision, recall and F1 requirements.

๐Ÿ› ๏ธ What are we looking for?

- Strong professional experience with Python.

- Experience with Django for backend development and data-driven applications.

- Experience working with pandas and openpyxl.

- Experience processing or extracting information from PDF and XML documents.

- Knowledge of data normalization, schema mapping and canonical data models.

- Experience with fuzzy matching techniques, ideally using rapidfuzz.

- Strong knowledge of SQL and PostgreSQL.

- Experience with SQLAlchemy ORM.

- Practical experience with scikit-learn classification workflows.

- Understanding of classification metrics such as precision, recall and F1-score.

- Experience with algorithms such as KNN, Random Forest or SVM.

โž• Nice to have

- Experience with AWS Aurora or other cloud-managed relational databases.

- Experience designing data ingestion or ETL/ELT pipelines.

- Experience working with heterogeneous datasets and complex source-data reconciliation.

- Knowledge of model threshold tuning and classification performance optimization.

๐Ÿ’ป Technology Stack

Python ยท pandas ยท openpyxl ยท pypdf ยท defusedxml ยท rapidfuzz ยท PostgreSQL ยท SQLAlchemy ยท scikit-learn ยท KNN ยท Random Forest ยท SVM ยท AWS Aurora

If you enjoy solving complex data-processing problems and working at the intersection of Data Engineering, Data Quality and Machine Learning, we would love to hear from you.

๐Ÿ“ฉ Apply or contact us to learn more about the project.

More jobs at Infoser Technological solutions, Hardware, Big Data Experts