AI Jobs Map

VCBay · India

Data Scientist

entry_levelfull timePosted 25 days ago
Apply on LinkedInLinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

etlpythonseleniumplaywrightmongodbapache-airflowmlopsci/cdpandasnumpytensorflowpytorchmlflowdockerkubernetesfastapiflaskawsgcpazure

Role Overview
We are looking for a versatile Data Scientist who can build robust data pipelines — from web scraping to AI/ML deployment. The ideal candidate is comfortable working across the full data and AI lifecycle: extracting data at scale, transforming and storing it reliably, and using it to train, deploy, and maintain machine learning models in production.

Key Responsibilities

- Design, build, and maintain scalable web scraping scripts/pipelines using Python (e.g., BeautifulSoup, Scrapy, Selenium, Playwright)

- Handle dynamic websites, pagination, anti-bot mechanisms, proxies, and rate-limiting strategies

- Clean, transform, and normalize scraped data (structured/unstructured) before storage

- Design MongoDB schemas and collections optimized for the type of data being handled

- Implement logic to identify and update unique/duplicate records efficiently (upserts, deduplication strategies)

- Schedule and monitor data pipelines/jobs (via cron, Airflow, or similar orchestration tools)

- Ensure data quality, consistency, and integrity across pipelines, including error logging, retries, and failure recovery

- Support end-to-end AI/ML lifecycle: data collection, preprocessing, feature engineering, model selection, training, and validation

- Fine-tune machine learning/deep learning models and evaluate performance against business requirements

- Package and deploy models into production environments (APIs, batch pipelines, etc.)

- Implement and maintain MLOps practices — versioning (models & data), CI/CD for ML, monitoring model performance/drift

- Collaborate with cross-functional teams to integrate AI models with existing data pipelines, including scraped/transformed data as model input

Required Skills

Core:

- Strong proficiency in Python (writing clean, modular, production-grade code)

- Hands-on experience with MongoDB (schema design, aggregation pipelines, indexing, upsert/dedup logic)

- Web scraping tools/libraries: BeautifulSoup, Scrapy, Selenium, Playwright, or similar

- Data transformation/manipulation using Pandas / NumPy

AI/ML:

- Understanding of the end-to-end AI/ML workflow — data prep, training, evaluation, deployment

- Familiarity with ML/DL frameworks: Scikit-learn, TensorFlow, PyTorch

- Exposure to MLOps tools: MLflow, DVC, Docker, Kubernetes (basic), CI/CD pipelines

- Experience with model deployment (REST APIs via FastAPI/Flask, or cloud ML services)

Good to Have:

- Experience with cloud platforms (AWS / GCP / Azure) for storage, compute, and ML services

- Familiarity with LLMs / NLP (e.g., Hugging Face, LangChain) if relevant to use case

- Knowledge of proxy rotation, CAPTCHA-handling, and anti-scraping evasion techniques

- Experience with workflow orchestration (Airflow, Prefect)

- Version control (Git) and Agile development practices

Soft Skills

- Strong problem-solving mindset, especially around handling unstructured/messy data

- Ability to independently manage multiple workstreams across data engineering and AI/ML

- Good documentation habits for pipelines and models

- Comfortable working in a fast-paced, evolving tech environment

Qualifications

- Bachelor's/Master's degree in Computer Science, Data Science, Engineering, or related field

- Prior experience with production-grade web scraping and/or ML systems is a strong plus

- Portfolio/GitHub showcasing scraping projects and/or ML models is preferred