AI Jobs Map

AIMLEAP · Bengaluru, Karnataka, India

Web scraping Engineer (Immediate joiners preferred)

entry_levelfull timePosted 2 days ago
Apply on LinkedInOpens the original posting. AI Jobs Map never asks for your details.

Stack mentioned

javascriptplaywrightseleniumetlobservabilitypythonnode.jshtmlredisrabbitmqapache-kafkasqlpostgresqlmysqldockerawsgcpazures3llm

Company Description

AIMLEAP is a global AI development and technology consulting company specializing in end-to-end full-stack solutions at the intersection of AI, data engineering, and software development. With over 14 years of industry experience, AIMLEAP has co-created more than 50 successful products and enabled over 750 businesses worldwide in their digital transformation journeys. The organization is CMMI Level 3 apprised and ISO/IEC 27001:2022 certified, underscoring a strong commitment to quality, performance, scalability, and security. Recognized as a Great Place to Work®, AIMLEAP operates global delivery centers in the USA, Canada, India, and Australia, offering services that range from AI-powered data management and web scraping to advanced analytics and automation. The company builds and maintains AI and data platforms such as RADAR, PRICELEAP, and Rapid Extract, providing real-time insights and clean web data for smarter decision-making.

About the Role

We are looking for a hands-on Web Scraping / Crawling Engineer with 3–5 years of experience in web scraping, browser automation, and scalable data extraction.

The ideal candidate should have strong experience working with dynamic and JavaScript-heavy websites, building reliable crawling workflows, handling crawling failures, and working with distributed processing systems.This is a highly technical, hands-on role. You will be expected to design, develop, debug, optimize, and maintain web crawling and data extraction systems.

Responsibilities

- Design, develop, and maintain scalable web crawling and scraping systems for dynamic and JavaScript-heavy websites.

- Develop browser automation workflows using Playwright, Selenium, Puppeteer, or similar frameworks.

- Investigate and resolve crawling issues such as 403/429 responses, redirects, timeouts, rendering failures, session issues, and anti-bot challenges.

- Develop robust retry, fallback, and failure-handling mechanisms to improve crawler reliability.

- Work with cookies, sessions, browser contexts, headers, proxies, and related crawling mechanisms to maintain state and improve crawling reliability.

- Build and optimize concurrent and distributed scraping workflows using asynchronous processing, queues, and worker-based architectures.

- Design data pipelines covering URL processing, crawling, extraction, validation, transformation, and storage.

- Optimize crawler performance, including concurrency, browser lifecycle, resource utilization, timeouts, and request handling.

- Implement monitoring and observability for crawl success rates, failure types, latency, retries, and worker performance.

- Debug complex crawling problems and identify root causes rather than relying only on one-off fixes.

- Develop reusable crawling components and frameworks that can support multiple websites and use cases.

- Collaborate with data engineering, AI, backend, and product teams to deliver reliable and structured web data.

- Evaluate and adopt new web crawling, browser automation, and data extraction technologies where appropriate.

Required Skills & Qualifications

- 3–5 years of hands-on experience in web scraping, web crawling, browser automation, or web data engineering.

- Strong programming experience in Python. JavaScript/Node.js is a plus.

- Strong hands-on experience with Scrapy, Playwright, Selenium, Puppeteer, or similar browser automation frameworks.

- Strong understanding of HTTP, HTML, DOM, JavaScript rendering, redirects, cookies, sessions, headers, and browser contexts.

- Experience with browser fingerprinting, WAFs, and modern anti-bot mechanisms.

- Experience working with dynamic, JavaScript-heavy, and asynchronous websites.

- Good understanding of asynchronous programming, concurrency, and parallel processing.

- Experience with job queues, distributed workers, or message-based processing systems such as Redis, RabbitMQ, Kafka, Celery, or similar technologies.

- Understanding of retry mechanisms, error handling, idempotency, rate limiting, and backpressure in distributed systems.

- Hands-on experience with proxy management, session handling, and anti-bot challenges.

- Strong debugging and analytical skills, with the ability to investigate and resolve complex crawling failures.

- Experience working with SQL databases such as PostgreSQL or MySQL.

- Familiarity with Docker and cloud platforms such as AWS, GCP, or Azure is a plus.

Good to Have

- Experience with Redis Streams, RabbitMQ, Kafka, Celery, or other distributed processing technologies.

- Experience processing large volumes of URLs or web data.

- Experience with AWS S3 or S3-compatible object storage.

- Experience building ETL/data processing pipelines.

- Experience with LLM-based data extraction or structured data extraction.

- Experience with crawler monitoring, logging, metrics, and observability.

What We’re Looking For

- Strong hands-on engineering mindset rather than purely project or delivery management experience.

- Ability to independently debug, investigate, and solve difficult crawling problems.

- Strong understanding of how browsers, HTTP requests, sessions, and websites interact.

- Ability to design systems that are reliable, scalable, and fault tolerant.

- Good understanding of how to distribute and coordinate large-scale scraping workloads.

- Ability to identify the root cause of crawling failures and develop generalized, reusable solutions.

- Willingness to work across scraping, backend services, distributed processing, data pipelines, and infrastructure when required.

- Strong problem-solving skills and curiosity to understand how websites behave rather than relying solely on existing scraping tools.

Education

- Bachelor's degree in computer science, Information Technology, Engineering, or a related field is preferred.

- Equivalent practical experience in software engineering or web scraping will also be considered.

More jobs at AIMLEAP