Python Web Scraping Engineer-4-6 Years
<p><strong>Python & JavaScript Developer – AI & Web Scraping</strong></p><p><br></p><p><strong>Experience:</strong> 4-7 Years </p><p><strong>Location:</strong> Remote </p><p><strong>Mode of Engagement:</strong> Full-time </p><p><strong>No of Positions:</strong> 3 </p><p><strong>Educational Qualifications:</strong> Bachelor's degree in computer science, Information Technology </p><p><strong>Industry:</strong> IT / Software Development </p><p><strong>Notice Period:</strong> Immediate Joiners Preferred</p><p><br></p><p><strong>About the Role</strong> </p><p><br></p><p>We are looking for a <strong>hands-on Web Scraping / Crawling Engineer</strong> with <strong>3–5 years of experience</strong> in web scraping, browser automation, and scalable data extraction. </p><p><br></p><p>The ideal candidate should have strong experience working with <strong>dynamic and JavaScript-heavy websites</strong>, building reliable crawling workflows, handling crawling failures, and working with distributed processing systems. </p><p><br></p><p>This is a highly technical, hands-on role. You will be expected to <strong>design, develop, debug, optimize, and maintain web crawling and data extraction systems</strong>. </p><p><br></p><p><strong>Responsibilities</strong> </p><p><br></p><ul><li>Design, develop, and maintain <strong>scalable web crawling and scraping systems</strong> for dynamic and JavaScript-heavy websites. </li><li>Develop browser automation workflows using <strong>Playwright, Selenium, Puppeteer, or similar frameworks</strong>. </li><li>Investigate and resolve crawling issues such as <strong>403/429 responses, redirects, timeouts, rendering failures, session issues, and anti-bot challenges</strong>. </li><li>Develop robust <strong>retry, fallback, and failure-handling mechanisms</strong> to improve crawler reliability. </li><li>Work with <strong>cookies, sessions, browser contexts, headers, proxies, and related crawling mechanisms</strong> to maintain state and improve crawling reliability. </li><li>Build and optimize <strong>concurrent and distributed scraping workflows</strong> using asynchronous processing, queues, and worker-based architectures. </li><li>Design data pipelines covering <strong>URL processing, crawling, extraction, validation, transformation, and storage</strong>. </li><li>Optimize crawler performance, including <strong>concurrency, browser lifecycle, resource utilization, timeouts, and request handling</strong>. </li><li>Implement monitoring and observability for <strong>crawl success rates, failure types, latency, retries, and worker performance</strong>. </li><li>Debug complex crawling problems and identify <strong>root causes rather than relying only on one-off fixes</strong>. </li><li>Develop reusable crawling components and frameworks that can support multiple websites and use cases. </li><li>Collaborate with data engineering, AI, backend, and product teams to deliver reliable and structured web data. </li><li>Evaluate and adopt new <strong>web crawling, browser automation, and data extraction technologies</strong> where appropriate. </li></ul><p><br></p><p><strong>Required Skills </strong></p><p><br></p><ul><li><strong>3–5 years of hands-on experience</strong> in web scraping, web crawling, browser automation, or web data engineering. </li><li>Strong programming experience in <strong>Python</strong>. JavaScript/Node.js is a plus. </li><li>Strong hands-on experience with Scrapy, <strong>Playwright, Selenium, Puppeteer, or similar browser automation frameworks</strong>. </li><li>Strong understanding of <strong>HTTP, HTML, DOM, JavaScript rendering, redirects, cookies, sessions, headers, and browser contexts</strong>. </li><li>Experience with <strong>browser fingerprinting, WAFs, and modern anti-bot mechanisms</strong>. </li><li>Experience working with <strong>dynamic, JavaScript-heavy, and asynchronous websites</strong>. </li><li>Good understanding of <strong>asynchronous programming, concurrency, and parallel processing</strong>. </li><li>Experience with <strong>job queues, distributed workers, or message-based processing systems</strong> such as Redis, RabbitMQ, Kafka, Celery, or similar technologies. </li><li>Understanding of <strong>retry mechanisms, error handling, idempotency, rate limiting, and backpressure</strong> in distributed systems. </li><li>Hands-on experience with <strong>proxy management, session handling, and anti-bot challenges</strong>. </li><li>Strong debugging and analytical skills, with the ability to investigate and resolve complex crawling failures. </li><li>Experience working with <strong>SQL databases</strong> such as PostgreSQL or MySQL. </li><li>Familiarity with <strong>Docker and cloud platforms</strong> such as AWS, GCP, or Azure is a plus. </li><li>Strong <strong>hands-on engineering mindset</strong> rather than purely project or delivery management experience. </li><li>Ability to independently <strong>debug, investigate, and solve difficult crawling problems</strong>. </li><li>Strong understanding of how browsers, HTTP requests, sessions, and websites interact. </li><li>Ability to design systems that are <strong>reliable, scalable, and fault tolerant</strong>. </li><li>Good understanding of how to distribute and coordinate large-scale scraping workloads. </li><li>Ability to identify the root cause of crawling failures and develop <strong>generalized, reusable solutions</strong>. </li><li>Willingness to work across <strong>scraping, backend services, distributed processing, data pipelines, and infrastructure</strong> when required. </li><li>Strong problem-solving skills and curiosity to understand how websites behave rather than relying solely on existing scraping tools. </li></ul><p><br></p><p><strong>Education</strong> </p><p><br></p><ul><li>Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field is preferred. </li><li>Equivalent practical experience in software engineering or web scraping will also be considered. </li></ul>