Scrapper developers

Brand:  Avolta
Country: 

IN

Location:  Bangalore Office
Job Type:  Definite

At Avolta (SIX: AVOL), our people are at the driving force behind our success. With a team of over 76,000 individuals representing more than 150 nationalities, we are a truly global company driven by passion, innovation, and excellence.

Born from the combination of Dufry and Autogrill, Avolta is redefining the travel experience through the dedication and expertise of our diverse workforce. Across 73 countries and 1,000 locations, our teams bring energy, creativity, and commitment to delivering world-class travel retail and food & beverage experiences.

We operate across multiple channels - including airports, motorways, cruise ships, ports, railways, and more - offering endless opportunities for collaboration and growth. Our people are empowered to make an impact, supported by a culture that values teamwork, development, and innovation.

Sustainability and social responsibility are embedded in our strategy, ensuring we grow in a way that benefits both our employees and the communities we serve.

Are you looking for a dynamic, international career where your contributions truly matter? Join Avolta and be part of a team that’s shaping the future of travel - together.

 

KEY RESPONSIBILITIES
Web Scraping & Competitor Onboarding - 65%
• Analyse assigned websites using Chrome DevTools (Network, Elements and Console) to determine whether data is
best extracted from HTML, XHR/Fetch requests, REST/JSON APIs, embedded JSON, or browser-rendered content.
• Develop production-grade spiders using the Scrapy framework, including Spiders, Requests/Responses, Items,
Item Loaders, Pipelines, Middlewares and project/settings configuration.
• Use Selenium WebDriver or Playwright for Python when required data cannot reliably be obtained through Scrapy
or direct HTTP requests.
• Handle JavaScript-rendered content and browser interactions such as dropdowns, pop-ups, location/airport
selection, load-more buttons and infinite scrolling.
• Implement and maintain location-specific crawls where product assortment, pricing, promotions, availability or
content varies by country, region, airport, store or selected location. Handle location selection through URLs, APIs,
cookies, sessions or browser interactions as required.
• Handle product variants accurately, including size, volume, pack size, colour or other product options, ensuring
relevant variants and their associated prices, availability, identifiers and attributes are extracted and mapped
correctly.
• Build robust XPath and CSS selectors capable of handling nested and changing HTML structures.
• Implement numbered, next-page, load-more, infinite-scroll and API/cursor-based pagination.
• Extract structured product information such as product name, brand, category, regular/promotional price, currency,
availability, product URL, image URL, SKU/product identifiers and other required attributes.
• Work with requests/httpx and direct APIs where browser automation is unnecessary; parse nested JSON and
embedded JavaScript/JSON objects.
• Implement request/session management including headers, cookies, retries, timeouts, redirects, throttling and proxy
configuration.
• Develop reusable scraping components and avoid duplicated website-specific logic wherever practical.
Data Quality, Testing & Validation - 20%
• Validate scraper output against the source website and defined data schema.
• Implement field validation, normalization and cleaning for prices, currencies, identifiers, availability and text fields.
• Handle missing fields, malformed responses, duplicate records and unexpected page structures gracefully.
• Write unit tests using pytest for parsers, extraction logic and reusable components.
• Perform record-count, pagination, location and variant validation to ensure complete and accurate extraction.
• Investigate QA-reported discrepancies, identify root causes and implement fixes with appropriate logging and error
handling.
Maintenance & Production Support - 15%
• Troubleshoot and repair scrapers affected by website structure, frontend or API changes.
• Diagnose HTTP errors including 403, 429, redirects, session failures and rate limiting.
• Identify blocking/anti-bot behaviour and implement approved mitigation approaches such as realistic headers,
cookies, session handling, throttling and proxy rotation.
• Debug browser/network requests to identify API and frontend behaviour changes.
• Maintain clean, reusable and well-documented code; participate in peer reviews and follow established Git
workflows.

TECHNICAL SKILLS - REQUIRED
Python - Strong hands-on experience
• Python programming, OOP, exception handling, modules/packages and reusable code design.
• JSON/XML processing, regular expressions and data cleaning/normalization.
Scrapy Framework - Strong hands-on experience
• Spiders, Requests/Responses and callbacks; Items / Item Loaders; Pipelines.
• Downloader/Spider Middleware, settings and per-spider configuration.
• Retries, error handling, concurrency, AutoThrottle/download delays, cookies, sessions, proxies, logging and
debugging.
Selenium & Playwright - Hands-on experience
• Browser automation for dynamic and JavaScript-heavy websites.
• Element identification and interaction, wait strategies, dropdowns, pop-ups and dynamic elements.
• Load-more/infinite scroll, cookie/session handling, headless execution and browser automation troubleshooting.
Web & API Extraction
• Strong XPath and CSS selectors; HTML/XML parsing using lxml and/or BeautifulSoup.
• Chrome DevTools, XHR/Fetch and REST/JSON API analysis, nested JSON and HTTP fundamentals.
• Pagination, API/cursor pagination, headers, cookies, sessions, status codes and redirects.
Location & Variant Handling
• Experience extracting location-dependent product, pricing and availability data.
• Ability to identify how websites maintain location context through APIs, URLs, cookies, sessions or browser
selections.
• Experience extracting products with multiple variants and correctly mapping variant-level attributes, identifiers, prices
and availability.
Engineering Practices
• Git branching, commits, pull requests and code reviews.
• pytest/unit testing, logging, exception handling and debugging.
• Basic Linux command-line skills and ability to independently analyse unfamiliar websites.
TECHNICAL SKILLS - ADVANTAGE
• Anti-bot/blocking analysis and mitigation; proxy pools and rotating proxies.
• CAPTCHA identification and third-party CAPTCHA-solving integrations.

 

• Scrapy custom downloader middleware, signals and extensions.
• Browser/network interception using Playwright.
• Docker and containerised scraper execution.
• SQL for validating and analysing extracted datasets.
• CI/CD pipelines, cron/scheduled crawler execution, cloud environments and distributed scraping.
• Monitoring and alerting for production crawlers.
• Experience scraping e-commerce, retail, travel, marketplaces or comparison-shopping websites.
• Familiarity with downstream data pipelines and data warehouses.
EXPERIENCE & QUALIFICATIONS
• 4-6 years of professional software/data engineering experience, with significant hands-on web scraping experience.
• Strong practical experience developing production scrapers using Python and Scrapy.
• Demonstrable experience building and maintaining multiple production web crawlers/spiders.
• Experience extracting data from a combination of static websites, dynamic websites and APIs.
• Experience troubleshooting scraper failures caused by website, frontend or API changes.
• Experience working with Git-based collaborative development.
• Bachelor's degree in Computer Science, Information Technology, Engineering or equivalent practical experience.
• GitHub/portfolio or examples of previous scraping projects are advantageous where available.
• Strong analytical and debugging skills with a focus on data accuracy, maintainability and reliability.

 

Avolta Logo

 

 

 

 

 

Due to certain email system settings, some of our messages may occasionally land in your junk or spam folder. To ensure you don’t miss any important updates regarding your application, please check these folders regularly and mark our emails as ‘Not Spam’ if needed.

We look forward to connecting with you soon!