Scraper Developer

Brand:  Avolta
Country: 

IN

Location:  Bangalore Office
Job Type:  Indefinite

At Avolta (SIX: AVOL), our people are at the driving force behind our success. With a team of over 76,000 individuals representing more than 150 nationalities, we are a truly global company driven by passion, innovation, and excellence.

Born from the combination of Dufry and Autogrill, Avolta is redefining the travel experience through the dedication and expertise of our diverse workforce. Across 73 countries and 1,000 locations, our teams bring energy, creativity, and commitment to delivering world-class travel retail and food & beverage experiences.

We operate across multiple channels - including airports, motorways, cruise ships, ports, railways, and more - offering endless opportunities for collaboration and growth. Our people are empowered to make an impact, supported by a culture that values teamwork, development, and innovation.

Sustainability and social responsibility are embedded in our strategy, ensuring we grow in a way that benefits both our employees and the communities we serve.

Are you looking for a dynamic, international career where your contributions truly matter? Join Avolta and be part of a team that’s shaping the future of travel - together.

 

ROLE SUMMARY

The Scraper Developer role is focused on a single mission: implementing data extraction pipelines for new competitor websites at the rate required to meet Avolta's competitive intelligence coverage targets. Working within the framework designed by the Lead Scraping Engineer, you will onboard new competitor sites — one to three per week at steady state — writing the parsers, selectors and extraction logic to turn raw HTML and API responses into clean, structured product and pricing data.

This is a coding-heavy, execution-focused role. You will spend most of your day writing Python, reading browser DevTools network traces, and debugging extraction issues. The right candidate is a developer who genuinely enjoys the puzzle of getting data out of websites that don't want to give it to you — and who takes pride in writing clean, well-documented, maintainable code.

 

KEY RESPONSIBILITIES

Competitor Onboarding (Primary — 70% of time)

  • Receive a list of assigned competitor URLs and implement scrapers that extract: product name, brand, category, price (regular and promotional), currency, availability status, product URL and any other fields specified in the data schema.
  • Analyse each target site using Chrome DevTools before writing any code: identify whether data is served as static HTML, AJAX-loaded JSON, or via a private API endpoint — and choose the extraction approach accordingly.
  • Write HTML parsers using BeautifulSoup4 and lxml: construct robust CSS and XPath selectors, handle nested elements, extract text with normalisation (strip whitespace, handle encoding, remove special characters).
  • Use Scrapy as the primary scraping framework: implement Spider classes, define item schemas, configure per-spider settings (download delay, concurrent requests, retry policy).
  • For JavaScript-rendered sites, use Playwright or Selenium to automate browser interaction: wait for network idle, click through tabs or load-more buttons, extract from the rendered DOM.
  • Handle pagination systematically: detect pagination type (URL parameter, next-button link, infinite scroll, cursor-based API), implement accordingly, and verify total page count is extracted correctly.
  • Implement basic evasion measures within team guidelines: set appropriate User-Agent headers, add randomised delays, use proxy rotation via team middleware.

 

Data Quality & Validation (20% of time)

  • Write pytest unit tests for all parser functions: test against saved HTML fixtures, cover edge cases (missing field, null price, non-standard currency format).
  • Run scrapers in development and verify output against a visual check of the source website before handing off to QA.
  • Complete scraper documentation for every new implementation: target URL pattern, fields extracted, extraction method (HTML/API/headless), known issues, estimated run time and record count.
  • Respond to QA-flagged data quality issues: investigate root cause, fix parser logic, rerun to confirm resolution.

 

Maintenance Support (10% of time)

  • Fix scrapers assigned to you when they break due to site changes: update selectors, handle new page structures, adapt to changed API endpoints.
  • Escalate to Lead Scraping Engineer when breakage is caused by new anti-bot measures rather than structural changes.

 

TECHNICAL SKILLS — REQUIRED

▪  Python 3.7+ (solid working knowledge)

▪  XPath (basic to intermediate: axes, predicates, functions)

 

▪  BeautifulSoup4 (HTML parsing — primary tool)

▪  Chrome DevTools (Network tab, Elements tab, Console)

 

▪  lxml (XPath-based parsing, element tree traversal)

▪  requests library (manual HTTP requests, session handling)

 

▪  Scrapy (spider development, items, pipelines)

▪  JSON parsing (json module, nested structure navigation)

 

▪  Selenium WebDriver (dynamic page interaction)

▪  Regular expressions (re module — for text extraction/cleaning)

 

▪  Playwright for Python (headless browser automation)

▪  pytest (writing unit tests for parsers)

 

▪  CSS selectors (authoring, specificity, debugging)

▪  Git (clone, branch, commit, PR workflow)

 

 

TECHNICAL SKILLS — ADVANTAGE

  • Experience with anti-bot evasion basics: setting realistic headers, cookie handling, referrer management.
  • Rotating proxy usage: how to configure a scraper to route through a proxy pool.
  • Basic JavaScript: reading page source JS to find embedded JSON data objects (e.g. window.__INITIAL_DATA__).
  • CAPTCHA solving basics: integrating a third-party solver API.
  • Experience scraping e-commerce, travel, retail or comparison shopping sites.
  • SQL (basic SELECT queries for verifying extracted data).
  • Docker basics: running a scraper inside a container.

 

EXPERIENCE & QUALIFICATIONS

  • 1-3 years of professional software engineering or data engineering experience.
  • Minimum 6 months of hands-on web scraping experience — this must be real, demonstrable work (not a single tutorial project). You should be able to show examples during interview.
  • Bachelor's degree in Computer Science, Information Technology or equivalent; strong practical portfolio considered.
  • GitHub profile with scraping examples is strongly preferred.
  • Methodical, detail-oriented approach to work — data accuracy matters as much as code execution.

 

Avolta Logo

 

 

 

 

 

Due to certain email system settings, some of our messages may occasionally land in your junk or spam folder. To ensure you don’t miss any important updates regarding your application, please check these folders regularly and mark our emails as ‘Not Spam’ if needed.

We look forward to connecting with you soon!