Design a Selenium web scraper to compare travel sites
Problem Design and implement a Selenium-based scraper that compares two travel sites (e.g. Expedia and Booking.com): run the same search on both, extract results and images, and print a side-by-side comparison.
Requirements
- Drive both sites: navigate, execute a search (destination, dates), harvest result cards (name, price, rating, image URLs); normalize and align listings across sites; output a comparison table.
Core design
- PAGE OBJECT MODEL — the pattern this question fishes for: per-site page classes (SearchPage, ResultsPage) encapsulating selectors and actions behind a common interface (TravelSiteScraper with search() and getResults()), so site-specific DOM details never leak into comparison logic; a Result DTO (site, name, price, currency, rating, imageUrl) as the normalized unit.
- Comparator service matching listings across sites (by normalized hotel name / fuzzy match) and rendering the side-by-side output.
- Robustness: explicit waits (WebDriverWait on expected conditions — never sleep()), retries on stale elements, headless mode, graceful per-listing failure (skip, don't crash the run).
Discussion points
- Dynamic content: lazy-loaded images and infinite scroll (scroll-and-wait loops), iframes, pop-ups/cookie banners as the first hurdle.
- Fragile selectors: prefer data-* attributes over brittle XPath chains; centralizing selectors in the page objects is the maintainability argument.
- Ethics/legality guardrails worth voicing: robots.txt, rate limiting, ToS; and when an official API beats scraping.
asked …