SelectorwatchGuide

How to know your scraper broke before the data goes wrong

Updated 2026-10-10 ยท by Hieu Tran, written with AI agents and checked against the sources below

Most scrapers don't crash when a site changes. They keep running and return empty, partial or shifted data, and someone notices days later in a report. A few cheap checks turn that silent failure into an alert.

Free tool: Check which of your CSS selectors still match a page

Paste a URL and your selectors; see match counts and the first match, as a plain HTTP scraper sees the page. It doesn't get around bot protection.

The four ways scrapers fail

Checks to add to every scraper

  1. Count assertions: for each selector, the expected number of matches (exactly 1 title, at least 10 rows). Fail the run if it's 0.
  2. Field validation: each extracted value must parse (price as a number, date as a date). Alert on the share that fails, not on single values.
  3. Status and size checks: log the HTTP status and response size per page; a sudden drop in size often means a block or a JavaScript shell.
  4. Row-count trend: compare today's total with the last 7 runs; a drop of more than, say, 30% should page someone.
  5. A canary page: one known page whose values you know; check it first on every run.

When a selector breaks

Prefer selectors tied to meaning over layout: [itemprop=price], meta[property='og:title'], data-* attributes and structured data (JSON-LD) change less often than class names generated by a front-end build. Test any replacement against several pages before deploying it.

Stay on the right side of the site

Respect robots.txt and terms of use, rate-limit yourself, and don't bypass bot protection. If a site blocks automated access, ask for an API or a data feed instead.

Free tool: Check which of your CSS selectors still match a page

Paste a URL and your selectors; see match counts and the first match, as a plain HTTP scraper sees the page. It doesn't get around bot protection.