Web scraping that keeps working when the site changes
We extract data from websites that do not give you an API: event listings, catalogs, prices, classifieds, official bulletins. Each source gets the strategy it needs, a test that calls it for real every few days and an alert the day it is redesigned, before anyone notices missing data.
Web scraping is the automated extraction of data published on websites that offer neither an API nor a download file. At Soamee we build custom scrapers in Node.js and TypeScript: for each source we choose between its internal API, its RSS feed, the HTML the server returns or a headless browser with Playwright, and every result goes through the same normalizer, so it reaches your database in a common format and without duplicates. When a site blocks automated requests (403, CAPTCHAs, IP bans) we integrate Bright Data: Web Unlocker, Scraping Browser or residential proxies. Every source has a test against the live site that runs every few days and an email alert with a diagnosis when it fails.
From the source inventory to clean data in your database
A scraper that works today takes an afternoon to write. The work is keeping it running six months later, with thirty sources that each change at their own pace.
Inventory and feasibility per source
Before writing code we review each site: whether it has a public or internal API, whether it is a WordPress with /wp-json open, whether it loads data via XHR, what its robots.txt says and whether it blocks automated requests. The output is one sheet per source with strategy, risks and frequency.
One strategy per source type
API, RSS, static HTML or headless browser behind a single interface. Adding a source usually means writing its configuration (URL, selectors, field mapping) and its test, without coding a scraper from scratch.
Normalization and deduplication
Spanish dates like «del 10 al 28 de junio», prices written as text, embedded HTML, the same category under five names. Everything goes through a normalizer and is stored with a unique key per source and external ID, so re-running updates instead of duplicating.
Sites with anti-bot protection
For sources that return 403 or a CAPTCHA we integrate Bright Data: Web Unlocker for HTTP requests, Scraping Browser for pages that need JavaScript and residential proxies when the problem is the IP. Only on those sources; the rest go out directly and use no unblocking.
Queues, retries and scheduling
Each source is a job in a BullMQ queue on Redis, with its own frequency, three retries with exponential backoff and a log per run: how many items it found, how many were new and how long it took.
Smoke tests and alerts
One test per source that calls the live site every three days in CI, plus unit tests against a saved snapshot of the response. If a run fails, an email arrives with the HTTP status, a snippet of the response and what to check first.
Four strategies, from most stable to most fragile
We use the first one on the list that the source allows. Each step down costs speed and stability, so we only go down when the source forces us to.
| Strategy | When we use it | Cost | What breaks |
|---|---|---|---|
| Public or internal API | The site loads its data from JSON: open data portals, WordPress with the REST API, event plugins, endpoints visible in the Network tab | ~200 ms per request | Schema changes, which are rare |
| RSS / Atom | The source publishes a feed and title, link and summary are enough | ~200 ms per feed | The feed date is the publication date, not the date of the event or the offer |
| Static HTML | The content comes in the HTML the server returns | One request per listing, plus one per detail page if a field is missing | The CSS selectors, with every redesign |
| Headless browser | Content rendered with JavaScript, infinite scroll, “load more” buttons or blocking of HTTP clients | 3-5 s per page and 100-200 MB of RAM per browser | Wait times and the browser binary in each environment |
What we have running
Figures from an aggregation platform we maintain, with public sources that have little in common: open data portals, WordPress, Drupal, React sites and hand-built pages.
sources, each with its own smoke test against the live site
extraction strategies behind a single interface
between each smoke test run in CI
retries with exponential backoff before a run is marked as failed
The most common failure is a source redesign. The smoke test catches it within three days at most. For sites that block automated requests we work with Bright Data.
Bright Data integration →From a list of URLs to the first record in your database
First one source of each type, so the real problems show up early. Then the rest in batches.
Sheet per source
We review each site, choose a strategy and note the odd parts: date formats, pagination, blocking, fields that only exist on the detail page.
Pilot per strategy
We implement one source of each type with its unit test, smoke test and document. That is where the problems the inventory misses show up.
Remaining sources
With the model validated, each new source is usually configuration plus a test. The ones that need a browser or unblocking go in a separate batch.
Operation
Scheduling, alerts, run reports and maintenance when a source changes. If you prefer to run it with your own team, we hand it over documented source by source.
Web scraping FAQ
Is web scraping legal in Spain and the EU? +
Extracting data published without access restrictions is common practice, but there are three limits we check for every source: the site’s terms of use, the EU sui generis database right (you cannot extract a substantial part of someone else’s database to reuse it) and the GDPR when personal data is involved. We do not access areas behind a login without the owner’s permission. If the case is unclear, we recommend a legal review before starting.
What happens when the site is redesigned? +
That source’s smoke test fails on the next run, within three days at most, and an alert arrives with the error. Redesigns almost never affect API or RSS sources. For HTML or headless browser sources, updating the selectors in the source configuration is usually enough, with no code deploy.
When do you need Bright Data? +
When the site returns 403 to any automated client, shows CAPTCHAs or blocks by IP. A self-hosted headless browser is not enough there, because the block depends on the IP and the browser fingerprint. We enable it only on the sources that need it; the rest still go out directly.
How often is the data updated? +
Each source has its own schedule: hourly for prices or availability, once a day for event listings or catalogs that change little. It is also a matter of courtesy to the source site: requesting every five minutes something that changes once a day only adds load.
In what format do I receive the data? +
In your database, usually PostgreSQL, with a common schema for all sources, or through a REST API. We can also generate periodic CSV or JSON exports or push changes via webhook.
Can I use the data to feed an AI model? +
Yes, and it is one of the most common uses: normalized data can be the base for semantic search, a RAG system or an agent. The same legal limits apply as for any other use, and we store the source URL and date of every record so the source can be cited.
You may also be interested in
Send us your list of sources
Send us the URLs you need data from and which fields you care about. We will return a sheet per source with the strategy, risks and frequency we recommend.
Tell us which sources you needTell us your challenge. We'll propose a solution.
No commitment. Within 24 hours, you'll receive a proposal with scope, timeline and budget. No fine print.