Skip to main content
Web scraping · Data extraction

Web scraping that keeps working when the site changes

We extract data from websites that do not give you an API: event listings, catalogs, prices, classifieds, official bulletins. Each source gets the strategy it needs, a test that calls it for real every few days and an alert the day it is redesigned, before anyone notices missing data.

Web scraping is the automated extraction of data published on websites that offer neither an API nor a download file. At Soamee we build custom scrapers in Node.js and TypeScript: for each source we choose between its internal API, its RSS feed, the HTML the server returns or a headless browser with Playwright, and every result goes through the same normalizer, so it reaches your database in a common format and without duplicates. When a site blocks automated requests (403, CAPTCHAs, IP bans) we integrate Bright Data: Web Unlocker, Scraping Browser or residential proxies. Every source has a test against the live site that runs every few days and an email alert with a diagnosis when it fails.

What we do

From the source inventory to clean data in your database

A scraper that works today takes an afternoon to write. The work is keeping it running six months later, with thirty sources that each change at their own pace.

Inventory and feasibility per source

Before writing code we review each site: whether it has a public or internal API, whether it is a WordPress with /wp-json open, whether it loads data via XHR, what its robots.txt says and whether it blocks automated requests. The output is one sheet per source with strategy, risks and frequency.

One strategy per source type

API, RSS, static HTML or headless browser behind a single interface. Adding a source usually means writing its configuration (URL, selectors, field mapping) and its test, without coding a scraper from scratch.

Normalization and deduplication

Spanish dates like «del 10 al 28 de junio», prices written as text, embedded HTML, the same category under five names. Everything goes through a normalizer and is stored with a unique key per source and external ID, so re-running updates instead of duplicating.

Sites with anti-bot protection

For sources that return 403 or a CAPTCHA we integrate Bright Data: Web Unlocker for HTTP requests, Scraping Browser for pages that need JavaScript and residential proxies when the problem is the IP. Only on those sources; the rest go out directly and use no unblocking.

Queues, retries and scheduling

Each source is a job in a BullMQ queue on Redis, with its own frequency, three retries with exponential backoff and a log per run: how many items it found, how many were new and how long it took.

Smoke tests and alerts

One test per source that calls the live site every three days in CI, plus unit tests against a saved snapshot of the response. If a run fails, an email arrives with the HTTP status, a snippet of the response and what to check first.

How we choose

Four strategies, from most stable to most fragile

We use the first one on the list that the source allows. Each step down costs speed and stability, so we only go down when the source forces us to.

StrategyWhen we use itCostWhat breaks
Public or internal APIThe site loads its data from JSON: open data portals, WordPress with the REST API, event plugins, endpoints visible in the Network tab~200 ms per requestSchema changes, which are rare
RSS / AtomThe source publishes a feed and title, link and summary are enough~200 ms per feedThe feed date is the publication date, not the date of the event or the offer
Static HTMLThe content comes in the HTML the server returnsOne request per listing, plus one per detail page if a field is missingThe CSS selectors, with every redesign
Headless browserContent rendered with JavaScript, infinite scroll, “load more” buttons or blocking of HTTP clients3-5 s per page and 100-200 MB of RAM per browserWait times and the browser binary in each environment
In production

What we have running

Figures from an aggregation platform we maintain, with public sources that have little in common: open data portals, WordPress, Drupal, React sites and hand-built pages.

35+

sources, each with its own smoke test against the live site

4

extraction strategies behind a single interface

3 days

between each smoke test run in CI

3

retries with exponential backoff before a run is marked as failed

The most common failure is a source redesign. The smoke test catches it within three days at most. For sites that block automated requests we work with Bright Data.

Bright Data integration →
How we work

From a list of URLs to the first record in your database

First one source of each type, so the real problems show up early. Then the rest in batches.

01

Sheet per source

We review each site, choose a strategy and note the odd parts: date formats, pagination, blocking, fields that only exist on the detail page.

02

Pilot per strategy

We implement one source of each type with its unit test, smoke test and document. That is where the problems the inventory misses show up.

03

Remaining sources

With the model validated, each new source is usually configuration plus a test. The ones that need a browser or unblocking go in a separate batch.

04

Operation

Scheduling, alerts, run reports and maintenance when a source changes. If you prefer to run it with your own team, we hand it over documented source by source.

FAQ

Web scraping FAQ

Is web scraping legal in Spain and the EU? +

Extracting data published without access restrictions is common practice, but there are three limits we check for every source: the site’s terms of use, the EU sui generis database right (you cannot extract a substantial part of someone else’s database to reuse it) and the GDPR when personal data is involved. We do not access areas behind a login without the owner’s permission. If the case is unclear, we recommend a legal review before starting.

What happens when the site is redesigned? +

That source’s smoke test fails on the next run, within three days at most, and an alert arrives with the error. Redesigns almost never affect API or RSS sources. For HTML or headless browser sources, updating the selectors in the source configuration is usually enough, with no code deploy.

When do you need Bright Data? +

When the site returns 403 to any automated client, shows CAPTCHAs or blocks by IP. A self-hosted headless browser is not enough there, because the block depends on the IP and the browser fingerprint. We enable it only on the sources that need it; the rest still go out directly.

How often is the data updated? +

Each source has its own schedule: hourly for prices or availability, once a day for event listings or catalogs that change little. It is also a matter of courtesy to the source site: requesting every five minutes something that changes once a day only adds load.

In what format do I receive the data? +

In your database, usually PostgreSQL, with a common schema for all sources, or through a REST API. We can also generate periodic CSV or JSON exports or push changes via webhook.

Can I use the data to feed an AI model? +

Yes, and it is one of the most common uses: normalized data can be the base for semantic search, a RAG system or an agent. The same legal limits apply as for any other use, and we store the source URL and date of every record so the source can be cited.

Let’s get started

Send us your list of sources

Send us the URLs you need data from and which fields you care about. We will return a sheet per source with the strategy, risks and frequency we recommend.

Tell us which sources you need
Web scraping

Tell us your challenge. We'll propose a solution.

No commitment. Within 24 hours, you'll receive a proposal with scope, timeline and budget. No fine print.

Book a free call →