Introducing dpkit Terminal and TypeScriptSee the announcement

Data & storage

Web data acquisition

Polite, resumable crawls that fail loudly instead of quietly going empty.

  • Crawlee
  • Playwright
  • Cheerio
  • Firecrawl

What it is

Collecting data from sites and services that offer no usable API — on a schedule, at volume, and without getting the client blocked.

Why we chose it

Crawlee handles the parts that are tedious to get right: request queues, retries, proxy rotation, concurrency limits and polite rate limiting. We reach for Cheerio when the markup is served whole and Playwright only when a page genuinely needs a browser, because a headless browser per page is the difference between a crawl that costs cents and one that costs hundreds. The rule we hold to is that a crawler must fail loudly: a selector that stops matching raises an error instead of writing an empty result over good data.

What we build with it

Data platforms & pipelines

Getting data in, checking it, and keeping it correct — from a handful of spreadsheets to millions of rows a day.

Data processing · Data standards & validation · Transactional databases · Web data acquisition

Where we use it

Behind marketplace and search products, where inventory has to be assembled from many sources before it can be normalised into one catalogue.
capturemycarro.app
The Mycarro car search engine as it runs today, searching across its live Portuguese listings.

Mycarrocar marketplace and dealer CRM

A used-car marketplace live in Portugal. Consumer search over listings gathered from across the market, a dealer back office for leads, inventory and invoicing, and a mobile app — built end to end.

we built

Get in touch

Tell us what you need built.

We are a small team in Portugal, working London hours. A site, an app, an internal tool, or the platform behind them — start with a sentence about the problem and we will take it from there.

All technologies
Get a quote