Introducing dpkit Terminal and TypeScriptSee the announcement

Beyond the build

Data platforms & pipelines

Getting data in, checking it, and keeping it correct — from a handful of spreadsheets to millions of rows a day.

starts fromdatist · pricing
€5,000
fixed price from
2 weeks
delivery from
Both agreed in writing before anyone starts

What it is

Everything between data arriving and data you can trust: collecting it, giving it a shape, checking it against that shape, and keeping it correct as it changes — whether it comes from a handful of spreadsheets or a few million rows a day.

What you get

Pipelines that fail loudly instead of quietly writing something wrong. A written schema your data is checked against on every run, so a bad file is caught at the door rather than by the person who consumes it a month later. And formats that outlive us — open, documented, and readable without our software.

How we build it

Columnar processing, so the cost of a million rows is closer to the cost of a thousand than you would expect, and a schema that travels with the data rather than living in a document beside it. This is our oldest line of work: we author the Fairspec exchange format, we wrote the TypeScript implementation of the Data Package standard under an NLnet grant, and the Python framework behind it is downloaded around 700,000 times a month.

Where we have built it

captureapplication.fairspec.org
The Fairspec site, the data exchange format we author, with its Python and TypeScript implementations.

Fairspecdata framework in TypeScript and Python

A specification for describing tabular datasets, with reference implementations in TypeScript and Python.

we author

capturemycarro.app
The Mycarro car search engine as it runs today, searching across its live Portuguese listings.

Mycarrocar marketplace and dealer CRM

A used-car marketplace live in Portugal. Consumer search over listings gathered from across the market, a dealer back office for leads, inventory and invoicing, and a mobile app — built end to end.

we built

What we build it on

  • Columnar dataframes instead of ad-hoc scripts, so a million rows costs what a thousand does.

  • Schemas that travel with the data, and one validation contract from database to browser.

  • One Postgres per product, with the schema and the queries deliberately owned by different tools.

  • Polite, resumable crawls that fail loudly instead of quietly going empty.

Get in touch

Tell us what you need built.

We are a small team in Portugal, working London hours. A site, an app, an internal tool, or the platform behind them — start with a sentence about the problem and we will take it from there.

All services
Get a quote