Projects

Two products built under DataBlock, in the open. Status here is read straight from each project's repository, so it is what is true today rather than what was true when this page was written.

DataBolt

A public benchmark for data engineers.

Visit
A reproducible pipeline benchmark on an open dataset, scored on dollars per terabyte and terabytes per hour. Correctness is a gate, not a score — the competitive surface is which architecture you chose.

Data-Citizen

An agentic overlay on NYC Open Data.

Visit
A normalised catalog of NYC open data with five-dimension quality scores and a dataset relationship graph, persisted so another application can query the layer instead of re-solving ingestion against raw Socrata endpoints.

Working on something similar?

These are the same problems that show up in client work — pipeline cost, data quality, and platforms that have to survive an audit.