The Lab

Two systems that run without me. They enter competitions, form hypotheses, spend real money on compute, and decide for themselves what to keep — inside budgets they cannot exceed.

The numbers below are read from each system's own repository, not typed in here. They are the operational kind on purpose: how much a run costs, how much of a finite budget is left, how often a promising result failed to survive an honest test. That is the part of this work that transfers to a production data platform.

Auto-Kaggle

An unattended loop that enters competitions, trains, learns from its own results, and keeps climbing — with no human in the path.

Four independent planes

Code, control, agents and compute fail separately; the whole system is recoverable from git.

Budget-guarded compute

Spend runs on a twice-weekly clock with a daily budget guard, not an open tap.

Live figures for this system aren't reachable right now, so the design is shown instead of the numbers.

EdwardPlata/auto-kaggle

Auto-Numerai

A research engine that proposes hypotheses, tests them against a finite selection budget, and refuses to promote a model that only looks good on the search split.

A finite selection budget

Reads of the holdout split are capped and counted, so the loop cannot quietly overfit its own scoreboard.

Claimed against realised

Every promotion records the search score next to the holdout score — the check that catches manufactured progress.

Live figures for this system aren't reachable right now, so the design is shown instead of the numbers.

Numerai research engine

Why this is on a consulting site

An unattended system is an honesty test. Anything that runs on a schedule with money attached will eventually find the gap between what you measured and what you claimed — and if nothing is watching for that gap, it will quietly widen.

So both systems are built around the same discipline a data platform needs: a budget that stops the process rather than a person who remembers to, a record of what was predicted against what actually happened, and a failure mode that halts instead of guessing. The competition scores are incidental. The machinery is the portfolio.

Need this kind of rigour on your stack?

Available for contract data-engineering work, and open to forward-deployed roles.