How BitBank Improves Itself: Monitoring Agents and Autoresearch
2026-09-23
BitBank is run by a very small team. It serves live crypto forecasts, a trading game, a token launchpad and on-chain analytics. Two groups of AI agents keep it working and make it better. Monitoring agents watch production and repair what breaks. Autoresearch agents run experiments on our forecasting and trading models and ship a change only when the evidence is strong. This post explains how both loops work, what they got right this week, and where they still need a human.
Loop 1: the production repair agent
Every five minutes a supervisor collects three kinds of signal: browser errors reported by our pages, errors from the Go backend and the forecasting sidecar's system journal, and health checks on the site and its session API. When it finds new problems, it starts a coding agent in the repository and gives it a bounded repair task.
- Collect and filter. Most error reports on the public internet are not bugs. SQL-injection and XSS probes, browser extensions, analytics beacons and scripts injected by in-app browsers all end up in the error log. The supervisor drops these before any agent sees them.
- Fingerprint. One broken route hit by a fuzzer can produce hundreds of distinct URLs. Issues are grouped by route, method and status, so one fault is one issue.
- Triage before repair. The agent first searches our incident history, then checks whether the failure still happens, locally and over public HTTPS. Only then does it change code.
- Fix, test, deploy, retest. A minimal fix, a focused regression test, the documented deploy procedure, then a retest of the original failure.
- Per-issue verdicts. The agent closes each issue on its own line, as
FIXEDwith evidence orHARMLESSwith a reason. Anything it cannot close is escalated to a human by email.
What went wrong, and how we fixed the monitor itself
The first version was too strict in a way that made it useless. An incident closed only if the agent proved every issue in the batch fixed. One unprovable item was enough to hold the whole batch open: a script error injected by the Twitter in-app browser on an old iPhone, referencing variables that do not exist in our code. The same 40 issues came back every six hours, each run ended in an email, and the queue grew to 153 entries.
| Change | Effect |
|---|---|
| Noise filter for probes, extensions, beacons, in-app webviews and SQL echo lines | 153 pending issues fell to 19 |
| Route-level fingerprints | A fuzzed endpoint counts as one issue, not 20 |
| Per-issue verdicts instead of all-or-nothing | Fixed items close, and only the rest escalate |
| Transient errors must recur 3 times | One upstream timeout no longer starts a repair |
| Go panics carry their stack trace | The agent can find the faulting line |
| Freshest issues first, and a 7-day expiry | Agent time goes to current problems |
The first run after these changes received three issues, found prior evidence for each, retested them,
and closed all three without sending an email. The same day it caught a real browser bug. Some early
Chromium builds on Android, reported from Honor phones, reject the array form of the canvas
roundRect radius argument, so the homepage chart crashed for those users. The chart now
falls back to a numeric radius and then to a plain rectangle.
Loop 2: autoresearch on the forecasting stack
Our price forecasts come from Chronos-2, a time-series foundation model, served from a GPU sidecar with custom CUDA kernels. Autoresearch agents get research contracts: a fixed incumbent, a data boundary, a metric, and a promotion gate. They write harnesses, run walk-forward experiments, and report what survives.
This week's run found two real defects in how we used the model:
- A precision bug. One step in our accelerated inference path converted the final price forecasts back to bfloat16. At a Bitcoin price near $86,000, bfloat16 values are 512 dollars apart. Every forecast snapped to that grid, an error of about 0.4% for every asset. Output now stays in float32, and the accelerated model agrees with the reference to within about 0.01%.
- A compounding path. The forecast line was built by chaining per-step open-to-close predictions, which compounded the noise in each step. It is now built from the median close forecast, damped toward the last price. The damping strength came from 5-fold cross-validation, and every fold chose a value between 0.05 and 0.10.
| Forecast line | MAE before | MAE after | Change | Folds improved |
|---|---|---|---|---|
| Chart line | 3.22% | 1.79% | -44% | 45 / 45 |
| Homepage and snapshots | 3.06% | 2.11% | -31% | 45 / 45 |
| Flat line (no-change baseline) | 1.79% |
This was measured end to end over HTTP against shadow copies of the live sidecar: 9 trading pairs, 5 chronological folds, January 2024 to September 2026. The change has a one-line rollback.
What the agents told us that we did not want to hear
Look at the last row of that table. After the fix, the chart line is only about as accurate as a flat line that assumes the price will not move. On hourly crypto data, the raw model's median forecast was 4 to 7% worse than that flat baseline for horizons of 1 to 24 hours, across 57 pairs, and no context length fixed it. The best stacked model, gradient-boosted trees on lags, volatility and Chronos quantiles, beat the flat line by about 0.16%. Short-horizon crypto prices are close to a random walk, and we would rather publish that than hide it. See our accuracy page for live numbers.
The autoresearch agents also searched variations of our live rotation strategy: regime gates, other horizons, and Chronos features. Removing the trailing stop raised average fold returns, but it won only 20 to 23 of 38 paired folds, with 40 to 50% drawdowns, which fails our promotion gate. Nothing was promoted. The search also showed that the current profile's headline backtest is a fragile peak: most single-parameter changes land well below it. A result like this counts as a success, because it stops a bad change from reaching production.
Why two loops, not one
The two loops are built alike. Each has a narrow contract, a verifier the agent does not control, and a record of past incidents or experiments to read before starting. Each can end in "no change" or "escalate". The monitor keeps the system correct, and autoresearch makes it better. Both depend on the agent being unable to mark its own work as done. Production retests and walk-forward folds make that call.
Next, we will feed monitor findings into research tasks. For example, a cluster of forecast-API errors could automatically open an experiment on the model behind them. We will also add live calibration drift as a monitored signal.