BitBank
Blog > How BitBank Improves Itself

How BitBank Improves Itself: Monitoring Agents and Autoresearch

BitBank is run by a very small team. It serves live crypto forecasts, a trading game, a token launchpad and on-chain analytics. Two groups of AI agents keep it working and make it better. Monitoring agents watch production and repair what breaks. Autoresearch agents run experiments on our forecasting and trading models and ship a change only when the evidence is strong. This post explains how both loops work, what they got right this week, and where they still need a human.

Loop 1: the production repair agent

Every five minutes a supervisor collects three kinds of signal: browser errors reported by our pages, errors from the Go backend and the forecasting sidecar's system journal, and health checks on the site and its session API. When it finds new problems, it starts a coding agent in the repository and gives it a bounded repair task.

  1. Collect and filter. Most error reports on the public internet are not bugs. SQL-injection and XSS probes, browser extensions, analytics beacons and scripts injected by in-app browsers all end up in the error log. The supervisor drops these before any agent sees them.
  2. Fingerprint. One broken route hit by a fuzzer can produce hundreds of distinct URLs. Issues are grouped by route, method and status, so one fault is one issue.
  3. Triage before repair. The agent first searches our incident history, then checks whether the failure still happens, locally and over public HTTPS. Only then does it change code.
  4. Fix, test, deploy, retest. A minimal fix, a focused regression test, the documented deploy procedure, then a retest of the original failure.
  5. Per-issue verdicts. The agent closes each issue on its own line, as FIXED with evidence or HARMLESS with a reason. Anything it cannot close is escalated to a human by email.
Bounded authority. Repair agents cannot change trading policy, move funds, edit business data, rotate credentials or run destructive migrations. One lock serializes all repairs, a runtime cap kills a stuck agent, and a cooldown limits how often an issue is retried. If the primary coding model is rate limited, the job falls back to a second and then a third provider before it gives up and emails us.

What went wrong, and how we fixed the monitor itself

The first version was too strict in a way that made it useless. An incident closed only if the agent proved every issue in the batch fixed. One unprovable item was enough to hold the whole batch open: a script error injected by the Twitter in-app browser on an old iPhone, referencing variables that do not exist in our code. The same 40 issues came back every six hours, each run ended in an email, and the queue grew to 153 entries.

ChangeEffect
Noise filter for probes, extensions, beacons, in-app webviews and SQL echo lines153 pending issues fell to 19
Route-level fingerprintsA fuzzed endpoint counts as one issue, not 20
Per-issue verdicts instead of all-or-nothingFixed items close, and only the rest escalate
Transient errors must recur 3 timesOne upstream timeout no longer starts a repair
Go panics carry their stack traceThe agent can find the faulting line
Freshest issues first, and a 7-day expiryAgent time goes to current problems

The first run after these changes received three issues, found prior evidence for each, retested them, and closed all three without sending an email. The same day it caught a real browser bug. Some early Chromium builds on Android, reported from Honor phones, reject the array form of the canvas roundRect radius argument, so the homepage chart crashed for those users. The chart now falls back to a numeric radius and then to a plain rectangle.

Loop 2: autoresearch on the forecasting stack

Our price forecasts come from Chronos-2, a time-series foundation model, served from a GPU sidecar with custom CUDA kernels. Autoresearch agents get research contracts: a fixed incumbent, a data boundary, a metric, and a promotion gate. They write harnesses, run walk-forward experiments, and report what survives.

This week's run found two real defects in how we used the model:

  • A precision bug. One step in our accelerated inference path converted the final price forecasts back to bfloat16. At a Bitcoin price near $86,000, bfloat16 values are 512 dollars apart. Every forecast snapped to that grid, an error of about 0.4% for every asset. Output now stays in float32, and the accelerated model agrees with the reference to within about 0.01%.
  • A compounding path. The forecast line was built by chaining per-step open-to-close predictions, which compounded the noise in each step. It is now built from the median close forecast, damped toward the last price. The damping strength came from 5-fold cross-validation, and every fold chose a value between 0.05 and 0.10.
Forecast lineMAE beforeMAE afterChangeFolds improved
Chart line3.22%1.79%-44%45 / 45
Homepage and snapshots3.06%2.11%-31%45 / 45
Flat line (no-change baseline)1.79%

This was measured end to end over HTTP against shadow copies of the live sidecar: 9 trading pairs, 5 chronological folds, January 2024 to September 2026. The change has a one-line rollback.

What the agents told us that we did not want to hear

Look at the last row of that table. After the fix, the chart line is only about as accurate as a flat line that assumes the price will not move. On hourly crypto data, the raw model's median forecast was 4 to 7% worse than that flat baseline for horizons of 1 to 24 hours, across 57 pairs, and no context length fixed it. The best stacked model, gradient-boosted trees on lags, volatility and Chronos quantiles, beat the flat line by about 0.16%. Short-horizon crypto prices are close to a random walk, and we would rather publish that than hide it. See our accuracy page for live numbers.

The autoresearch agents also searched variations of our live rotation strategy: regime gates, other horizons, and Chronos features. Removing the trailing stop raised average fold returns, but it won only 20 to 23 of 38 paired folds, with 40 to 50% drawdowns, which fails our promotion gate. Nothing was promoted. The search also showed that the current profile's headline backtest is a fragile peak: most single-parameter changes land well below it. A result like this counts as a success, because it stops a bad change from reaching production.

Why two loops, not one

The two loops are built alike. Each has a narrow contract, a verifier the agent does not control, and a record of past incidents or experiments to read before starting. Each can end in "no change" or "escalate". The monitor keeps the system correct, and autoresearch makes it better. Both depend on the agent being unable to mark its own work as done. Production retests and walk-forward folds make that call.

Next, we will feed monitor findings into research tasks. For example, a cluster of forecast-API errors could automatically open an experiment on the model behind them. We will also add live calibration drift as a monitored signal.

Further reading