BitBank
Blog > Reading Forecasts and Bots Honestly

Reading Crypto Forecasts and Trading Bots Honestly

BitBank publishes forecasts, accuracy figures, a risk view, a backtester and a trading bot. Each answers a narrower question than its title suggests. Every figure below was read from live JSON endpoints on 30 September 2026 and has moved since; the point is the method, not the values.

What a confidence band is

The forecast pages at /coins and /coin/{pair} draw a median path from Chronos2 with an upper and lower band around it. In the sidecar code the model is asked for the 10th, 50th and 90th percentiles of each future candle. The 10-90 spread is then multiplied by about 1.53 (the ratio of the 95 percent to the 80 percent Gaussian quantile) and the result is widened so it always contains the forecast candle's own high and low. So the band is a nominal 95 percent interval built on a normal-shape assumption, not an empirically calibrated one.

Crypto returns have fat tails, so a Gaussian-shaped band will be breached more often than 1 time in 20 when the market gaps. I found no published coverage statistic for the bands in the handlers or docs, so treat them as the model's statement of its own uncertainty, not a guarantee. A band is a width, not a direction.

The width grows quickly with horizon. On the hourly forecast for BTC_USDT at retrieval:

HorizonMedian closeBandHalf-width (band_pct)
+1 hour$83,345$82,435 - $84,0871.09%
+6 hours$83,373$81,912 - $84,8331.75%
+12 hours$83,353$81,507 - $85,1982.21%
+24 hours$83,458$80,270 - $85,9173.82%

The daily forecast is wider still: day seven was $70,373 - $97,186 (16.00 percent) around a median of about $83,780. The median barely moves while the band covers roughly 16 percent either way. That is the model saying it has no strong view, which is the honest answer for most crypto horizons.

Check /accuracy before trusting a pair

The /accuracy page is backed by /api/trading-bot/accuracy. It reports mean absolute percentage error per horizon, the standard deviation of that error, and a sample count. At retrieval:

HorizonMAEStdev of errorSamples
1 hour0.326%0.376%1,425
7 hours0.945%0.942%1,407
1 day1.976%1.893%1,356
7 days6.429%4.298%924

Three caveats come from the code. First, the source field read live_poloniex_holdout: the handler had no stored observations to score and replayed the site's simple point forecaster over recent Poloniex candles, so this is not a scorecard of stored Chronos2 paths. Second, that fallback covers exactly three pairs: BTC_USDT (MAE 0.239 percent, 475 samples), ETH_USDT (0.314) and SOL_USDT (0.425). For any other pair there is no figure here to check, and that absence is information. Third, the stdev is about as large as the MAE, so typical and bad errors differ a lot.

The useful question is never "is 0.326 percent small" but "is it smaller than predicting the last close?" The point-forecast code carries a comment with its own measurement over 51 Binance pairs from 2021 to 2026: against a persistence baseline the current rule gains only a few hundredths of a percent beyond the first hour and effectively zero at 168 hours, and the comment's own summary is that hourly crypto is very close to a random walk. A low MAE at one hour mostly reflects that prices barely move in an hour. Always compare an error figure to a no-skill baseline before calling it skill. See also inside the forecasting models.

The /risk page is a different instrument

/risk computes an outlook and band per market from hourly price history. The docs say it is a heuristic, not a Chronos2 forecast or the bot's signal; the JSON carries signal_source: risk_heuristic and a stale flag for candles over two hours old. At retrieval the first seven cards (BTC, ETH, XRP, ZEC, SOL, BNB, SUI) all read Neutral and hold, BTC's band being $83,002 - $83,750 around $83,374. Use it to compare recent movement across markets, not to pick direction; see risk and uncertainty.

Backtest versus executed fills

Every bot result on the site is simulated, and the pages say so. The public performance endpoint states its execution model: signals are generated at candle close, and simulated fills execute at the next candle open with fees plus 5 bps of slippage per side. In code, the public bot charges a 0.06 percent fee plus 0.05 percent slippage per side, which is 22 bps round trip. The /backtest form defaults to a 0.1 percent fee, and the simulator floors slippage at 5 bps even if you type zero, so a default round trip there costs about 30 bps.

What a simulation cannot know is documented too: the research notes list spreads, queue depth and exit capacity as unmodelled, reported opens that may be stale, and assumed full fills. An hourly delay there is a stress scenario, not a measured latency. The rotation endpoint's own warning reads: forward paper account, assumed costs and full fills, not real orders or historical backtest returns, with live_orders_enabled false. See honest backtesting for the mechanics.

Read the bot's numbers as a distribution, not a headline

The public summary for the trailing 30 days showed a total return of 3.24 percent, a maximum drawdown of 3.17 percent and a win rate of 23.1 percent across 26 closed simulated trades, equal-weighted over six pairs. The per-pair breakdown is more informative than the total:

PairReturnTradesWin rate
ZEC_USDT+23.09%250%
LINK_USDT+6.48%250%
ADA_USDT+2.54%333.3%
ETH_USDT-2.28%616.7%
SOL_USDT-5.17%812.5%
BTC_USDT-5.24%520%

Three of six pairs lost money. The basket's average of 3.24 percent is carried by one coin, and with 26 trades in total, nothing here separates skill from a good month. The same pattern appears in the research docs: in the delayed-execution study, removing ZEC turned the declared primary strategy from a positive result to negative 6.55 percent, and the authors wrote that the sizing rule reduces drawdown but does not establish a diversified profitable edge.

The forward paper portfolio on /trading-bot is the cleanest check, because it cannot be tuned after the fact. Its published state showed equity of $10,164.71 on $10,000 (up 1.65 percent), maximum drawdown 2.43 percent, $34.82 of modelled fees and five completed trades. That is a short window and still paper; until the record is long it is an observation, not evidence.

Overfitting is the default, not the exception

Change enough parameters on a fixed history and something will look good. The repo's notes say so: one search tried 192 configurations and 8 passed every validation fold at doubled costs; a second tried 576 and 30 passed, and the notes conclude that repeated selection on inspected history is not evidence of a live edge. Every slider you move in the backtester is another configuration tried. Decide the rule first, test it once on untuned data, and double the fee and slippage inputs. The mixed basket's 2x gross exposure looked optimal on eight months of data; the longer backfill averaged negative 8.0 percent per three weeks at 2x against negative 3.9 percent at 1x, so it runs at 1x.

Fees and turnover

A 22 to 30 bps round trip has to be beaten by the average trade before it earns anything, and hourly strategies pay it hundreds of times a month. Compare trade counts first: fewer, longer holds are far less sensitive to a wrong cost assumption.

Position sizing

Sizing decides whether a bad streak is survivable, and a forecast says nothing about it. The documented research portfolio sizes each entry from volatility, capped at 15 percent of initial capital, keeps a 5 percent cash reserve and uses no leverage, yet its drawdowns still ran in the mid teens; the docs call it an entry budget, not a guarantee. For your own account, fix the maximum loss you accept per position, size from the distance to your exit rather than a forecast's conviction, and do not lever an unmeasured signal. More in risk management in trading.

When not to trade

  • When the band is wide relative to the expected move. A median a fraction of a percent off spot inside a 4 percent band carries almost no information.
  • When the pair has no accuracy figure, few samples, or an error not clearly better than last-close persistence.
  • When the evidence is a backtest you tuned yourself, a handful of simulated trades, or one coin's month.
  • When the risk card is stale or a page labelled simulated is being read as live.
  • Around scheduled macro events and expiries, when spreads and slippage run above the simulator's 5 bps.

A practical order of operations

See the band at /coins, find the pair on /accuracy and compare its error to last-close persistence, check /risk for recent movement, test any rule on /backtest with doubled costs, and watch the forward paper record on /trading-bot and /dashboard for as long as you can. Only then size something small. The tools show what the model believes and how wrong it has recently been; they do not tell you whether to be in the trade.

Not financial advice. Forecasts and simulated results are not guarantees and real fills will differ. Figures read from bitbank.nz endpoints on 30 September 2026 and change continuously; check live data before acting.