Does signed aggressor flow predict short-horizon returns? Yes, clearly — and it is still not tradeable. Both halves matter.
python run_alpha.py # sign check -> dev grid -> single test eval -> cost check
python plot_alpha.py # -> results/alpha_summary.png
Raw tick data is not committed (~110 MB). Fetch the three days from Binance's
public market-data archive into data/:
BASE=https://data.binance.vision/data/futures/um/daily/aggTrades/BTCUSDT
for d in 2026-08-13 2026-08-14 2026-08-15; do
curl -O $BASE/BTCUSDT-aggTrades-$d.zip && unzip -o BTCUSDT-aggTrades-$d.zip -d data/
done
run_alpha.py expects data/BTCUSDT-aggTrades-2026-08-13.csv (dev) and
-2026-08-14.csv (test). Requires numpy, pandas, scipy, matplotlib.
| Data | Binance USDT-M aggTrades, tick level, ~708K trades/day |
| Bars | 1 second, VWAP-priced, ~79.8K active bars/day |
| Dev | Aug 13 — all 120 feature x horizon combinations scored here |
| Test | Aug 14 — touched once, with the single dev-selected config |
| Held back | Aug 15 (partial day), never opened |
| Selected on dev | ofi (signed volume) at a 1-second horizon |
| Dev IC | +0.2533 |
| Test IC | +0.2586, 95% block-bootstrap CI [+0.2495, +0.2672] |
| Shrinkage dev → test | −2.1% (i.e. none) |
| Gross edge | +0.0335 bps per trade |
| Round-trip taker cost | ~9 bps |
| Breakeven | ~268x the observed edge |
1. The trade-sign convention is verified, not assumed.
is_buyer_maker == True means the buyer was resting, so the aggressor sold.
Getting this backwards flips the sign of every downstream result and nothing
else in the pipeline would complain. So it is asserted against price impact:
IC(signed volume, same-bar return) = +0.5195. Buyer-initiated flow pushes
price up within the bar. If that check failed the run aborts.
2. A first version found the wrong answer, and the reason was the price series. The initial mid proxy was built from the last buyer- and seller-initiated trade in each bar, forward-filled when a side was missing. That side was stale in 38.9% of 1-second bars, which smears each move across several bars and manufactures positive return autocorrelation:
| price series | lag-1 autocorrelation |
|---|---|
| ffilled mid proxy | +0.110 |
| plain last trade | +0.082 |
Momentum then "predicts" the catch-up, and mom_5 beat ofi on the dev grid —
an artifact of series construction, not an effect. Switching to VWAP, which
exists in every bar with a trade and so is never carried forward, ofi wins
instead. That is the economically sensible answer, and it only appeared after
the artifact was removed.
3. IC, not accuracy — and OFI is checked for redundancy against momentum. Directional accuracy is close to meaningless here: returns cluster near zero, so a predictor can be right 51% of the time on noise and lose on the few large moves. IC is scale-free and respects magnitude ordering.
OFI and short-horizon momentum are nearly the same quantity measured two ways, so a raw IC for OFI risks just re-reporting momentum. Partial Spearman, on the test day:
| horizon | raw IC(ofi) | ofi │ mom_5 | mom_5 │ ofi |
|---|---|---|---|
| 1s | +0.2586 | +0.1991 | +0.1525 |
| 5s | +0.1773 | +0.1192 | +0.1515 |
| 30s | +0.0674 | +0.0436 | +0.0596 |
OFI keeps ~77% of its IC after conditioning at 1s, so it is not redundant. The dominance flips by 30s, where momentum carries more of the signal — consistent with flow imbalance being a very short-lived microstructure effect.
4. A positive IC is not money. 120 configurations were searched on dev. Under the null, the largest |IC| among many draws is not centred at zero — selection alone inflates it (simulated E[max |IC|] over 20 noise configs ≈ 0.041). The test number is readable at face value only because the test day was never used for selection.
And even a real IC of 0.26 loses: the gross edge is 0.0335 bps against ~9 bps of round-trip taker fees, so it needs 268x to break even. Capturing this would require posting passively and earning the spread rather than crossing it, which is a market-making problem, not a prediction problem.
- Trade flow, not order-book flow. Binance publishes tick-level trades free; its book snapshots are sampled ~every 25s, far too coarse. This is the Lee-Ready / Kyle strand (signed aggressor volume), not queue imbalance at the touch. Worth stating precisely — "order book imbalance" would be a different and stronger claim than the data supports.
- Two days of data. Nothing here establishes stability across regimes.
- One instrument, one venue.
- The cost model is a flat taker fee. No queue position, no slippage, no latency, no adverse-selection modelling.
- VWAP within a bar is still not the true mid; a genuine L1 book feed would be the correct series.
