Markets are self-improving, as models trade better, the market inefficiencies disappear, so trading well gets harder the better you get. It's a never-ending hill to climb.
I ran SOTA reasoning LLMs on it. TL;DR: they suck, and reasoning doesn't help. No model came close to a simple static benchmark (experienced human Quant chose parameters);More reasoning ≠ better trading; Fun fact: when losing, LLMs trade less rather than smarter.
Why quant trading is a great intelligence test (minimal jargon): a) requires ML research that generalizes OOS (alpha); b) needs robustness to regime shifts and new rules; c) trains long-horizon planning under tradeoffs (how to trade now if AAPL is +2% tomorrow but −5% the day after). And d) it's self-correcting: alphas decay, trading too much gets you adversely selected, and trading well makes the market more efficient — so it never stops being hard.
Background: Ex-quants from G-Research (ML algo trading) & TransMarketGroup (Crypto options); + LLM inference kernels at Etched.
Please poke holes in the setup - where does "markets as an eval" break down? Especially keen on views from outside quant.