Prediction markets look like the easiest thing in the world to backtest. Price lives in [0, 1]. Payoff is a dollar or nothing. No leverage, no funding, no basis, no perpetual-swap weirdness at 00:00 UTC. We had a backtester that already handled crypto perps with fees, funding and impact, and the plan was to spend two days adapting it and then start generating strategies.
It took eight. Here's the log, roughly in order, including the day the equity curve looked incredible and was.
Day 1: the adapter that should have been trivial
The initial mapping was one-to-one and felt clean. A market is an instrument. YES price is the price. Buying YES is long, buying NO is short. Resolution is a forced close at 1.00 or 0.00. Feed the hourly price series into the same event loop, done.
That mapping survives about ninety minutes of contact with real data. The first thing to go was the returns series. Our whole risk layer — position sizing, vol targeting, the Sharpe reporting — assumed log returns of a continuing instrument. A market that goes from 0.04 to 0.00 has an undefined log return and a defined, total loss. And a market sitting at 0.04 three days before resolution is a categorically different object from one sitting at 0.04 four minutes before, even though both rows say 0.04.
We ended up reparametrising the whole series on time-to-resolution rather than wall clock, and treating PnL as terminal rather than marked. That was the correct call, and it invalidated every piece of reporting code downstream of it.
Day 2: the fee formula is the strategy
On Kalshi the trading fee isn't a basis-point haircut. It's quadratic in price: roughly 0.07 × contracts × P × (1−P), rounded up to the cent at the order level. Which means the fee is largest exactly where a coin-flip market is most interesting, and nearly free out in the tails.
| Price | Fee per contract | Edge needed (prob. points) | Fee as % of stake |
|---|---|---|---|
| 5¢ | 0.33¢ | 0.33 | 6.7% |
| 10¢ | 0.63¢ | 0.63 | 6.3% |
| 25¢ | 1.31¢ | 1.31 | 5.3% |
| 50¢ | 1.75¢ | 1.75 | 3.5% |
| 75¢ | 1.31¢ | 1.31 | 1.8% |
| 90¢ | 0.63¢ | 0.63 | 0.7% |
| 95¢ | 0.33¢ | 0.33 | 0.35% |
Read the third column as the thing that matters: to break even buying at 50¢ and holding to resolution, your probability estimate has to beat the market's by 1.75 points. Not 1.75 percent relative. Absolute. If you exit before resolution instead of holding, you pay it twice, plus the spread.
Polymarket's CLOB has historically carried a very different fee profile, closer to zero on the trade itself, which moves the entire cost into spread and depth. So the same strategy logic has two completely different economics depending on venue, and any cross-venue comparison that uses one fee model is fiction. We now carry a per-venue fee callable, not a scalar.
If you read one line of a prediction-market backtest report, make it the fee line in probability points. A strategy with a claimed 1.2-point forecasting edge on 50¢ markets is a losing strategy on Kalshi before a single other cost. That fact should be visible on page one, not derived by the reader.
Day 3: the book is a staircase, and it's short
Our impact model was a square-root function calibrated on BTC perp order books. It's the wrong shape. With a 1¢ tick on a $1 instrument, there are only 99 possible price levels in the entire book, and a mid-liquidity market might have a couple of hundred contracts resting at the top level and then a gap.
Here's a snapshot from an econ-data market we were sampling, YES side, about 30 hours before resolution:
| Level | Size (contracts) | Cumulative cost of buying |
|---|---|---|
| 42¢ | 310 | 310 @ 42.0¢ avg |
| 43¢ | 85 | 395 @ 42.2¢ avg |
| 45¢ | 140 | 535 @ 42.9¢ avg |
| 49¢ | 600 | 1,135 @ 46.1¢ avg |
A 1,000-contract order — five hundred dollars of risk, a rounding error in crypto — moves the effective price four cents. Four cents is more than twice the fee. There's no smooth function to fit here; you walk the book level by level or you're making it up. We deleted the sqrt model for this asset class and wrote a literal book-walker, which is slower and correct.
Somewhere in here I caught myself being annoyed that these markets are "too small to matter." They're small because the thing being priced is a single real-world question with a deadline, and there are only so many people with an opinion about US initial jobless claims on a Wednesday. The size is the market. Complaining about it is like complaining that a poker table only seats nine.
Day 4: the day the equity curve was beautiful
Sharpe 4.1 on 1,400 markets. Smooth. Every researcher's stomach should drop at this and mine did, eventually, about forty minutes after I'd already screenshotted it.
The bug was in the resampler. Price history for illiquid markets comes back at irregular timestamps, so we reindexed onto a uniform hourly grid. The final point of every archived series is the settlement value, 1.00 or 0.00. Our reindex used a nearest-neighbour fill with no direction constraint, so for any market whose last trade was hours before resolution, the settlement value propagated backwards across the gap. The model was reading tomorrow's answer in today's row for exactly the illiquid markets where it could size up without hitting depth limits.
Nearest-neighbour fill is fine on a continuous instrument. On an instrument whose last observation is the ground truth, it's a lookahead machine. We now assert that every fill is forward-only and that the settlement row is tagged and excluded from any feature computation. That assert is nine lines long and it is the most valuable code we wrote all week.
Day 5–6: 1,412 markets, roughly 180 bets
Sample size on prediction markets is a trap with a friendly face. You resolve 1,412 markets, you feel like you have 1,412 observations, and you compute a t-stat accordingly. But fourteen NFL games on the same Sunday share a weather system, an injury-report release and a common bettor flow. Fifty-one state-level election markets share one national swing. "Will X happen by March 31" and "will X happen by June 30" are the same bet with different expiries and they resolve together.
We clustered markets by underlying event source and resolution date and got an effective independent count somewhere near 180. That's not enough to distinguish a real edge from noise at the effect sizes we were looking at, and no amount of per-market bootstrapping fixes it, because the bootstrap has to resample clusters rather than rows. Getting that wrong inflates your significance by a factor of two or three, quietly.
Day 7: resolution is a probability distribution too
The things we hadn't modelled at all, discovered the hard way:
- Early resolution. A "by December 31" market can resolve on August 4 when the event actually happens. Your capital comes back sooner and your holding period was never what the contract name implied.
- Disputes. Oracle-resolved markets have a challenge window. A position can sit in limbo for days past the event, and in a small number of cases the outcome flips from the obvious reading.
- Voids. Ambiguous source data gets a market cancelled and everyone refunded. In our sample that's under 1% of markets, but they cluster in exactly the ambiguous questions a model finds mispriced.
- Capital lockup. A 3-point edge realised over 90 days and a 3-point edge realised over 6 hours are not comparable, and a backtest that reports total edge without a capital-time denominator will always prefer the slow one.
We added a holding-cost term and a resolution-date jitter to the simulator. Neither is precise. Both are better than the implicit assumption that resolution happens exactly on the stated date with certainty.
Day 8: what we'd skip, and what we'd do first
- Skip the returns-based scaffolding entirely. Don't adapt a vol-targeting risk layer to an instrument with terminal payoff. Write a bet-sizing layer that speaks in probability and stake. We lost two days trying to make the old one fit.
- Build the fee function before the strategy. Print the break-even edge curve for every venue you intend to trade and pin it above the desk. It kills about half of the ideas before they cost you a backtest run.
- Walk the book from day one. Any parametric impact model imported from a continuous market will be wrong in a way that flatters you.
- Cluster before you count. Decide the clustering key for effective sample size at the same time you decide the universe, not after you have a Sharpe you like.
- Assert on fill direction. One line per join. If a series ends in ground truth, treat backward fill as a bug by construction rather than something you'll notice in review.
The eight days were worth it, mostly because prediction markets strip away the ambiguity that makes crypto backtests so easy to fool yourself with. There's no funding rate to hand-wave, no mark-price convention to argue about. Just a question, a deadline and an answer. When the backtest is wrong here, it's wrong in a way you can point at.
← All posts


