Market Blog

Oil-Stock Backtests: Too Many Signal Failures, One Absurd False Positive

Sometimes the most useful thing a quant research platform can produce is a list of things that clearly do not work. That is the vibe from the latest batch of published backtests on trades.run, where energy stocks and their macro triggers have been put through a cold, statistical wringer. The results do not make a neat bull or bear case — they make a humility case.

The Return That Is Just a Glitch

The most notable number in the batch is also the least believable. A backtest buying MPC when Brent falls more than 1% but MPC closes firm generated a return of 40527180621470278898662919509896109434173652992% on $100,000 across 260 closed trades. If taken literally, that would mean the strategy turned a hundred grand into more wealth than exists in the known universe. It also managed a 69% win rate with a best individual trade of only +8.19% and a worst of -12.81%. No sane combination of those trades gets anywhere near that figure. The data leans strongly toward a leverage or compounding artifact, not a tradable pattern. Anyone eyeballing that outcome as a green light for a commodity-conviction strategy should probably look again at the curve and ask what broke.

Mean Reversion in Oil Stocks Keeps Failing

While that one backtest appears to be a bug, the rest of the batch is boring in a different way: the signals just do not work. A buy-COP-after-RSI-crater setup, with Brent crude closing below a certain level as part of the trigger, lost -42.65% over 27 trades with a 22% win rate. Over the same window, SPY buy-and-hold rose +76.34%. That is not a subtle miss; it is a strong case that this version of dip-buying is fighting the tape. Similarly, a buy-XOM-after-gap-down-but-reversing-up setup churned out only +2.78% across 13 trades, while SPY again gained +76.34%. Even the best trade in that sequence, +3.39%, was far too small to compensate for the frequency of losses. These two mean-reversion ideas sound plausible in a meeting; in historical testing, they are the kind of edge that quietly bleeds.

Relative Strength Is Not Persistent Either

What about the more sophisticated angle — that one refiner or pipeline operator should outperform another after a macro move? The findings pour cold water there too. After a 10-day stretch where VLO outran MPC with Brent flat or down, the following 10 days did not favour VLO: MPC actually edged ahead, with a mean spread of about -0.62 points and a median closer to -1.2. VLO won only 8 of 19 qualifying follow-ups. The idea that XOM leads next-session Brent more strongly on top-quintile volume days also fails the significance test — the difference between high-volume and low-volume days has a p-value around 0.34, which is basically a coin flip. And in the broader macro setup, when 10-year yields drop more than 20bps over 20 sessions and Brent sits above its 50-day SMA, WMB has actually underperformed XLE over the subsequent 20 sessions. WMB's mean forward return was -0.56% versus +2.38% for XLE, with WMB winning only 6 of 22 sessions. So even the relative-value trades, the ones that feel structurally sound, have a poor recent history.

What ties these findings together is a clear point of view: the oil patch is a graveyard for pattern-based conviction. Strong narratives — buy the oversold bounce, catch the reversal, ride relative strength after a macro cross — all sound good until they run into actual data. With no major headline catalyst shifting energy sentiment right now, the quant output suggests the market's intraday and daily moves in these names are close to noise. The lean here is not toward a short or a long; it is toward scepticism. The most defensible read of these backtests is that they are doing their real job: killing bad ideas before a human can put real money behind them.