AI Research XLEXLE_newsmacro:brent_daily

XLE vs Brent around Iran/Hormuz/sanctions/tanker headline spikes

4
Spike-intensity days

Only four days qualified. Across roughly three years of daily headline tracking, that is the entire top-quintile sample of Iran/Hormuz/sanctions/tanker news spikes — a reminder that geopolitical risk often arrives in sharp bursts, not sustained runs. With a sample that thin, every mean difference carries a wide error margin.

The test was simple: on spike days, did XLE lag Brent's same-day move and then close the gap over five sessions? No. XLE ran about 1.9 percentage points ahead of Brent on those days, then trailed by roughly 5.3 points over the following week — the opposite of the expected catch-up. Both gaps sit in statistical noise, with same-day p≈0.49 and five-session p≈0.26.

The honest read is inconclusive rather than a clean rejection. The full breakdown below shows how the headline regimes were built, how the returns were aligned, and why four sessions cannot settle how XLE prices Hormuz risk.

The research question

Over the past ~3 years, when Iran/Hormuz/sanctions/tanker headline intensity spikes into the top quintile of its prior-20-session distribution, does XLE's same-day return lag Brent crude's same-day move by more than on low-intensity days, and does that lag close over the next five sessions? I expect energy equities to underprice the initial supply-shock risk premium, so XLE falls behind Brent on the spike day and then catches up as the Hormuz premium gets built into energy equity valuations.

How this was measured

Built daily XLE and Brent close-to-close returns from XLE minute bars and the global brent_daily_df macro frame, then aligned both series on common trading dates. For news intensity, parsed XLE headline text (title, summary, topics) for a fixed keyword set (Iran, Hormuz, sanction(s), tanker, Strait of Hormuz), summed daily relevance-weighted article counts, and classified each day against the trailing 20-session distribution. A spike is top-quintile intensity with positive news count; low-intensity is bottom-quintile intensity. We then measured same-day Brent-minus-XLE return and forward five-session XLE-minus-Brent return by regime, comparing means with Welch t-tests. The first 20 sessions are dropped because the trailing distribution is not yet available.

The key numbers

Spike-intensity days
4
Top-quintile keyword intensity with positive news count
Low-intensity days
698
Bottom-quintile keyword intensity
Mean same-day lag (Brent-XLE) — spike
-1.8916%
N=4
Mean same-day lag — low
-0.0361%
N=698
Same-day lag: spike minus low
-1.8555%
Positive means larger XLE lag on spike days
Same-day lag Welch t
-0.788
Positive favors spike days lagging more
Same-day lag Welch p
0.4878
p=0.4878 >= 0.05 -> no statistically-clear difference
Mean 5d catch-up (XLE-Brent) — spike
-5.2581%
N=4
Mean 5d catch-up — low
0.2352%
N=693
5d catch-up: spike minus low
-5.4933%
Positive means XLE catches up to Brent over the next 5 sessions
5d catch-up Welch t
-1.383
Positive favors catch-up after spike days
5d catch-up Welch p
0.2603
p=0.2603 >= 0.05 -> no statistically-clear catch-up difference

Reading the numbers

Only 4 headline-spike days vs 698 quiet ones, so comparisons are fragile. Spike-day Brent-vs-XLE gap averaged -1.89% vs -0.04% quiet (p=0.488), and 5-day catch-up averaged -5.26% vs +0.24% (p=0.260). The p-values mean neither difference is statistically clear.

The charts

Same-day XLE-Brent lag by headline intensity regime
What this chart says

The spike-day box is built from just four sessions and sits at a mean of -1.89% on the Brent-minus-XLE scale, while the low-intensity box of 698 sessions is near -0.04% but ranges much wider, from about -13.5% to +11.8%. The detail to note is that the spike-day mean is not a positive gap of the kind that would show XLE clearly falling behind Brent; with p=0.488, this difference is easily explained by random noise.

5-day XLE-Brent catch-up by headline intensity regime
What this chart says

For the five-day window, XLE minus Brent averages -5.26% after spike days versus +0.24% after low-intensity days, with the four spike observations ranging from -15.4% to +2.7%. A positive value would mean XLE catching up to Brent, so the negative spike average is the opposite of that catch-up, and p=0.260 says the apparent shortfall is not statistically distinguishable from ordinary variation.

Average cumulative XLE-Brent forward differential
What this chart says

The line chart follows the average cumulative gap from T+1 through T+5. After spike days the line starts at -0.44% and drifts down to -5.26%, while after low-intensity days it stays almost flat between +0.04% and +0.24%. The downward slope after spikes is the key detail: the gap widens over the next five sessions instead of closing, so these data do not show energy equities catching up to Brent after headline spikes.

Intensity-regime contrast

RegimeNMean same-day lagMedian same-day lagMean 5d catch-upMedian 5d catch-up
Spike (top quintile)4-0.0189-0.0376-0.0526-0.0418
Low (bottom quintile)698-0.0004-0.00130.00240.0057

The takeaway

No — the data do not show XLE lagging Brent on headline-spike days or closing the gap over the next five sessions, so the expected underprice-then-catch-up pattern is not supported. In the roughly three-year window only 4 days actually qualified as top-quintile Iran/Hormuz/sanctions/tanker spikes; on those days XLE's same-day return averaged about 1.9 percentage points ahead of Brent, not behind, and over the following five sessions XLE trailed Brent by roughly 5.3 percentage points on average — the opposite direction of a catch-up. Both contrasts are firmly in noise territory: the same-day difference had about a 49-in-100 chance of being luck (p≈0.49) and the five-day difference about 26-in-100 (p≈0.26), with just 4 spike days behind each mean. The practical takeaway is that this sample is far too thin to conclude anything reliable about how XLE prices Hormuz risk — the sign of the gap even ran against the hypothesis, but the honest read is inconclusive rather than a tested rejection.

The fine print