The methodology page documents how the signals are defined and scored. This page answers the question underneath it: where each construct comes from in the published literature and in desk practice — and, separately, which thresholds are house convention. The distinction matters, because a monitoring system that blurs it is asking to be trusted rather than checked.
As of August 2026. Reviewed adversarially before publication, with every citation below venue-verified. The catalog changes; this page is dated and revised with it.
The design principle: the platform computes standard, documented measures; the signal layer applies named thresholds to them. Every firing signal reports its rule, its input value, and a rendered trigger sentence — "crude stocks seasonal percentile of 0 is below the 20 threshold" — and signal_backtest scores each rule against its own history. Nothing here is a black box, and nothing below is claimed as proprietary insight when it is standard practice.
Comparing current stocks to the same-week multi-year benchmark is the institutional framing. EIA's Weekly Petroleum Status Report states stocks relative to the five-year average for the time of year, and the companion This Week in Petroleum charts plot inventories against the five-year range; the IEA Oil Market Report benchmarks OECD industry stocks against the five-year average — the metric that anchored the OPEC+ rebalancing narrative.
The economic foundation is the theory of storage — Working (1933, 1949), Kaldor (1939), Brennan (1958) — with the classic empirical treatment in Fama & French (1987) and the modern restatement in Gorton, Hayashi & Rouwenhorst (2013), who show that inventory deviation from normal — exactly what a seasonal percentile measures — is the state variable organizing commodity risk premia and basis behavior.
House convention: the 20/80 quintile cut-offs, and 90 for products. These are hand-chosen, not fitted, and the backtest exists to score them. The window is a real limitation: a five-year base includes regime distortion — EIA footnoted the same problem when 2020 entered their five-year averages — and on a five-year same-week base there are five observations per week, so "above the 80th percentile" is close to "above the five-year maximum". The measure is coarse by construction.
Cushing is the NYMEX WTI physical delivery point. Its working storage is finite and its effective bounds — tank bottoms and operational minimums at the low end, shell capacity at the top — are operationally documented. On 20 April 2020 the May WTI contract settled at −$37.63/bbl (intraday low −$40.32) with Cushing's uncommitted storage effectively spoken for: the canonical exhibit of Cushing state dominating WTI structure at the extremes. The CFTC's Interim Staff Report (November 2020) documents the storage conditions and pointedly stops short of causal attribution, and this page does not go further than the report does.
The z-score form makes "full" and "tight" precise, with one honest caveat: the physical stress line traders actually watch is distance to operational minimums (~20 million barrels) and to tank tops. A z-score approximates that distance; it does not measure it. Cushing's shell capacity also grew materially over the 2010s, so a level-based z spans a capacity regime change.
The curve metric is deferred minus prompt (M12−M1), so negative means backwardation; the key name m1_m12 is legacy and the live schema says so. steep_backwardation fires below −$10 — the prompt trading $10 or more over the 12-month — and steep_contango above +$5. punishing_carry is roll yield (M12/M1 − 1) below −15%: negative in backwardation, and "punishing" from the point of view of storage economics, since the curve is paying the market not to store. That is the opposite sign of the long-investor "roll yield" convention. Read with the other convention in mind, these definitions look inverted — which is exactly why they are spelled out here. The signals were firing correctly; the labels were not saying so.
Term structure and roll yield are among the most documented objects in commodity finance: Erb & Harvey (2006), Gorton & Rouwenhorst (2006) and Koijen, Moskowitz, Pedersen & Vrugt (2018) all show that curve shape and carry organize returns, and the inventory-to-backwardation link is the theory-of-storage literature of §1. Brent–WTI as a location and infrastructure spread has its own literature from the 2011–2013 dislocation — Fattouh (2011), An Anatomy of the Crude Oil Pricing System, OIES WPM 40; Kilian (2016) on shale-era logistics.
One structural caveat a spread trader will raise first: WTI Midland has been deliverable into the Brent basket since June 2023, tethering Brent–WTI to freight plus quality. The spread's semantics changed at that point, and the signal's history spans the break.
The Generalized Supremum ADF test is Phillips, Shi & Yu (2015a, 2015b — two companion papers, International Economic Review 56(4)), the standard econometric test for explosive episodes. It has been applied to oil specifically by Caspi, Katzke & Gupta (2018, Energy Economics) and by Fantazzini (2016, Energy Policy), who identifies the 2014–15 crash as a statistically significant negative bubble. This is the signal with the least distance between the published method and the implementation.
The framework is Kilian (2009, AER): an oil price move means different things depending on whether it is supply- or demand-driven. The equity leg is the companion paper, Kilian & Park (2009, International Economic Review), which finds demand-driven oil moves co-moving positively with stocks. The specifically VIX-based rendering — using volatility and risk innovations to separate demand from supply shocks — belongs to the later risk-on/risk-off literature, of which Ready (2018, Review of Finance) is the clean reference.
The catalog's rolling-correlation rule is a house simplification of that line of work, not a result from it, and the correlation window materially affects when it fires. It is Kilian-style shock discrimination, operationalized. The VIX rendering is not Kilian's, and this page does not credit it to him.
Tanker rates as the binding constraint on crude arbitrage — a wide Brent–WTI means nothing if the freight to move the barrels consumes it — is bread-and-butter physical trading, visible weekly in Baltic Exchange assessments and broker reports. The grounding here is practice rather than journals, and the z-score form is house. Kilian's global activity index is built from dry-bulk freight rates, which is evidence that freight carries macroeconomic information, but his use is demand-side and the transfer to tankers is indirect.
What the freight proxy actually is. BWET is the Breakwave Tanker Shipping ETF (NYSE Arca, 2023 inception) — an exchange-traded wrapper on tanker freight futures that cash-settle against Baltic assessments, with route exposure concentrated in VLCC Middle East–to–East. It is a liquid daily proxy, but it is not a Baltic index: it carries roll and fund-flow effects, and roughly three years of history, so a z>3 trigger rests on a short and regime-heavy sample. Route-specific Baltic assessments — TD3C, and TD25 for the US Gulf–to–Europe crude leg — are the higher-fidelity upgrade path.
Percentile-ranked managed-money positioning is a standard sell-side and desk monitor; every major bank publishes a CoT monitor in this exact framing. The academic record is less flattering. Sanders, Boris & Manfredo (2004, Energy Economics) find noncommercial positions in energy futures follow returns rather than predict them; Sanders, Irwin & Merrin (2009) reach the same conclusion across a broader panel; Büyükşahin & Harris (2011, The Energy Journal) find speculator positions do not Granger-cause crude prices. Büyükşahin & Robe (2014, JIMF) add the complementary finding that heavy speculative participation changes cross-market co-movement — so crowding matters for market character even where it fails as a timing tool. Kang, Rouwenhorst & Tang (2020, Journal of Finance) is the modern generation, identifying short-horizon premia in position changes as liquidity provision.
Taken together: positioning extremes are risk context — crowdedness, vulnerability to a squeeze — and not a directional signal. The catalog uses them that way. Its fundamental_rally contradiction fires on median positioning, using the absence of crowding to qualify a move rather than using an extreme to predict one. Brent managed-money data come from ICE Futures Europe's COT report, not the CFTC's.
The 3-2-1 crack is the standard refining margin proxy, documented as such by EIA and CME and used as the canonical refiner hedge structure. A screen 3-2-1 is not a full refinery margin: RINs and RVO obligations, natural gas costs, and in Europe carbon all move the real economics, and the signal should be read as a margin proxy rather than a margin.
On utilization: sustained rates above roughly 92–93% of EIA operable capacity mean the system has little spare conversion capacity. That is necessary but not sufficient for tight product markets. The demand-led episodes — 2005, 2022–23 — pair those rates with genuinely tight products; the honest counterexample is the late 1990s, when utilization above 95% coexisted with weak margins and preceded a wave of closures. High utilization measures the absence of slack, not the presence of margin. Turnaround seasonality in spring and autumn is documented in EIA reporting and industry trackers, and a runs-trough z-score marks it — though the refining family is the one group in the catalog still keyed to raw levels rather than seasonal percentiles, so refinery_weak and maintenance_detected co-fire through every normal maintenance season.
Heating and cooling degree days measured against normals are NOAA and CPC standard products, and EIA's own STEO demand models are explicitly degree-day driven. Anomaly z-scores are the textbook form. The ten-year base is a house choice rather than the standard one — NOAA and CPC normals are thirty-year, currently 1991–2020 — chosen deliberately to limit warming-trend bias in the heating-degree baseline.
The "call on OPEC" — now the call on DoC crude — is a standard construct published in the agencies' own reports. The wedge between direct-communication and secondary-source production figures is a real and known artifact: secondary sources exist precisely because self-reported figures are subject to interpretation, and tracking the per-member gap is legitimate compliance analysis, though the gap also reflects condensate boundary definitions and months where members do not submit. The OPEC–IEA demand divergence was one of the defining market debates of 2023–2025 and is tracked by every macro desk. The thresholds applied to all of these are house conventions on documented quantities.
The sixteen contradictions do not come from the literature. They are the house's trade-thesis logic — the one layer here with no citation behind it, and the one that adds something. Each is built from documented legs. refiner_margin_trap — crack extreme with gasoline oversupplied — is margin mean-reversion risk. freight_locked_arb — spread wide with freight extreme — is arbitrage economics. The Atlantic-basin family is location-spread fundamentals. forced_storage_build — Cushing building fast while carry is punishing, which by the §3 convention means building into backwardation — flags a theory-of-storage violation: the curve is paying the market not to store and stocks are building anyway, which points to involuntary builds, landlocked barrels, or a curve that is wrong. Building into contango would be unremarkable cash-and-carry; the sign convention is what makes the signal read correctly.
An analyst can disagree with any of them. The catalog names them precisely so that disagreement can be specific.
Every monitoring system has a scope. Here is where this one ends:
production_at_seasonal_high (renamed 2026-08-11 from production_at_record, which claimed more than the predicate tested) sits at the top of a secularly rising series and stays true for long stretches; opec_call_rising fires on any positive quarter-over-quarter change, which is the normal Q1→Q3 demand build; opec_compliance_gap_extreme is chronically satisfied by a couple of members; and the refining thresholds move with the turnaround calendar. They describe a standing condition rather than a new one, and should be read that way.cushing_tight currently returns exactly that warning. This affects 14 of the 53 signals: the 13 z-based predicates plus vix_oil_supply_shock, whose rolling correlation behaves the same way. Until the engine computes point-in-time statistics natively, backtest evidence certifies the point-in-time variant of a rule.The claim that the platform scores its own rules is checkable. Three runs from 9 August 2026, honesty guards intact, reported the way the guards require — independent cluster counts beside n, base rates beside conditionals, and the inconvenient results included:
Claim tested: record-low crude stocks are bullish crude. Target PET.RWTC.D, history from 1990, point-in-time. 15 firings → 15 independent clusters at the 5-, 10- and 21-day horizons, 14 at 63 days. Median forward WTI move −0.4% at 5 days against an unconditional +0.34% (edge −0.74 pp); −2.2% at 21 days against +0.81% (edge −2.96 pp). Hit rate in the thesis direction: 40%, 53%, 33%, 50% across 5/10/21/63 trading days.
Record-low stocks describe a state that spot and curve have already absorbed — which is what the theory of storage predicts. The signal is regime context, not a buy trigger, and the catalog treats it as such.
Claim tested: extreme cracks mean-revert. History from 2020. 4 firings → 4 independent clusters at 5 days, 3 at 10–21 days, 2 at 63. The crack kept rising: median +$2.45/bbl at 5 days, +$6.40 at 21, +$22.51 at 63, against unconditional moves of +$0.21, +$0.69 and +$1.18. Hit rate in the mean-reversion direction: 25%, 0%, 0%, 33%. Both flags fire — low confidence at n<10, and single-regime with three of the four episodes in 2022.
On this evidence the claim is not supported, and the sample is too small and too concentrated to support the opposite one either. That is the whole finding, printed.
The point-in-time replay returns a definition-mismatch warning instead of numbers: the live engine's full-sample z and the backtest's expanding-window z are different statistics, and the tool says so rather than scoring a signal the engine does not fire.
signal_backtest exists to score them rather than to justify them.derived_signals(schema=True) publishes every rule; every firing signal shows its arithmetic in a sentence; and the results above are published whether or not they flatter the catalog.The operational companion to this page — signal definitions as predicates, backtest guardrails, revision policy and the calibration ledger — is Methodology & Calibration. The full catalog, including every threshold and alias, is machine-readable via derived_signals(schema=True) on any connected client.