How the reports earn trust

A professional reader of any analytical publication asks four questions: How are the signals defined? Can I inspect the methodology? Can I reproduce the numbers? How well do they calibrate over time? This page answers all four — and every answer below is live-derivable from the platform, not marketing copy about it.

The design principle: the prose in an EnergyScope report sits on top of a structured analytical layer with explicit signal definitions, pre-registered scoring rules, a public revision trail, and (from July 2026) a Brier-scored calibration ledger. The report doesn't ask to be trusted — it shows its working and scores its own past calls, every edition. The essay version of this argument: Reliable Reports Are a Systems Problem — why the writer goes last.

1How are the signals defined?

Every signal is a named boolean predicate over a flat metrics dict — a key, an operator, and a threshold. Nothing is discretionary at firing time. The full catalog is machine-introspectable: derived_signals(schema=True) returns exactly what is rendered below (53 signals, 57 accepted input aliases, 16 contradictions).

SignalFires whenMeaning
Refined products & cracks
crack_extremecrack_321_z > 3.03-2-1 crack z-score above 3 (z window: 2020–)
gasoline_crack_extremecrack_gasoline_z > 3.0Gasoline crack z-score above 3
diesel_crack_extremecrack_diesel_z > 3.0Diesel crack z-score above 3
gasoline_oversupplied / _undersuppliedgasoline_seasonal_pctl > 90 / < 20Gasoline stocks vs 5-year same-week percentile
distillate_oversupplied / _undersupplieddistillate_seasonal_pctl > 90 / < 20Distillate stocks vs seasonal percentile
refinery_strong / refinery_weakrefinery_util > 92 / < 85US refinery utilization extremes. Unconditional levels, not seasonally adjusted — so these fire with the turnaround calendar (weak in spring/autumn, strong in mid-summer) as well as with genuine market states. The refining family is the one group still on raw levels rather than seasonal percentiles.
maintenance_detectedmaintenance_trough_z < −2.0Structural turnaround trough in refinery runs — expected to co-fire with refinery_weak through the spring and autumn maintenance seasons
Crude supply
crude_oversupplied / _undersuppliedcrude_seasonal_pctl > 80 / < 20US crude stocks (ex-SPR) vs seasonal percentile
cushing_full / cushing_tightcushing_z > 2.0 / < −2.0Cushing (WTI delivery point) inventory z-score
cushing_building_fastcushing_7d_pct > 5.0Cushing built more than 5% in 7 days
production_at_seasonal_highproduction_seasonal_pctl ≥ 99US crude production at or above the 99th percentile of the same-week distribution — a seasonal high, NOT an all-time record (renamed 2026-08-11: it fired at 13,804 kb/d while the all-time weekly high was 13,862). A state descriptor on a secularly rising series, so it stays true for long stretches; not an event
Market structure
steep_backwardation / steep_contangom1_m12 < −10 / > 512-month curve, M12−M1 ($/bbl; negative = backwardation — the key name is legacy). Below −$10 the prompt sits $10+ over the 12-month; above +$5 is contango.
punishing_carryroll_yield_pct < −1512-month roll yield (M12/M1 − 1, negative in backwardation) worse than −15% — the curve pays the market not to store.
freight_extreme / freight_collapsedbwet_z > 3.0 / < −2.0Crude tanker freight z-score extremes. BWET is the Breakwave Tanker Shipping ETF (NYSE Arca, 2023 inception) — a freight-futures wrapper settling against Baltic assessments, not a Baltic index; ~3 years of history, so these thresholds sit on a short sample.
spread_wide / spread_invertedbrent_wti_zscore > 2.0 / spread < 0Brent–WTI spread stretched, or Brent below WTI
crude_price_extremewti_price_pctl > 95WTI flat price above its 95th historical percentile
bubble_detectedbubble_gsadf_excess > 0GSADF test statistic exceeds its critical value (explosive behavior)
Positioning & macro
wti_specs_extended / wti_specs_short / wti_positioning_medianwti_mm_pctl > 90 / < 10 / 30–70WTI managed-money positioning percentile (CFTC). Read as risk context — crowding and squeeze vulnerability — not as a directional call; why.
brent_specs_extended / brent_specs_short / brent_positioning_medianbrent_mm_pctl > 90 / < 10 / 30–70Brent managed-money positioning percentile (ICE Futures Europe COT, not CFTC)
vix_oil_supply_shockvix_corr > 0.5VIX–WTI correlation positive — a co-movement pattern consistent with a supply shock, which does not by itself identify one; the name is shorthand
Cross-Atlantic (EU/US composition)
eu_oil_stocks_tight / _comfortableeu_oil_seasonal_pctl < 20 / > 80EU27 oil stocks vs seasonal percentile
eu_gas_stocks_tight / _comfortableeu_gas_seasonal_pctl < 20 / > 80EU27 gas stocks vs seasonal percentile
nl_stocks_tightnl_oil_seasonal_pctl < 20Netherlands national oil stocks vs seasonal percentile — a national aggregate on the EU reporting cycle, used as a proxy for the ARA region. It is not the weekly Insights Global independent tank survey that "ARA stocks" usually means on a European desk.
eu_oil_imports_declining / _surgingeu_oil_imports_3m_pct < −10 / > 10EU27 oil import flow shift over 3 months
eu_oil_exports_risingeu_oil_exports_3m_pct > 10EU27 product re-export flows rising
eu_electricity_price_extremede_electricity_price_z > 2.0German power price extreme — demand-destruction / energy-crisis signal
eu_renewables_acceleratingeu_renewables_yoy_change > 2.0EU27 renewable share up >2pp YoY — long-term displacement
Agency layer (OPEC MOMR · IEA OMR)
opec_demand_upgraded / _downgradedopec_world_demand_revision > 0.1 / < −0.1OPEC month-over-month world demand revision (mb/d)
opec_production_gapopec_call_production_gap > 2.0Call on crude minus reported production > 2 mb/d — an implied stock draw on OPEC's own balance, not a measured shortfall. Read the aggregates before trading it: the MOMR's call is on DoC crude (the wider Declaration of Cooperation group), so pairing it with an OPEC-only production figure would overstate the gap by construction.
opec_call_risingopec_call_qoq_change > 0Call on DoC crude up quarter-over-quarter — sequential, not seasonally adjusted, so it fires through the normal Q1→Q3 demand build most years
opec_compliance_gap_extremeopec_compliance_gap_max_kbd > 200Largest direct-communication vs secondary-source production divergence across members > 200 kb/d — a reporting divergence. Quota behaviour is one possible cause; condensate boundary definitions and months where a member does not submit are others, and the threshold is chronically satisfied by a couple of members
opec_iea_spread_widening / _invertedopec_iea_demand_spread_2026 > 1.5 / < −0.5OPEC vs IEA demand-forecast divergence (mb/d)
Weather
weather_cold_extreme / weather_mild_winterhdd_anomaly_z > 2.0 / < −2.0US heating degree days vs 10-year average
weather_hot_extremecdd_anomaly_z > 2.0US cooling degree days extreme — gasoline + power-gen demand

Reading the catalog honestly

Contradictions — the headline layer

Single signals describe; contradictions — named combinations that a desk reads as one thesis-level statement, whether the legs pull against each other (refiner_margin_trap, freight_locked_arb) or reinforce (global_stocks_tight) — are what a desk actually argues about, so they are the report headlines. Each has a required signal set, optional exclusions, a score (average of component scores), and a fixed narrative. Fourteen of the sixteen require two or more signals; two (diesel_clean_long, flow_freight_divergence) are a single required signal with an exclusion. (us_comfortable_eu_tight was rebuilt 2026-08-11: it previously paired one requirement with an exclusion, which together with global_stocks_tight exhaustively partitioned a boolean — a guaranteed headline whenever the EU leg fired. Its US leg is now an assertion, crude_oversupplied.) A sample:

ContradictionRequiresThe reading
regime_curve_destructionsteep_backwardation + punishing_carryBackwardation steep, roll yield punishing — storage economics destroyed. Inventory holders liquidate into the squeeze, deepening it. Its two legs are near-algebraic restatements of each other (they cross at roughly $67 flat), so it counts as one observation about the curve, not two independent ones.
fundamental_rallyspread_wide + wti_positioning_medianPrices/spread at extremes but spec positioning at median — the rally is fundamental, not crowded. (The stretched leg is the Brent–WTI location spread, not flat price; and median positioning is also consistent with short-covering already done.)
freight_locked_arbspread_wide + freight_extremeSpread wide while freight sits at an extreme — watch freight alongside the spread, because a wide spread with expensive shipping may not be an executable arb. Whether it is actually shut depends on the spread in $/bbl against the freight rate; this conjunction compares z-scores, not economics.
supply_shock_signaturevix_oil_supply_shock + freight_extremeVIX–WTI correlation positive while tanker freight is at an extreme — the pattern associated with supply-side shocks rather than demand panic. A PATTERN MATCH, not an identification: neither leg observes a supply disruption, and the freight leg is an ETF measuring cost, not availability.
diesel_clean_longdiesel_crack_extreme − distillate_oversuppliedDiesel crack extreme without inventory overhang — the cleanest long signal in the complex.
global_stocks_tightcrude_undersupplied + eu_oil_stocks_tightUS crude AND EU oil stocks both in the bottom fifth of their seasonal ranges — tightness on both sides of the Atlantic at once, so a wider arb cannot relieve one from the other. Not tested here: freight, whether cargoes actually move, or flat price.
opec_revision_vs_priceopec_demand_downgraded + crude_price_extremeOPEC cut demand while price sits at extremes — bearish divergence; the biggest demand-side bull turned cautious.

How the contradiction layer is scored — and why it mostly isn't. Most contradictions combine conditions that are individually uncommon (2σ z-thresholds, or 20th/80th/90th-percentile cuts), so the joint base rate keeps the sample in single digits however long the history runs. The two single-signal-plus-exclusion forms are scoreable to the same extent as their required signal, and inherit its sample. So the signal layer is scored against its own firings; the contradiction layer is largely argued. Each one is named, its component signals are published, and its reading is fixed in advance — so it can be disagreed with precisely, which is the honest alternative to a statistic that will never exist.

Full catalog — all 53 predicates, the 57-alias input map, and all 16 contradiction definitions — at /signals (generated from the running server) or via derived_signals(schema=True) on any connected client.

2Can I inspect the methodology?

Yes — the rules the reports operate under are written down, versioned, and enforced by the tooling rather than by good intentions.

Backtest rules (signal_backtest)

Every claimed edge is an event study over the signal's own historical firings, under always-on honesty guards:

Revision policy

3Can I reproduce the numbers?

Every quantitative claim in a report traces to a named tool call against versioned data.

4How well do they calibrate?

The honest answer today: being measured — first resolution August 2026. The machinery is running and the ledger is open, but one unresolved entry is not a track record. This section is a promissory note with a date on it; it becomes evidence a year from now, and the reader is entitled to hold it to that.

The calibration ledger OPENED 24 JUL 2026  scores every edition's scenario probabilities. Priors are frozen at publication and never restated; each scenario set resolves against the realized Brent front-month at the stated horizon and receives a multi-category Brier score. Uniform guessing scores 0.667 on a 3-way set — that is a floor, not a skill bar: the honest comparison is the climatological base rate (how often that outcome occurs anyway), and resolved entries will be reported against both, because beating a forecaster who refuses to look at history proves nothing. Replayed or hindsight entries are marked as such and never Brier-scored. Entry 1 — the July 2026 update's 50/30/20 scenario prior — RESOLVES 24 AUG 2026

The Report performance panel

Every edition ends by scoring the previous one, under one non-negotiable condition: scoring is pre-registered, never retrospective. Trigger-based calls score mechanically against thresholds frozen at publication; prose claims are editorially scored and flagged as such; and every new call ships with its scoring rule attached, so the next edition cannot choose how to grade this one. In the panel's first outing, the mechanical calls held — and none of the three editorially-scored prose claims survived the backtest as stated. That result was printed.

See it in action

The demo reports below were produced under exactly these rules — revision trails, appendices, and performance panels included.

Demo Reports    Analyst Workflows