Technology

Data, models, and methods.

The data pipelines, valuation models, and machine-learning program behind our siting and power-cost work.

01  The data foundation

We maintain pipelines over ERCOT settlement data: day-ahead and real-time prices at the individual settlement-point and bus-node level, sourced from commercial market-data providers and ERCOT's public datasets. When we look at a project site, we use the price history of that specific settlement point, hour by hour, across years of market history, including scarcity and congestion events.

We also cross-check node locations against ERCOT's settlement-point registry to confirm each node's location before it is used in analysis.

02  The valuation engine

Our valuation models simulate project revenue hourly against historically grounded price scenarios, preserving volatility structure, seasonality, and tail events, and roll the results into cashflow forecasts and present-value analysis.

  • Scenario-based: results are reported as ranges across scenarios rather than a single path.
  • Basis-aware: the spread between a node and its hub is modeled explicitly.
  • Auditable: forecasts trace back to observable market history.

03  Hedges & revenue collars

Fixed-price swaps, floors, caps, and collar structures are priced against our simulated hourly revenue scenarios. Outputs include expected payments under each structure across scenarios, and the risk remaining at the node after a contract settles at the hub.

  • Hedges: swap and offtake structures evaluated across simulated outcomes, including residual basis risk.
  • Revenue collars: floor-and-cap structures modeled hourly, quantifying expected payments under the floor and revenue forgone above the cap.

04  Capacity expansion

Long-lived assets are exposed to changes in the grid around them. We model buildout scenarios for new generation, storage, and load entering the system, and estimate the effect of each on congestion, basis, and price shape at the nodes we cover.

05  The machine-learning program

The program under development combines two layers: traditional capacity expansion modeling to produce long-run power price scenarios, and machine-learning ingestion models that keep the expansion model's assumptions current.

Capacity expansion models forecast prices by simulating what gets built: generation, storage, and load, dispatched against transmission constraints. Their outputs are only as good as their assumptions, and those assumptions change monthly. Projects advance or die in the interconnection queue, and market rules get rewritten. Updating the assumptions by hand is slow, and most of the source material is unstructured text.

The ingestion engine reads that text:

  • Developer planning materials: interconnection queue filings, planning documents, and project constraints. Which projects are advancing, where they connect, at what size, and on what timeline.
  • Market policy and rules: ERCOT protocol revisions, regulatory filings, and legislation, including upcoming filings and Texas senate bills, mapped to the model parameters they affect.

The output is structured buildout and policy assumptions that feed the expansion model directly, so price scenarios reflect current buildout and market rules. Alongside this, we are developing GPU-trained models for probabilistic price forecasting at the node level: distributions of outcomes per settlement point rather than single estimates.

Both workstreams are in active development and prototyping; they are not yet in production.

See the pipeline design →

Principles

  • Probabilistic outputs. Forecasts are reported as ranges, not single numbers.
  • Nodal resolution. Projects settle at nodes, so analysis is done at nodes.
  • Backtesting. Forecasts are compared against realized prices.

Have a site to evaluate?

Send us the settlement point and we'll take a look.

info@barkerinfra.com