Skip to content
A datacentre aisle, racks receding into light

AI Infrastructure

Capacity you provision for is capacity somebody has to predict.

The operation

Inference load is not smooth. It is diurnal, bursty, and it saturates against a ceiling you already paid for. The expensive failure is rarely the spike itself. It is provisioning for a spike that never arrives, every day for a quarter, or missing the one that did.

What changes

Forge reads the request load the fleet already produces and projects it forward, so the capacity decision is made against a number instead of a nerve.

Tested here

Tested on a production Azure Functions trace, a 300-function cohort at five-minute grain. Floor skill of +12.0 to +21.5% over persistence, and breach-risk Brier of 0.031 to 0.035 against a 0.049 base rate. This was the industry we predicted would fit before we had the data to check.