The Meta Level

The pipeline

How the layers fit together. One directory per layer; this ties them.

How the layers fit together. One directory per layer; this ties them.

AREA MAPPING --> GEOLOGY --> WATER --> DEMAND --> [SCHEME] --> [OPTIMISATION]
layer directory models status
Area mapping (in core/, build_catchments.py) terrain, catchments, rivers, sites works
Geology / aquifer papers/aquifers what the rock is, where, how much it holds built
Water papers/water how water arrives, moves, leaves built, holes named
Demand papers/demand who needs how much, and when 3 chapters built — licensed range, daily series, useful-water tiers
Scheme papers/scheme a scheme imposed on the natural layers 8 chapters built; one open defect on portfolios
Optimisation - which scheme is best not started

Demand is its own layer, and it sits before the scheme. It was previously one line inside a scenario file - a flat 3 Mm3/yr with no source - which is not enough to size anything against. The reason it needs a layer of its own is a counterfactual: what was abstracted in a drought year is what was available, not what was needed, so demand records measure the shortage rather than the requirement. It will be built from the abstraction licence register plus modelling of what would have been used had water been there, with environmental demands layered on top rather than reduced to a single threshold.

The first three layers are NATURAL. They describe how the place works with nobody intervening. That baseline has to exist before a scheme can be assessed, because a scheme is a perturbation of it.

Within the scheme layer the order is also fixed, and it is not the obvious one: availability -> aquifer -> reservoir. A reservoir is a buffer, not a purpose; its volume is the residual left when the river and the ground cannot match supply to demand between them. Sizing it first means picking a number and then building a scheme to justify it. See papers/scheme.


The original scoping note

Scope of the whole chain, end to end. Written 2026-08-11.

The point of this document is NOT that every stage works. Most do not yet. The point is that the chain is named end to end, so that effort can be aimed at whichever link carries the most uncertainty rather than at whichever link is most interesting. A pipeline with stated assumptions beats a subset with hidden ones.

Read the last section first if you read nothing else.


The chain

# stage what turns into what status
1 Rock identity map colour -> named formation works
2 Rock geometry formation -> located bodies, area, extent works
3 Aquifer properties body -> thickness, Sy, T weak
4 Depth & saturation body -> saturated thickness partial
5 Storage volume the above -> drainable Mm3 bounded, not known
6 Rainfall gauges -> areal daily per catchment works
7 Soil moisture rain -> AET, effective rainfall partial
8 The split effective rain -> runoff vs recharge MISSING
9 Runoff routing runoff -> river flow at a point partial
10 River stage flow -> level at any point MISSING
11 GW <-> river exchange head difference -> baseflow, signed structure only
12 Diversion flow + constraints -> divertible volume exists
13 Surface storage divertible -> reservoir stage/volume works
14 Injection divertible -> aquifer head rise rate-limited, seasonal headroom, settling
15 Recovery head -> abstractable rate, sustained pump- and buffer-limited
16 Return flows leakage -> river baseflow credit assumed 80%, swept: 1-3% effect
17 Flood attenuation storage -> peak reduction downstream partial
18 Tidal boundary river -> sea, through gates MISSING

Stage notes, and the assumption each one rests on

1-2 Rock. 87% of the 50k map is a named formation; 335 located aquifer bodies within 40 km of the site box, each with a grid reference and the point where it comes nearest. Solid. Assumption: the 625k maps aquifers AT OUTCROP, so depth-to-top is 0 by construction. That is the source’s assumption, not a measurement.

3 Properties. 9 units of 32 have parameters of their own; the rest share one fracture-flow default. Chalk Sy is now analogue from WD/97/34 Table 4.1.6 rather than a guess. Assumption: thickness is unread for most units. It is now the binding constraint on the Chalk, having replaced Sy.

4-5 Volume. 15,801 Mm3 drainable, range 2,033-92,140. 36% of rock is dry once the water table is applied. Assumption, and it is load-bearing: drainable is a CEILING NOBODY CAN REACH. Water moves ~16 m through the Sherwood in the 229-day lag, so recovery is set by well count and placement, not by polygon area. The volume model deliberately stops one step short of “accessible” so the refusal is visible.

6 Rainfall. Areal, Thiessen-weighted, per catchment, 2010-. Rebuilt from 2012 with --start 2010-01-01: the automatic common-period rule chose 2012 by a margin of three gauges (65 vs 62), and two extra years containing the worst drought in the record beat three gauges out of sixty-five. The 2010-11 areal rainfall therefore rests on a slightly thinner network than the rest.

7-8 THE SPLIT. Soil-bucket overflow currently ALL becomes recharge. There is no quick-runoff pathway. This is the largest structural hole in the chain and it propagates: it makes recharge an upper bound, and since head identifies only recharge/Sy, it makes fitted Sy an upper bound too.

9-10 Routing and stage. Flow is modelled at gauged outlets; level is not modelled anywhere. 33 stations of 15-minute stage are now on disk to fix this, including 7 sites where a flow gauge sits inside a RECORDED flood outline - the strongest calibration evidence available.

11 Exchange. Built SIGNED - k*(h - h_base), so direction falls out of state and the same water cannot be counted as both baseflow and storage. h_base is a calibrated constant until stage 10 lands, at which point it becomes a series and nothing else changes.

13 Surface storage. The most complete part of the model. Reservoir stage-volume curves from LiDAR, sited candidates, runs 00-06.

14-15 Injection and recovery. Not built, and NOT VALIDATABLE from the record: nothing in 14 years of observation contains an injection. Amplitude bounds it - an aquifer that naturally swings 12 m will not absorb a 50 m mound without pushing back - but that is a bound, not a model.

16 Return flows. 80% of leakage is credited back as river baseflow. Still an assumption with no source - but it has now been SWEPT from 0 to 1 across every scheme configuration (papers/scheme/synthesis.md) and moves each answer by only 1-3%, changing no ranking. It should still become a routed quantity once stage 10 exists; it is no longer a risk to any conclusion.

18 TIDAL BOUNDARY - THE MISSING END OF THE PIPE. The Levels drain to the Parrett estuary and the Bristol Channel through gates and sluices that can only discharge when the tide is low enough. Drainage capacity here is TIDE-LIMITED, not channel-limited: the Levels flood partly because at high tide they cannot drain at all, whatever the channel could carry. Huntspill Sluice - which dominated our own data-quality flags precisely because it is operated - is one of these structures. Nothing in this project models the sea end. A flood attenuation claim that ignores the tidal gate is a claim about a system with an open outlet, which this one does not have.


Where the uncertainty actually is

Ranked by how much of the final answer each governs, not by how uncertain it is in itself. A wide error bar on something small is cheap; a hidden structural omission is not.

rank uncertainty kind why it ranks here
1 the runoff/recharge split (8) structural governs recharge AND Sy AND baseflow. Everything downstream inherits it
2 the tidal boundary (18) structural the pipe has no modelled end. Caps every attenuation claim
3 injection/recovery (14-15) structural the scheme’s entire premise, and unvalidatable from the record
4 accessible vs drainable (5) conceptual a factor of unknown size between a ceiling and a yield
5 unit thicknesses (3) parameter 5,179 Mm3 on Bracklesham alone, all guess
6 datum mixing (10) data mAOD vs local stage, would corrupt silently
7 station coverage data the prime candidate aquifer has ONE monitoring borehole

The top three are structural, not parametric. More data does not fix them; building the missing stage does. That is the argument for scoping the pipe before refining any single number in it.


Deferred, deliberately

Small, known, and not worth blocking on:

  • despike screens - the datum test works (3 mAOD, 19 stage, 13 ambiguous); the stuck-sensor test is measuring sensor RESOLUTION not faults (65% false positives) and the structure test missed Kenn Pier because name-matching cannot see behaviour. Use only the 22 unambiguous stations until this is redone.
  • Bracklesham / Upper Greensand / Great Oolite thicknesses
  • per-aquifer water-table fits below the n>=8 threshold
  • Ilchester as a stage-flow pair - it is 3.4 km from its gauge

The layer being built now: water table / flood water / river

Stages 7-11 and 17. Known holes and their known fixes, so the next session starts from the list rather than rediscovering it.

hole consequence known fix size
no runoff/recharge split recharge and Sy are both upper bounds; only their ratio (~29.5 m/yr) is identified partition bucket overflow, calibrate against flood peaks now that 15-min stage exists the main job
h_base is a constant GW<->river exchange is a signed linear reservoir, not a coupling substitute modelled stage; ONE line, structure already signed small
13 of 35 stations ambiguous datum mAOD and local stage cannot be pooled use local MINIMUM elevation in a window, and the record’s RANGE (stage rarely spans >5 m) small
regional-mean rainfall in the coupled model spatial detail absent from recharge assign boreholes to catchments - the areal series already exist small
prime candidate has 1 monitoring borehole no fitted response for the Sherwood more stations, or accept it is uncalibrated and say so data-limited
sub-daily peaks vs daily model a daily model under-fits peaks; tuning runoff to compensate hides a timestep error in a physical parameter run the peak calibration sub-daily; the 15-min data is on disk medium
no injection anywhere in the record stage 14-15 unvalidatable amplitude BOUNDS it (~12 m natural swing); present as bounded, never predicted unfixable from data

Screens repaired 2026-08-11: the stuck-sensor test was measuring sensor quantisation (65% false positives, now 7%) and structure detection was name-only (missed Kenn Pier, the second-largest flag source). Both now behave. Datum classification was never broken, only incomplete - 22 of 35 stations are usable as-is.


← All notes · More from Aquifer Storage