The pipeline
How the layers fit together. One directory per layer; this ties them.
How the layers fit together. One directory per layer; this ties them.
AREA MAPPING --> GEOLOGY --> WATER --> DEMAND --> [SCHEME] --> [OPTIMISATION]
| layer | directory | models | status |
|---|---|---|---|
| Area mapping | (in core/, build_catchments.py) |
terrain, catchments, rivers, sites | works |
| Geology / aquifer | papers/aquifers |
what the rock is, where, how much it holds | built |
| Water | papers/water |
how water arrives, moves, leaves | built, holes named |
| Demand | papers/demand |
who needs how much, and when | 3 chapters built — licensed range, daily series, useful-water tiers |
| Scheme | papers/scheme |
a scheme imposed on the natural layers | 8 chapters built; one open defect on portfolios |
| Optimisation | - | which scheme is best | not started |
Demand is its own layer, and it sits before the scheme. It was previously one line inside a scenario file - a flat 3 Mm3/yr with no source - which is not enough to size anything against. The reason it needs a layer of its own is a counterfactual: what was abstracted in a drought year is what was available, not what was needed, so demand records measure the shortage rather than the requirement. It will be built from the abstraction licence register plus modelling of what would have been used had water been there, with environmental demands layered on top rather than reduced to a single threshold.
The first three layers are NATURAL. They describe how the place works with nobody intervening. That baseline has to exist before a scheme can be assessed, because a scheme is a perturbation of it.
Within the scheme layer the order is also fixed, and it is not the
obvious one: availability -> aquifer -> reservoir. A reservoir is a
buffer, not a purpose; its volume is the residual left when the river
and the ground cannot match supply to demand between them. Sizing it
first means picking a number and then building a scheme to justify it.
See papers/scheme.
The original scoping note
Scope of the whole chain, end to end. Written 2026-08-11.
The point of this document is NOT that every stage works. Most do not yet. The point is that the chain is named end to end, so that effort can be aimed at whichever link carries the most uncertainty rather than at whichever link is most interesting. A pipeline with stated assumptions beats a subset with hidden ones.
Read the last section first if you read nothing else.
The chain
| # | stage | what turns into what | status |
|---|---|---|---|
| 1 | Rock identity | map colour -> named formation | works |
| 2 | Rock geometry | formation -> located bodies, area, extent | works |
| 3 | Aquifer properties | body -> thickness, Sy, T | weak |
| 4 | Depth & saturation | body -> saturated thickness | partial |
| 5 | Storage volume | the above -> drainable Mm3 | bounded, not known |
| 6 | Rainfall | gauges -> areal daily per catchment | works |
| 7 | Soil moisture | rain -> AET, effective rainfall | partial |
| 8 | The split | effective rain -> runoff vs recharge | MISSING |
| 9 | Runoff routing | runoff -> river flow at a point | partial |
| 10 | River stage | flow -> level at any point | MISSING |
| 11 | GW <-> river exchange | head difference -> baseflow, signed | structure only |
| 12 | Diversion | flow + constraints -> divertible volume | exists |
| 13 | Surface storage | divertible -> reservoir stage/volume | works |
| 14 | Injection | divertible -> aquifer head rise | rate-limited, seasonal headroom, settling |
| 15 | Recovery | head -> abstractable rate, sustained | pump- and buffer-limited |
| 16 | Return flows | leakage -> river baseflow credit | assumed 80%, swept: 1-3% effect |
| 17 | Flood attenuation | storage -> peak reduction downstream | partial |
| 18 | Tidal boundary | river -> sea, through gates | MISSING |
Stage notes, and the assumption each one rests on
1-2 Rock. 87% of the 50k map is a named formation; 335 located aquifer bodies within 40 km of the site box, each with a grid reference and the point where it comes nearest. Solid. Assumption: the 625k maps aquifers AT OUTCROP, so depth-to-top is 0 by construction. That is the source’s assumption, not a measurement.
3 Properties. 9 units of 32 have parameters of their own; the rest
share one fracture-flow default. Chalk Sy is now analogue from
WD/97/34 Table 4.1.6 rather than a guess.
Assumption: thickness is unread for most units. It is now the
binding constraint on the Chalk, having replaced Sy.
4-5 Volume. 15,801 Mm3 drainable, range 2,033-92,140. 36% of rock is dry once the water table is applied. Assumption, and it is load-bearing: drainable is a CEILING NOBODY CAN REACH. Water moves ~16 m through the Sherwood in the 229-day lag, so recovery is set by well count and placement, not by polygon area. The volume model deliberately stops one step short of “accessible” so the refusal is visible.
6 Rainfall. Areal, Thiessen-weighted, per catchment, 2010-.
Rebuilt from 2012 with --start 2010-01-01: the automatic
common-period rule chose 2012 by a margin of three gauges (65 vs 62),
and two extra years containing the worst drought in the record beat
three gauges out of sixty-five. The 2010-11 areal rainfall therefore
rests on a slightly thinner network than the rest.
7-8 THE SPLIT. Soil-bucket overflow currently ALL becomes
recharge. There is no quick-runoff pathway. This is the largest
structural hole in the chain and it propagates: it makes recharge an
upper bound, and since head identifies only recharge/Sy, it makes
fitted Sy an upper bound too.
9-10 Routing and stage. Flow is modelled at gauged outlets; level is not modelled anywhere. 33 stations of 15-minute stage are now on disk to fix this, including 7 sites where a flow gauge sits inside a RECORDED flood outline - the strongest calibration evidence available.
11 Exchange. Built SIGNED - k*(h - h_base), so direction falls
out of state and the same water cannot be counted as both baseflow and
storage. h_base is a calibrated constant until stage 10 lands, at
which point it becomes a series and nothing else changes.
13 Surface storage. The most complete part of the model. Reservoir stage-volume curves from LiDAR, sited candidates, runs 00-06.
14-15 Injection and recovery. Not built, and NOT VALIDATABLE from the record: nothing in 14 years of observation contains an injection. Amplitude bounds it - an aquifer that naturally swings 12 m will not absorb a 50 m mound without pushing back - but that is a bound, not a model.
16 Return flows. 80% of leakage is credited back as river
baseflow. Still an assumption with no source - but it has now been
SWEPT from 0 to 1 across every scheme configuration
(papers/scheme/synthesis.md) and moves each answer by only 1-3%,
changing no ranking. It should still become a routed quantity once
stage 10 exists; it is no longer a risk to any conclusion.
18 TIDAL BOUNDARY - THE MISSING END OF THE PIPE. The Levels drain to the Parrett estuary and the Bristol Channel through gates and sluices that can only discharge when the tide is low enough. Drainage capacity here is TIDE-LIMITED, not channel-limited: the Levels flood partly because at high tide they cannot drain at all, whatever the channel could carry. Huntspill Sluice - which dominated our own data-quality flags precisely because it is operated - is one of these structures. Nothing in this project models the sea end. A flood attenuation claim that ignores the tidal gate is a claim about a system with an open outlet, which this one does not have.
Where the uncertainty actually is
Ranked by how much of the final answer each governs, not by how uncertain it is in itself. A wide error bar on something small is cheap; a hidden structural omission is not.
| rank | uncertainty | kind | why it ranks here |
|---|---|---|---|
| 1 | the runoff/recharge split (8) | structural | governs recharge AND Sy AND baseflow. Everything downstream inherits it |
| 2 | the tidal boundary (18) | structural | the pipe has no modelled end. Caps every attenuation claim |
| 3 | injection/recovery (14-15) | structural | the scheme’s entire premise, and unvalidatable from the record |
| 4 | accessible vs drainable (5) | conceptual | a factor of unknown size between a ceiling and a yield |
| 5 | unit thicknesses (3) | parameter | 5,179 Mm3 on Bracklesham alone, all guess |
| 6 | datum mixing (10) | data | mAOD vs local stage, would corrupt silently |
| 7 | station coverage | data | the prime candidate aquifer has ONE monitoring borehole |
The top three are structural, not parametric. More data does not fix them; building the missing stage does. That is the argument for scoping the pipe before refining any single number in it.
Deferred, deliberately
Small, known, and not worth blocking on:
- despike screens - the datum test works (3 mAOD, 19 stage, 13 ambiguous); the stuck-sensor test is measuring sensor RESOLUTION not faults (65% false positives) and the structure test missed Kenn Pier because name-matching cannot see behaviour. Use only the 22 unambiguous stations until this is redone.
- Bracklesham / Upper Greensand / Great Oolite thicknesses
- per-aquifer water-table fits below the n>=8 threshold
- Ilchester as a stage-flow pair - it is 3.4 km from its gauge
The layer being built now: water table / flood water / river
Stages 7-11 and 17. Known holes and their known fixes, so the next session starts from the list rather than rediscovering it.
| hole | consequence | known fix | size |
|---|---|---|---|
| no runoff/recharge split | recharge and Sy are both upper bounds; only their ratio (~29.5 m/yr) is identified | partition bucket overflow, calibrate against flood peaks now that 15-min stage exists | the main job |
h_base is a constant |
GW<->river exchange is a signed linear reservoir, not a coupling | substitute modelled stage; ONE line, structure already signed | small |
| 13 of 35 stations ambiguous datum | mAOD and local stage cannot be pooled | use local MINIMUM elevation in a window, and the record’s RANGE (stage rarely spans >5 m) | small |
| regional-mean rainfall in the coupled model | spatial detail absent from recharge | assign boreholes to catchments - the areal series already exist | small |
| prime candidate has 1 monitoring borehole | no fitted response for the Sherwood | more stations, or accept it is uncalibrated and say so | data-limited |
| sub-daily peaks vs daily model | a daily model under-fits peaks; tuning runoff to compensate hides a timestep error in a physical parameter | run the peak calibration sub-daily; the 15-min data is on disk | medium |
| no injection anywhere in the record | stage 14-15 unvalidatable | amplitude BOUNDS it (~12 m natural swing); present as bounded, never predicted | unfixable from data |
Screens repaired 2026-08-11: the stuck-sensor test was measuring sensor quantisation (65% false positives, now 7%) and structure detection was name-only (missed Kenn Pier, the second-largest flag source). Both now behave. Datum classification was never broken, only incomplete - 22 of 35 stations are usable as-is.