INCA Revision evidence & claims

swipe left or right to change slide

Round-1 review · two reviewers, 22 numbered comments

What the reviewers asked, and what we ran

Click a row for the comments it covers and the experiments behind it. Six themes account for every critical comment; the rest are wording, availability and citations.

~900new calibration runs
8–15repeats per configuration
$605API cost of the revision
9claims removed or narrowed
Internal report · 16 September 2026

What the revised paper claims, and what backs it

Four claims replace the submitted version's framing, in the order the evidence is read: what happens, what carries it, where it scales, and why. Click a card to jump; on every slide the chart is the argument, underlined terms explain how an arm was run, and Detail holds the caveats.

2 basinsten parameters, 2018 calibration, 2019 hold-out
2 basinsmulti-site, 50 and 100 parameters
8–15repeats per configuration
6 + 1classical optimizers plus a rule-based twin
Claim 1 · sample efficiency

A usable calibration in tens of evaluations, not hundreds

click a legend entry to hide a method

INCA crosses NSE 0.7 first on both basins, in 8 of 8 repeats, then stops itself; classical optimizers overtake its calibration fit once hundreds of evaluations are affordable.

Detail · how the axis is counted, and what changed

The x-axis is CREST forward evaluations actually executed, not agent iterations. Because candidates are simulated in parallel batches, every runs-to-threshold number is a batch-consistent count. The submitted version plotted INCA rounds against DREAM model calls, which overstated the gap; the showcase calibration used 65 evaluations, not 24.

Claim 1 · sample efficiency

Four times fewer evaluations than the best classical method

Detail · hold-out behaviour is basin-dependent

In Georgia INCA (69 evaluations), EKI (500) and DREAM (10,662) are indistinguishable out of sample and DREAM's in-sample advantage disappears; SCE-UA stays ahead. In California, where 2018 was dry and 2019 wet, INCA and the 500-evaluation optimizers all lose 0.4–0.5 NSE out of sample while DREAM loses 0.16. No universal generalization advantage is claimed.

Claim 2 · mechanism

Semantics and feedback, not parameter names and not resampling

Bars are medians over repeats and every repeat is drawn as a dot; the dark mark is the same arm at 250 evaluations, where most have converged. Sample sizes differ by arm (8 for most INCA configurations, 5 for the matched-seed pairs, 15 for the classical methods) and "one-shot pooled" is not a set of repeats but the median over 2,000 random subsets of the pooled candidates. Drag the slider to see the advantage appear and disappear.

Detail · what each arm changes
INCA, signatures / imagesfull two-stage workflow; only the perception channel differs.
Direct LLMsame context, history and failure feedback, no proposal/verification split.
One-shota single call, eight candidates, no result ever fed back.
One-shot pooledindependent single calls pooled to the same evaluation budget.
Names + definitionsreal names, each parameter defined, nothing about how it moves the hydrograph.
Names + inverted equationthe same, with the routing relation as EF5's documentation writes it, which is inverted.
Anon. labels + relationsx1…x13 with the parameter–process relations restated on those labels.
Anon. labels onlyx1…x13, no physical content at all.

Headline contrasts: pooling 64 one-shot proposals at a full calibration's budget reaches 0.64 against 0.77 for the loop in Georgia, and never reaches NSE 0 in California where the loop reaches 0.87. Anonymised labels with the relations perform like real names; without them the agent fails.

Claim 2 · mechanism

The staged architecture is a robustness device, not a performance boost

Detail · how to read this chart, and what a stalled repeat is

Four cells, one basin. Each pair of bars is one model under one workflow. The solid bar is the median of the repeats in which the agent saw rendered hydrograph and flow-duration figures, the outlined bar the median when the same series arrived as numerical signatures. Every repeat is plotted as a dot beside its own bar, because two of these distributions are two-humped and the median falls in the gap between the humps.

Stalled repeats, not crashes. No run in this figure failed to execute. A repeat below 0.5 is one where the loop spent its full budget without leaving a poor parameter region: it accepted an early update that moved the wrong way, then circled inside that region. These stay in the medians, which is why the Claude cells sit low.

What the architecture buys. For GPT-5 the staged workflow adds little. For Claude Sonnet 4.5 it converts a two-humped outcome into a reliable one, lifting the image median from 0.40 to 0.72 and halving the number of stalls. The framework runs unchanged on both models: the model sets the lower tail, not the ceiling, and the single best calibration of the study came from Claude.

The two input channels are within noise here. Images and signatures differ by at most 0.15 in any cell and the ranking of the four cells is the same on both. This basin does not separate them.

Claim 3 · setup

One parameter block per gauged sub-basin

Detail · what a parameter block is, how the gauges were chosen, and why only two basins

What a block is. EF5 already supports one parameter block per gauge: a parameter set can carry several gauge=<id> sections, and the model assigns every grid cell to the first gauge it meets walking downstream. Listing the outlet plus k interior gauges therefore cuts the basin into k+1 zones with no extra machinery, and each zone gets its own copy of the ten free multipliers, so the dimension is 10 × (k + 1): 100 on the Raritan, 50 on Alameda. We checked this numerically — giving every block identical values reproduces the single-gauge run exactly, and editing only the outlet block leaves every interior gauge's series bit-identical.

Whose solution the coloured maps show. Both are INCA, not a classical optimiser, and neither is calibrated gauge by gauge. Each is a single joint run against one objective, the mean NSE over all gauges, with every block free in the same round and acceptance decided on the mean. The zoned map is the 100-parameter run on the Raritan and the 50-parameter run on Alameda, about 180 forward evaluations; the lumped map is the same agent on ten shared multipliers, about 40 and 64. Each cell is the median over five seeds at that gauge, so the Δ column compares two medians rather than two matched runs, and the median vector is not itself a solution any single run produced.

The one per-zone step inside a round. Winner-take-all throws away most of what a round found: on the Raritan, seven of eight candidates improved Lamington in a typical round and all seven were discarded because the candidate with the best mean was the one that left that zone alone, since a zone worth a tenth of the objective cannot outvote the two zones holding 55 per cent of the basin. So each round also assembles one extra candidate by taking every zone's block from whichever candidate fitted that zone's own gauge best. It is simulated like any other candidate and competes on equal terms, never accepted on assumption, because the zones are nested and a reassembled vector is not guaranteed to be better.

Why the partition is nested. Water only moves downstream, so a zone's parameters change that zone's gauge and every gauge below it, and nothing above it. Each block is therefore constrained by a record of its own, which is what makes 100 parameters identifiable at all. The objective is the mean NSE over all gauges, not the outlet alone.

Eight basins were screened for interior gauges with continuous hourly records: drainage area between 3% and 100% of the outlet, both years covered, at least 95% daily coverage, a sub-daily variability check, a D8 nesting check with a snap radius, and a regulation check by within-basin relative outliers. That leaves one basin supporting 100 parameters and one supporting 50. A partition into ungauged units can be built, but every method's solutions then scatter across seeds like uniform random draws, so those parameters carry no information.

The two basins behave differently, which is the point: on the Raritan one gauge (Lamington) cannot be fitted with the parameters that fit the others, so the extra blocks pay; on Alameda one parameter set fits all five gauges in the calibration year, and the zoned solution only wins out of sample.

Claim 3 · scaling

An early-budget advantage on 50–100 spatial parameters

Before about 200 evaluations only INCA is above zero, including against every method given the same zone map; after about 500 those schemes catch up and pass it, and they generalise at least as well.

Detail · what each of the key three actually does, and the tied-versus-zoned test

Two of the three share one mechanism; the third does not. INCA and the rule table both cut the basin into zones, produce a batch each round, and then build one extra candidate by taking every zone's block from whichever member of that batch, the incumbent included, scored best at that zone's own gauge. The assembled candidate is simulated and competes on equal terms. The only thing that differs between them is who writes the moves: the language model or the rule table. That is the like-for-like control, and the gap between the red lines and the grey one is what the writing has to account for.

The EKI line is a different scheme entirely. It has no assembly step. It calibrates the ten lumped multipliers with all zones tied, then visits the zones one at a time, spends a fixed number of evaluations refining that zone's block with the others held, and keeps the visit if the objective improves. That is block coordinate descent. It appears in this view because it is the strongest non-agent method at this budget, not because it is mechanism-matched.

Why the zone-map baselines exist. The agent is told which parameters belong to which gauged sub-basin, so comparing it only with optimisers handed the 100 numbers as one flat vector would be unfair. The schemes a hydrologist would build with the same information are all here. With a large budget the per-zone rule table is the strongest method on both basins. It is simply very slow to get there.

The grey diamonds are the same agent on ten lumped parameters. The extra spatial parameters pay where the gauges disagree (Raritan: 0.50 against 0.34) and improve the hold-out year on both basins, but not the calibration fit on Alameda.

Claim 4 · why it is fast early

The early advantage comes from construction, not from search

Detail · the decomposition, and the matched control

What the three rows count, and why the denominators differ. In the blind twin states we know exactly which parameters were broken, so every proposal can be scored against them. Row 1 asks, across every pairing of a candidate with a parameter that was actually broken, how often that parameter was changed at all: 80 per cent for the agent, 49 for the rule table. Row 2 narrows to just those changes and asks whether the sign was right: 97 and 100 per cent. Row 3 changes unit and counts whole candidates, asking what share of them moved every broken parameter in the right direction: 70 per cent against 25.

The obvious objection, and the control for it. If the agent simply moves more knobs per candidate, then breadth is a property of how each batch was built and not a capability, and the rule table could match it by spending its eight slots differently. Half of that is right. Grouping candidates by how many parameters they move shows that coverage is almost entirely mechanical: among candidates that move five or more, the rule table covers the broken parameter 100 per cent of the time against the agent's 95. The agent does not pick better parameters, it picks more of them.

But breadth alone does not reproduce the result. Comparing only the wide candidates, paired inside each twin state so the difficulty is held fixed, the agent scores higher in 12 of 15 states with a median gap of 0.26 and a signed-rank p of 0.007 — and this is with the rule table covering the broken parameter more often and never once getting a direction wrong. The difference is in the rest of the vector: on the parameters that were fine, a wide rule-table candidate moves twice as far as a wide agent candidate, 2.12 against 1.08 in log units. It fixes what is broken and breaks what was not. Breadth is necessary and the rule table can have it; spending it without collateral is the part that does not come for free. Note also that the rule table was never capped — a quarter of its own candidates already move five or more parameters, and on the multi-site problem the explicitly widened variant does not close the gap either.

The simplest proof that the construction is joint. A table of independent per-parameter responses has a fingerprint: the size of a step is set by the tier the rule fired at, so every parameter a candidate moves gets the same step, whether or not that parameter is the one that is broken. In the rule table's candidates the step on the broken parameter and the step on an intact one are exactly equal 54 per cent of the time, and the median ratio is 1.01. In the agent's candidates they are equal 2 per cent of the time; the broken parameter is moved 2.4 times further, in 78 per cent of candidates, and the agent was never told which parameter was broken. That is what "this one a lot, that one a little" looks like in the data, and it is why the agent's wide candidates carry half the collateral of the rule table's. It is the same fact as the coverage result, seen from the magnitude side rather than the on-off side.

Row 2 is the important null. Neither method makes direction mistakes. The domain knowledge that tells you to raise soil storage when the volume is too high is in the rule table as well, and it works there. So the advantage is not knowing which way to turn a knob. It is whether the knob that matters is in the batch at all, which is row 1, and that follows from writing a median of five coordinated changes per candidate rather than two.

This is construction, not diagnosis. The agent is not better at naming which process broke; it is better at writing down a wide, internally consistent multi-parameter hypothesis in one pass. Breadth costs: movement on parameters that were not deficient lowers a candidate's score for both methods. It pays here because the per-parameter directions are reliable and because acceptance is decided by simulation, so the cost falls on the seven candidates that are discarded.

In one hundred dimensions the argument becomes structural. A hundred forward evaluations cannot sample a hundred-dimensional box, so nothing handed the flat vector works: the best classical optimiser reaches 0.04 and most are negative.

The control holds everything mechanical fixed. It is given the same zone map, the same eight candidates per round, and the agent's own per-zone assembly step, the one that rebuilds a candidate each round from whichever proposal fitted each zone's gauge best. The rule table writes the moves instead of the language model. Nothing else differs. After a hundred forward evaluations it is at −0.12 on the 50-parameter problem and 0.00 on the 100-parameter one, against 0.39 and 0.48 for the agent. The zone map and the assembly machinery are not what produce the early advantage; what gets written into each candidate is.

Why no other baseline is shown here. Schemes that calibrate a lumped version first and then distribute reach a similar number, but they never search the 50- or 100-dimensional space at all, so comparing them on a 100-parameter axis would be misleading. They are on the scaling slide, where the budget axis makes clear what they are doing.

The limit. This is an early-budget statement, and the same rule table that sits at 0.00 here is the strongest method of all at two thousand evaluations. The ordering reverses once there is enough budget to search, as the scaling slide showed.

Claim 4 · what construction looks like

One round of construction, read line by line

Detail · how to read a row, and what the colours mean

Each row is one of the eight candidates the agent produced in one round, before any of them was simulated. A round is two calls: a proposal call writes eight candidates, a verification call checks and edits them and attaches a one-sentence rationale to each. The sentence shown is that rationale, because the edited list is the one that was actually simulated. The four numbers are the multipliers as applied. The last column is the full-year NSE the candidate scored once CREST ran it.

Green is a move of a deficient parameter toward the hidden truth, red away from it, amber a move of a parameter that was never perturbed and therefore should not have moved.

What to look for is the column, not the row. Construction shows up as a consensus down the green columns with variation everywhere else: the batch agrees on the parameters that matter and disagrees about the rest. That is what a wide, internally consistent hypothesis looks like when it is written eight ways in one round, and it is why the deficient parameter is in play 80 per cent of the time rather than 49.

The four cases mark the edges. Construction can be right while the accompanying sentence is wrong. It can be irrelevant when the deficient parameter barely moves the flow. It is not free: breadth produces wrecks alongside the winner, which is why every candidate is simulated and only the best is kept. And in the fourth case the sentence and the numbers agree, and the rule table's eight candidates for the same state sit underneath for contrast.

Scope

Where this applies, and what we are not claiming

Most likely useful

  • Limited forward-evaluation budgets.
  • Each evaluation expensive relative to LLM inference.
  • Parameters with informative, correctly stated response semantics.
  • Structured diagnostic information available.

Not claimed

  • Not posterior inference or uncertainty quantification.
  • Not a guaranteed causal diagnosis.
  • Not universally faster in wall-clock.
  • Not a better final optimum.
  • Not demonstrated for O(1000) parameters or other models.
48–98 sone GPT-5 call, the time of 20–34 CREST evaluations
44–73%of INCA's wall-clock is LLM latency
$605total API cost of the revision, 207 runs
8 of 10parameters active for discharge in this EF5 build
Detail · technical audit notes

Verifying the parameter guide by one-at-a-time perturbation showed that two searched interflow parameters have no effect on discharge in this EF5 build, and that the guide's routing equation reproduces an inverted form from EF5's documentation while its directional statement is correct. All methods searched the same space, so the comparative conclusions are unchanged; the active dimension and the released guide are corrected.

What this work actually is

An agent is a base model plus a harness

“Learn it late enough, and you can skip it entirely.”

“With more capable models, what used to require a lot of handholding and scaffolding no longer does.” — OpenAI Developers, Rethinking skills and prompts for GPT-6 Astra · see also OpenAI Academy, Skills

LLM
pretrained intelligence
+
harness — what we build, and what the revision measured
=
working agent
INCA
Detail · why this framing changes what the paper claims

Read this way, the contribution is not "an LLM can calibrate a hydrologic model". It is that a particular harness makes a general model useful on a calibration problem, and the revision measures which parts of that harness carry the result. Two parts are load-bearing: the stated parameter–process relations and the simulate-and-feed-back loop. Two are robustness devices whose value depends on the model and the channel: the verification stage and the failure feedback. One turned out not to matter: reading rendered hydrographs rather than numbers.

The quotation at the top is the same point made from the other side. Scaffolding that a weaker model needs becomes dead weight for a stronger one, and this deck contains a measured instance of it: the staged proposal-and-verification workflow lifts Claude Sonnet 4.5 from a two-humped outcome to a reliable one and does almost nothing for GPT-5. Which parts of a harness are load-bearing is a property of the model, not of the harness, and it has to be re-measured as models change.

It also explains why the comparison with a rule-based proposer is the honest control. That proposer is the same harness with the language model removed. It works, more slowly, and with a large budget it wins — which places the language model precisely: it is a better proposal mechanism inside the harness, not a new kind of solver.

Where this sits, and what comes next

Four stages of model calibration

Stages 1 to 3 are what this paper documents. Stage 4 is an outlook, not a result.

Detail · what would make stage 4 a research programme rather than a slogan

When INCA was built, every harness component had to be designed by hand through intuition and repeated trial and error: how the agent reads diagnostics, how the loop is organised, how failed proposals are handled, how information is fed back. General-purpose agent frameworks are now much stronger at planning, tool use and recovery from failure, which suggests specifying the calibration goal and letting the agent design and refine most of the workflow itself.

The interesting version of that is to treat the harness as something that can be learned or optimised rather than hand-written — for example with reinforcement learning against calibration outcomes across many basins and models. The research question then shifts: not only whether a language model can diagnose hydrologic behaviour, but whether it can learn better strategies for designing the scientific method it applies. That also gives the field a way to measure progress between stages instead of asserting it.