Research
From research to operational forecasting practice.
We study whether large language model agents can support frontline flood forecasting without weakening scientific accountability.
Why an agent for flood forecasting?
Climate change is driving more extreme floods, and forecasting is one of the first defenses. Operational forecasting starts with a hydrological model, but the final bulletin rarely comes straight from the model. Experienced forecasters stay in the loop, combining rainfall and water-regime information with local experience to revise the output. That judgment is often a major part of forecast quality.
This layer is tacit: hard to express, hard to audit, and slow to train. Machine learning scales, but often stays difficult to inspect. LLMs bring language, planning, and tool use, but most current uses stop at chat interfaces and miss the full operational process that real forecasting requires.
HydroAgent is built around that gap: the forecaster's work needs to be captured, reviewed, and run in the tools people actually use.
LLM Agent × Hydrology
Exploring how large language model agents can interface with hydrological models and operational data.
Forecaster-in-the-loop
Keeping human expertise central while automating routine steps in the forecast workflow.
Workflow automation
End-to-end orchestration from data ingestion to bulletin generation and review.
Papers
Research papers
HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows
† Corresponding author
Encoding forecaster judgment as explicit, rule-bounded skills lifted KGE by 0.023–0.154 over a 0.890 scheme-library baseline — and all five tested LLMs completed the same workflow.
Question
Can tacit forecaster expertise be formalized so that it is auditable and transferable, without letting a language model improvise the hydrology?
Approach
A skill-orchestrated agent framework in which each skill encodes explicit rules that bound LLM reasoning inside a model-driven forecasting workflow, with three forecaster-in-the-loop review checkpoints.
Result
Prior judgment captured observed peak flow and flood volume within a 5% tolerance in 10 and 11 of 14 events (5-fold cross-validation over 129 events: Pearson r = 0.62 and 0.84). Building on a scheme library already at mean KGE 0.890, guided scheme selection improved KGE by a further 0.023–0.154, placing simulated peak flow and volume inside the prior judgment ranges for all 14 and for 13 of 14 events. Judgment accuracy across the five LLMs ranged from 40% to 80%.
More papers will be listed here as they appear.
Start a focused discussion about product fit, workflow design, or research collaboration.
HydroAgent-Lab works with institutions, forecasting teams, and research partners that need operationally credible hydrologic systems.