How to Use ChatGPT to Learn Hydrological Modelling (II): conceptual understanding and critical thinking
Estimated reading time: 18-22 minutes
In Part I we looked at setting up ChatGPT, giving it context, writing clearer prompts and checking its answers. Now the goal changes.
Here ChatGPT is not an encyclopaedia. It becomes a tutor, examiner, critical interlocutor and case generator.
The theory component in IIO409 is designed precisely to reduce the value of memorising definitions and to increase the value of understanding, interpreting, diagnosing and arguing:
\[ Theory = 0.40\,ECT + 0.20\,LCA + 0.25\,DMC + 0.15\,PC \]where ECT stands for Conceptual Theory Assessments, LCA for Critical Reading of Scientific Articles, DMC for Modelling Diagnosis through Case Studies, and PC for Scientific Participation in Class.

The implication is important: if you use ChatGPT only to obtain ready-made answers, you are using it in the least useful way for this assessment design.
1. Your goal is not to remember: it is to build a mental model
Suppose you are studying equifinality.
You can ask:
What is equifinality?
You will get a definition.
But a definition is not the same thing as understanding.
Try this sequence instead:
Explain equifinality with an example from a conceptual rainfall-runoff model.
Then:
Now give me two parameter sets that could produce similar hydrographs
but different internal states. I do not need real numbers; I want to understand the mechanism.
Then:
Why might this problem remain hidden if I look only at NSE or KGE?
Then:
Now give me one plausible but incorrect statement about equifinality.
Do not tell me why it is wrong until I try to criticise it.
Finally:
Change the problem: now the aim is not total discharge reproduction, but low-flow analysis.
Does the importance of equifinality change? Ask me first what I think.
This turns a static definition into a structure of reasoning.
2. The Socratic ladder
OpenAI’s Study Mode is based on interactive questioning, scaffolded explanations and checks of understanding. Even without any special mode, you can reproduce that logic in your own prompting.

Use five levels.
Level 1: explain
Explain the concept with an example from a Chilean catchment.
Level 2: retrieve
Stop the explanation and ask me five questions, one at a time.
Do not show the answer until I reply.
Level 3: challenge
Look for weaknesses in my replies.
Do not correct only the outcome: identify the conceptual error.
Level 4: transfer
Change one important condition of the problem and ask whether my reasoning still holds.
Level 5: defend
Act as a lecturer in an oral defence.
Question my assumptions and ask for a justification of each decision.
If you can answer successfully at level 5, your understanding is probably much stronger than if you can merely repeat a definition.
3. ECT (40%): Conceptual Theory Assessments
ECT tasks will prioritise interpretation and reasoning.

To prepare well, you should not simply ask ChatGPT for “exam questions”. Instead, ask it for the type of cognitive work the assessment values.
3.1 Interpreting results
Prompt:
Generate a short hydrological modelling case with:
- performance in calibration and verification;
- bias;
- one efficiency metric;
- one visible feature of the hydrograph.
Ask me three interpretation questions.
Do not give the solutions until I answer.
Then ask:
Evaluate my reply by separating:
- observation;
- inference;
- conclusion that is not yet justified.
This distinction is extremely useful. Many student errors occur when an observation is turned immediately into a causal claim.
For example:
“The model underestimates peaks.”
is an observation.
“Precipitation must be underestimated.”
is a hypothesis, not a necessary conclusion.
3.2 Comparing methodological alternatives
Try:
Compare two calibration strategies:
A) optimise NSE;
B) optimise a metric that gives greater weight to low flows.
Do not tell me which one is better.
Set up a use case and ask me which I would choose and why.
Then change the use case.
The important idea here is that a methodological decision depends on the purpose of the model.
3.3 Working with scenarios
I have a model calibrated in a wet period.
Create a dry-period verification scenario.
Ask me what problems I would expect and what diagnostics I would perform.
Then:
Give me an alternative explanation to mine.
Learning to produce alternative explanations is essential if you want to avoid overconfident conclusions.
3.4 Finding assumptions
Ask:
Give me an apparently correct methodology for calibrating a conceptual model,
but hide five implicit assumptions.
Do not flag them.
My task will be to identify them.
Then compare your list with ChatGPT’s and discuss any differences.
4. How to study each main content block in IIO409
4.1 Hydrological models and conceptual models
Do not memorise only classifications.
You should be able to answer:
- Which process is each store trying to represent?
- Which processes are absent?
- Which observations could discriminate between two structures?
- Which parts of the discharge behaviour depend on which model components?
Prompt:
Describe a catchment with climate, seasonality, snow/no snow, vegetation,
geology and discharge regime.
Do not choose a model for me.
Ask me what processes I would include in my conceptual model and why.
Then critique my representation.
A more demanding exercise:
Propose two different conceptual models that could reproduce the same discharge behaviour.
Ask me what extra data I would use to discriminate between them.
4.2 Hydrometeorological data
The question is not only “What data exist?”, but:
Are these data suitable for the modelling question?
Try:
Give me a hypothetical data set for a Chilean catchment with five quality problems:
missing dates, inconsistent units, bias, station change or temporal-resolution problems.
Do not tell me the problems. Give me clues and ask me to diagnose them.
Or:
Compare an in situ source, a satellite product and a reanalysis for precipitation.
Build the comparison in terms of:
- resolution;
- representativeness;
- biases;
- coverage;
- implications for calibration.
Then give me one case where the apparently most precise option is not the best.
4.3 Sensitivity: from “which parameter matters” to “why”
For Sobol analysis you should understand at least the conceptual difference between first-order effects and interactions.
Prompt:
Generate hypothetical Sobol results for four parameters.
Include first-order and total indices.
Do not interpret the table.
Ask me:
1. which parameter dominates;
2. which one shows interaction;
3. what this implies for calibration;
4. what I cannot conclude from these indices.
A deeper question:
Can a weakly sensitive parameter still be physically important?
Make me argue both for and against.
This helps you separate sensitivity from “physical importance”, which are not the same thing.
4.4 Calibration and verification
Study calibration as a problem of inference and decision, not merely “finding the best parameter values”.
Prompt:
Give me a case in which a calibration with PSO obtains an excellent objective value,
but the model is still not adequate.
Include enough clues for me to diagnose the problem.
Do not give me the solution yet.
You should be able to discuss:
- objective function;
- warm-up period;
- dependence on the calibration period;
- parameter bounds;
- overfitting;
- stability across periods;
- multiple optima;
- equifinality;
- the relation between purpose and metric.
Make ChatGPT challenge you:
I am going to defend this claim:
"The parameter set with the best KGE is the true parameter set of the catchment".
Argue with me as a lecturer.
Do not accept claims without justification.
4.5 Uncertainty and GLUE
An uncertainty band does not speak for itself.
Prompt:
Create a conceptual GLUE example with:
- 10,000 simulations;
- an acceptance threshold;
- a percentage of behavioural simulations;
- a predictive band.
Ask me to interpret what each element means.
Then ask me which sources of uncertainty are NOT necessarily represented.
Another important question:
If I increase or reduce the GLUE threshold, what effects should I expect?
Do not answer: make me predict first.
4.6 Regionalisation
Prompt:
I have an ungauged catchment and three potential donor catchments.
Create climatic, topographic and hydrological-signature information for the four.
My task will be to choose a regionalisation strategy.
Then question my similarity criterion.
You need to understand that geographical proximity and hydrological similarity are not synonyms.
4.7 Climate change
Here the key is to think in terms of a modelling chain.
Prompt:
Build a conceptual chain from climate scenarios to future discharge:
climate model -> downscaling/bias correction -> meteorological variables -> hydrological model -> indicators.
Hide three sources of uncertainty.
Ask me to identify them and explain how they propagate.
Then:
Now change just one methodological decision in the chain.
Ask me how the conclusions might change.
5. LCA (20%): Critical Reading of Scientific Articles
ChatGPT can summarise papers. That is probably one of the least demanding ways to use it.
For LCA you need to go further.
Step 1: read first yourself
Before asking ChatGPT anything, identify:
- the problem;
- the question;
- the hypothesis;
- the data;
- the method;
- the results;
- the main conclusion.
Then use ChatGPT to compare with your interpretation.
I am going to give you a paper.
Do not summarise it first.
Ask me these questions one at a time:
1. What is the scientific question?
2. What is the hypothesis?
3. What evidence could falsify it?
4. Which methodological decisions are critical?
5. Which limitations affect the conclusion?
Step 2: separate result from interpretation
Build a table with:
- observed result;
- authors' interpretation;
- another plausible interpretation;
- evidence needed to distinguish between them.
Step 3: search for assumptions
Identify five methodological assumptions in the paper.
For each one, indicate which conclusion might change if the assumption fails.
Step 4: formulate a new question
Based only on the paper, propose five possible follow-up research questions.
Order them from a small extension to a genuinely new conceptual question.
Then decide yourself which one is scientifically meaningful.
Step 5: prepare an oral presentation
Do not ask:
Make my slides.
Instead:
Help me identify the three essential scientific messages.
Ask me first what I think is most important before proposing them.
The presentation should be a consequence of your understanding, not a substitute for it.
6. DMC (25%): Modelling Diagnosis through Case Studies
This component is especially suitable for studying with ChatGPT because it can generate endless variations of a case.
Base prompt:
Act as an IIO409 lecturer.
Create a realistic rainfall-runoff modelling case with at least four methodological problems.
The problems may involve data, warm-up, calibration, verification,
objective function, sensitivity, parameters, uncertainty or interpretation.
Do not mark the errors.
Give me only the case and then ask:
1. which problems I detect;
2. which one is most serious;
3. what consequences it would have;
4. how I would correct it.
After you answer:
Do not give me your full solution yet.
Tell me only which important problem I failed to detect.
Finally:
Now provide your full diagnosis and compare it with mine.
Make the cases progressively harder
Beginner level:
- calibration and verification in the same period;
- incorrect units;
- no warm-up period.
Intermediate level:
- objective function inconsistent with the modelling aim;
- parameters at the bounds;
- climate-change analysis with an inconsistent baseline.
Advanced level:
- good aggregate metric but structured errors;
- parameter compensation;
- inconsistency between internal states and discharge;
- parametric uncertainty presented as total uncertainty.
7. PC (15%): Scientific Participation in Class
Participation is not the same thing as talking often. It means making a contribution that improves the discussion.
ChatGPT can help before class.
Suppose the topic is parameter identifiability.
Tomorrow we will discuss parameter identifiability.
Give me:
- three conceptual questions that could open a discussion;
- two debatable statements;
- one example in which high sensitivity does not guarantee identifiability;
- one counter-argument to each statement.
Do not memorise the questions. Choose one that you actually understand and rephrase it in your own words.
You can also use ChatGPT to prepare hypotheses:
Give me five possible explanations for a parameter that changes strongly between two periods.
Classify them as:
- data;
- structure;
- optimisation;
- non-stationarity.
A scientifically useful classroom contribution might be:
“If the parameter changes between periods, how do we distinguish weak identifiability from structural compensation within the model?”
That connects concepts and opens discussion.
8. Feynman technique with ChatGPT
A good test of understanding is to explain a concept in your own words.
Prompt:
I am going to explain GLUE as if I were teaching it to a classmate who is just beginning.
Do not interrupt me.
When I finish:
1. identify what is correct;
2. find ambiguities;
3. detect important missing ideas;
4. ask me one question that my explanation cannot yet answer.
Then rewrite the explanation yourself.
Do not ask ChatGPT to rewrite it for you.
9. The “predict before you ask” method
This is one of the best study strategies for hydrological modelling.
Before consulting ChatGPT, write your own prediction:
If I increase a parameter associated with storage, I expect that...
Then ask:
Evaluate my prediction.
Do not focus only on whether the direction is correct.
Discuss under which conditions it could fail.
This prevents the passive habit of allowing ChatGPT to think first every time.
10. Build your own bank of mistakes
Each time you answer something incorrectly, record:
Concept:
My answer:
Error:
Why I got it wrong:
Correct principle:
New example:
After several weeks, ask:
Here is my error log.
Do not explain the topics again.
Identify patterns and create five new questions targeted specifically at my weaknesses.
This turns ChatGPT into an adaptive practice tool.
11. Master prompt for preparing an ECT
Act as a Civil Engineering lecturer in Hydrological Modelling.
Topic: [TOPIC].
I want to prepare for a conceptual assessment.
Do not test memorisation of definitions.
Generate five questions with increasing difficulty:
1. interpretation;
2. methodological comparison;
3. cause-effect;
4. diagnosis;
5. transfer to a new scenario.
Ask me one question at a time.
Wait for my answer.
Assess the reasoning, not only the conclusion.
When my answer is incomplete, use a hint before giving the solution.
At the end, identify my two main conceptual weaknesses.
12. Master prompt for DMC
Create a realistic hydrological diagnosis case for IIO409.
Include information about:
- data;
- model;
- warm-up;
- calibration;
- verification;
- objective function;
- parameters;
- results;
- uncertainty.
Introduce between three and five methodological problems of different severity.
Do not label them.
Ask me to:
1. identify them;
2. prioritise them;
3. explain the consequences;
4. propose corrections;
5. defend one alternative.
Then question my diagnosis as in an oral defence.
13. What not to do when studying theory with ChatGPT
Avoid turning your studying into:
- Summary -> passive reading
- Flashcards -> repetition without context
- Automatic answers -> false confidence
- AI-generated essay -> no practice of reasoning
- Unquestioned explanations -> learning mistakes
The key question is always:
Can I apply this idea in a new case without ChatGPT telling me what to do?
If the answer is no, you are still studying too passively.
14. Before an assessment
Try this 45-minute routine:
10 min: retrieval
Without ChatGPT, write down everything you remember about the topic.
15 min: Socratic questioning
Use ChatGPT to identify gaps.
10 min: new case
Ask for a problem you have not seen before.
5 min: oral explanation
Explain it out loud without looking.
5 min: error log
Write down exactly what you did not understand.
This is far more useful for IIO409 than asking ChatGPT for “a short summary for the test”.
15. The goal: learn to think like a modeller
A good modeller is not defined by the ability to recall the longest definition.
A good modeller asks:
- What does this model represent?
- What does it not represent?
- What evidence supports this decision?
- What alternatives exist?
- What assumption am I making?
- How might I be wrong?
- Which observation would distinguish between competing explanations?
- Does the conclusion depend on the purpose?
- What uncertainty remains?
ChatGPT can help you practise those questions again and again.
But only if you use it deliberately: make it question you, contradict you, change the conditions and demand that you justify your choices.
In Part III we will take the same logic into the practical side of the course: base R, vibe coding, reproducibility, debugging, sensitivity, calibration, uncertainty and preparation for the practical challenges.
Recommended references and further reading
- OpenAI. Study Mode in ChatGPT. https://openai.com/index/chatgpt-study-mode/
- OpenAI. Using Study Mode. https://help.openai.com/en/articles/11780217-using-study-mode-in-chatgpt
- OpenAI. A student’s guide to writing with ChatGPT. https://openai.com/chatgpt/use-cases/student-writing-guide/
- Nield, D. 28 tips to take your ChatGPT prompts to the next level. https://www.wired.com/story/28-tips-to-take-your-chatgpt-prompts-to-the-next-level/
- Jones, E. I use these 3 ChatGPT prompts to turn AI into my personal tutor. https://www.tomsguide.com/ai/i-use-these-3-chatgpt-prompts-to-turn-chatgpt-into-my-personal-tutor-and-learning-became-much-more-engaging