Part II - Critical thinking

How to Use ChatGPT to Learn Hydrological Modelling (II): conceptual understanding and critical thinking

Estimated reading time: 18-22 minutes

In Part I we looked at setting up ChatGPT, giving it context, writing clearer prompts and checking its answers. Now the goal changes.

Here ChatGPT is not an encyclopaedia. It becomes a tutor, examiner, critical interlocutor and case generator.

The theory component in IIO409 is designed precisely to reduce the value of memorising definitions and to increase the value of understanding, interpreting, diagnosing and arguing:

\[ Theory = 0.40\,ECT + 0.20\,LCA + 0.25\,DMC + 0.15\,PC \]

where ECT stands for Conceptual Theory Assessments, LCA for Critical Reading of Scientific Articles, DMC for Modelling Diagnosis through Case Studies, and PC for Scientific Participation in Class.

Visual summary of Part II

The implication is important: if you use ChatGPT only to obtain ready-made answers, you are using it in the least useful way for this assessment design.


1. Your goal is not to remember: it is to build a mental model

Suppose you are studying equifinality.

You can ask:

What is equifinality?

You will get a definition.

But a definition is not the same thing as understanding.

Try this sequence instead:

Explain equifinality with an example from a conceptual rainfall-runoff model.

Then:

Now give me two parameter sets that could produce similar hydrographs
but different internal states. I do not need real numbers; I want to understand the mechanism.

Then:

Why might this problem remain hidden if I look only at NSE or KGE?

Then:

Now give me one plausible but incorrect statement about equifinality.
Do not tell me why it is wrong until I try to criticise it.

Finally:

Change the problem: now the aim is not total discharge reproduction, but low-flow analysis.
Does the importance of equifinality change? Ask me first what I think.

This turns a static definition into a structure of reasoning.


2. The Socratic ladder

OpenAI’s Study Mode is based on interactive questioning, scaffolded explanations and checks of understanding. Even without any special mode, you can reproduce that logic in your own prompting.

Socratic ladder

Use five levels.

Level 1: explain

Explain the concept with an example from a Chilean catchment.

Level 2: retrieve

Stop the explanation and ask me five questions, one at a time.
Do not show the answer until I reply.

Level 3: challenge

Look for weaknesses in my replies.
Do not correct only the outcome: identify the conceptual error.

Level 4: transfer

Change one important condition of the problem and ask whether my reasoning still holds.

Level 5: defend

Act as a lecturer in an oral defence.
Question my assumptions and ask for a justification of each decision.

If you can answer successfully at level 5, your understanding is probably much stronger than if you can merely repeat a definition.


3. ECT (40%): Conceptual Theory Assessments

ECT tasks will prioritise interpretation and reasoning.

Theory assessment map

To prepare well, you should not simply ask ChatGPT for “exam questions”. Instead, ask it for the type of cognitive work the assessment values.

3.1 Interpreting results

Prompt:

Generate a short hydrological modelling case with:
- performance in calibration and verification;
- bias;
- one efficiency metric;
- one visible feature of the hydrograph.

Ask me three interpretation questions.
Do not give the solutions until I answer.

Then ask:

Evaluate my reply by separating:
- observation;
- inference;
- conclusion that is not yet justified.

This distinction is extremely useful. Many student errors occur when an observation is turned immediately into a causal claim.

For example:

“The model underestimates peaks.”

is an observation.

“Precipitation must be underestimated.”

is a hypothesis, not a necessary conclusion.


3.2 Comparing methodological alternatives

Try:

Compare two calibration strategies:
A) optimise NSE;
B) optimise a metric that gives greater weight to low flows.

Do not tell me which one is better.
Set up a use case and ask me which I would choose and why.
Then change the use case.

The important idea here is that a methodological decision depends on the purpose of the model.


3.3 Working with scenarios

I have a model calibrated in a wet period.
Create a dry-period verification scenario.
Ask me what problems I would expect and what diagnostics I would perform.

Then:

Give me an alternative explanation to mine.

Learning to produce alternative explanations is essential if you want to avoid overconfident conclusions.


3.4 Finding assumptions

Ask:

Give me an apparently correct methodology for calibrating a conceptual model,
but hide five implicit assumptions.
Do not flag them.
My task will be to identify them.

Then compare your list with ChatGPT’s and discuss any differences.


4. How to study each main content block in IIO409

4.1 Hydrological models and conceptual models

Do not memorise only classifications.

You should be able to answer:

  • Which process is each store trying to represent?
  • Which processes are absent?
  • Which observations could discriminate between two structures?
  • Which parts of the discharge behaviour depend on which model components?

Prompt:

Describe a catchment with climate, seasonality, snow/no snow, vegetation,
geology and discharge regime.

Do not choose a model for me.
Ask me what processes I would include in my conceptual model and why.
Then critique my representation.

A more demanding exercise:

Propose two different conceptual models that could reproduce the same discharge behaviour.
Ask me what extra data I would use to discriminate between them.

4.2 Hydrometeorological data

The question is not only “What data exist?”, but:

Are these data suitable for the modelling question?

Try:

Give me a hypothetical data set for a Chilean catchment with five quality problems:
missing dates, inconsistent units, bias, station change or temporal-resolution problems.

Do not tell me the problems. Give me clues and ask me to diagnose them.

Or:

Compare an in situ source, a satellite product and a reanalysis for precipitation.
Build the comparison in terms of:
- resolution;
- representativeness;
- biases;
- coverage;
- implications for calibration.

Then give me one case where the apparently most precise option is not the best.

4.3 Sensitivity: from “which parameter matters” to “why”

For Sobol analysis you should understand at least the conceptual difference between first-order effects and interactions.

Prompt:

Generate hypothetical Sobol results for four parameters.
Include first-order and total indices.

Do not interpret the table.
Ask me:
1. which parameter dominates;
2. which one shows interaction;
3. what this implies for calibration;
4. what I cannot conclude from these indices.

A deeper question:

Can a weakly sensitive parameter still be physically important?
Make me argue both for and against.

This helps you separate sensitivity from “physical importance”, which are not the same thing.


4.4 Calibration and verification

Study calibration as a problem of inference and decision, not merely “finding the best parameter values”.

Prompt:

Give me a case in which a calibration with PSO obtains an excellent objective value,
but the model is still not adequate.

Include enough clues for me to diagnose the problem.
Do not give me the solution yet.

You should be able to discuss:

  • objective function;
  • warm-up period;
  • dependence on the calibration period;
  • parameter bounds;
  • overfitting;
  • stability across periods;
  • multiple optima;
  • equifinality;
  • the relation between purpose and metric.

Make ChatGPT challenge you:

I am going to defend this claim:
"The parameter set with the best KGE is the true parameter set of the catchment".

Argue with me as a lecturer.
Do not accept claims without justification.

4.5 Uncertainty and GLUE

An uncertainty band does not speak for itself.

Prompt:

Create a conceptual GLUE example with:
- 10,000 simulations;
- an acceptance threshold;
- a percentage of behavioural simulations;
- a predictive band.

Ask me to interpret what each element means.
Then ask me which sources of uncertainty are NOT necessarily represented.

Another important question:

If I increase or reduce the GLUE threshold, what effects should I expect?
Do not answer: make me predict first.

4.6 Regionalisation

Prompt:

I have an ungauged catchment and three potential donor catchments.
Create climatic, topographic and hydrological-signature information for the four.

My task will be to choose a regionalisation strategy.
Then question my similarity criterion.

You need to understand that geographical proximity and hydrological similarity are not synonyms.


4.7 Climate change

Here the key is to think in terms of a modelling chain.

Prompt:

Build a conceptual chain from climate scenarios to future discharge:
climate model -> downscaling/bias correction -> meteorological variables -> hydrological model -> indicators.

Hide three sources of uncertainty.
Ask me to identify them and explain how they propagate.

Then:

Now change just one methodological decision in the chain.
Ask me how the conclusions might change.

5. LCA (20%): Critical Reading of Scientific Articles

ChatGPT can summarise papers. That is probably one of the least demanding ways to use it.

For LCA you need to go further.

Step 1: read first yourself

Before asking ChatGPT anything, identify:

  • the problem;
  • the question;
  • the hypothesis;
  • the data;
  • the method;
  • the results;
  • the main conclusion.

Then use ChatGPT to compare with your interpretation.

I am going to give you a paper.
Do not summarise it first.

Ask me these questions one at a time:
1. What is the scientific question?
2. What is the hypothesis?
3. What evidence could falsify it?
4. Which methodological decisions are critical?
5. Which limitations affect the conclusion?

Step 2: separate result from interpretation

Build a table with:
- observed result;
- authors' interpretation;
- another plausible interpretation;
- evidence needed to distinguish between them.

Step 3: search for assumptions

Identify five methodological assumptions in the paper.
For each one, indicate which conclusion might change if the assumption fails.

Step 4: formulate a new question

Based only on the paper, propose five possible follow-up research questions.
Order them from a small extension to a genuinely new conceptual question.

Then decide yourself which one is scientifically meaningful.

Step 5: prepare an oral presentation

Do not ask:

Make my slides.

Instead:

Help me identify the three essential scientific messages.
Ask me first what I think is most important before proposing them.

The presentation should be a consequence of your understanding, not a substitute for it.


6. DMC (25%): Modelling Diagnosis through Case Studies

This component is especially suitable for studying with ChatGPT because it can generate endless variations of a case.

Base prompt:

Act as an IIO409 lecturer.

Create a realistic rainfall-runoff modelling case with at least four methodological problems.
The problems may involve data, warm-up, calibration, verification,
objective function, sensitivity, parameters, uncertainty or interpretation.

Do not mark the errors.
Give me only the case and then ask:
1. which problems I detect;
2. which one is most serious;
3. what consequences it would have;
4. how I would correct it.

After you answer:

Do not give me your full solution yet.
Tell me only which important problem I failed to detect.

Finally:

Now provide your full diagnosis and compare it with mine.

Make the cases progressively harder

Beginner level:

  • calibration and verification in the same period;
  • incorrect units;
  • no warm-up period.

Intermediate level:

  • objective function inconsistent with the modelling aim;
  • parameters at the bounds;
  • climate-change analysis with an inconsistent baseline.

Advanced level:

  • good aggregate metric but structured errors;
  • parameter compensation;
  • inconsistency between internal states and discharge;
  • parametric uncertainty presented as total uncertainty.

7. PC (15%): Scientific Participation in Class

Participation is not the same thing as talking often. It means making a contribution that improves the discussion.

ChatGPT can help before class.

Suppose the topic is parameter identifiability.

Tomorrow we will discuss parameter identifiability.

Give me:
- three conceptual questions that could open a discussion;
- two debatable statements;
- one example in which high sensitivity does not guarantee identifiability;
- one counter-argument to each statement.

Do not memorise the questions. Choose one that you actually understand and rephrase it in your own words.

You can also use ChatGPT to prepare hypotheses:

Give me five possible explanations for a parameter that changes strongly between two periods.
Classify them as:
- data;
- structure;
- optimisation;
- non-stationarity.

A scientifically useful classroom contribution might be:

“If the parameter changes between periods, how do we distinguish weak identifiability from structural compensation within the model?”

That connects concepts and opens discussion.


8. Feynman technique with ChatGPT

A good test of understanding is to explain a concept in your own words.

Prompt:

I am going to explain GLUE as if I were teaching it to a classmate who is just beginning.

Do not interrupt me.
When I finish:
1. identify what is correct;
2. find ambiguities;
3. detect important missing ideas;
4. ask me one question that my explanation cannot yet answer.

Then rewrite the explanation yourself.

Do not ask ChatGPT to rewrite it for you.


9. The “predict before you ask” method

This is one of the best study strategies for hydrological modelling.

Before consulting ChatGPT, write your own prediction:

If I increase a parameter associated with storage, I expect that...

Then ask:

Evaluate my prediction.
Do not focus only on whether the direction is correct.
Discuss under which conditions it could fail.

This prevents the passive habit of allowing ChatGPT to think first every time.


10. Build your own bank of mistakes

Each time you answer something incorrectly, record:

Concept:
My answer:
Error:
Why I got it wrong:
Correct principle:
New example:

After several weeks, ask:

Here is my error log.
Do not explain the topics again.
Identify patterns and create five new questions targeted specifically at my weaknesses.

This turns ChatGPT into an adaptive practice tool.


11. Master prompt for preparing an ECT

Act as a Civil Engineering lecturer in Hydrological Modelling.

Topic: [TOPIC].

I want to prepare for a conceptual assessment.
Do not test memorisation of definitions.

Generate five questions with increasing difficulty:
1. interpretation;
2. methodological comparison;
3. cause-effect;
4. diagnosis;
5. transfer to a new scenario.

Ask me one question at a time.
Wait for my answer.
Assess the reasoning, not only the conclusion.
When my answer is incomplete, use a hint before giving the solution.
At the end, identify my two main conceptual weaknesses.

12. Master prompt for DMC

Create a realistic hydrological diagnosis case for IIO409.

Include information about:
- data;
- model;
- warm-up;
- calibration;
- verification;
- objective function;
- parameters;
- results;
- uncertainty.

Introduce between three and five methodological problems of different severity.

Do not label them.
Ask me to:
1. identify them;
2. prioritise them;
3. explain the consequences;
4. propose corrections;
5. defend one alternative.

Then question my diagnosis as in an oral defence.

13. What not to do when studying theory with ChatGPT

Avoid turning your studying into:

  1. Summary -> passive reading
  2. Flashcards -> repetition without context
  3. Automatic answers -> false confidence
  4. AI-generated essay -> no practice of reasoning
  5. Unquestioned explanations -> learning mistakes

The key question is always:

Can I apply this idea in a new case without ChatGPT telling me what to do?

If the answer is no, you are still studying too passively.


14. Before an assessment

Try this 45-minute routine:

10 min: retrieval

Without ChatGPT, write down everything you remember about the topic.

15 min: Socratic questioning

Use ChatGPT to identify gaps.

10 min: new case

Ask for a problem you have not seen before.

5 min: oral explanation

Explain it out loud without looking.

5 min: error log

Write down exactly what you did not understand.

This is far more useful for IIO409 than asking ChatGPT for “a short summary for the test”.


15. The goal: learn to think like a modeller

A good modeller is not defined by the ability to recall the longest definition.

A good modeller asks:

  • What does this model represent?
  • What does it not represent?
  • What evidence supports this decision?
  • What alternatives exist?
  • What assumption am I making?
  • How might I be wrong?
  • Which observation would distinguish between competing explanations?
  • Does the conclusion depend on the purpose?
  • What uncertainty remains?

ChatGPT can help you practise those questions again and again.

But only if you use it deliberately: make it question you, contradict you, change the conditions and demand that you justify your choices.

In Part III we will take the same logic into the practical side of the course: base R, vibe coding, reproducibility, debugging, sensitivity, calibration, uncertainty and preparation for the practical challenges.

docs