Methodology & Science

How health risk
is measured

This page explains how neighbourhood data is transformed into theme scores, how the model combines them, and how present and future heat-health vulnerability should be interpreted.

01

What do the risk scores represent?

The score is a relative vulnerability score, not a medical prediction.

It tells us how vulnerable one neighbourhood looks compared with other Dutch neighbourhoods when heat exposure, population sensitivity, social conditions and the built environment are considered together.

A higher score means that more heat-related risk factors are present in the same place. It does not mean that a certain percentage of residents will become ill.

0–30 Lower relative vulnerability
30–70 Moderate relative vulnerability
70–100+ Higher relative vulnerability

Overall risk score

The overall score is the final number shown on the map. It combines the four theme scores into one heat-health vulnerability estimate.

  • Use it to compare neighbourhoods with each other.
  • Higher means more overall vulnerability.
  • Scores above 100 can appear in future scenarios when projected risk goes beyond today’s highest baseline value.

Theme scores

Theme scores explain what is driving the overall score. They show whether risk comes more from heat exposure, population sensitivity, socioeconomic conditions or the built environment.

  • Use them to understand why a neighbourhood scores high or low.
  • Higher means more vulnerability within that specific theme.
  • Two neighbourhoods can have the same overall score for very different reasons.
02

From neighbourhood data to one score

The score is built in four steps. First, the different variables are made comparable. Then they are grouped into four broader vulnerability themes. The model uses these themes to estimate risk and the final result is converted into the 0–100 score shown on the map.

Step 01
Make variables comparable

The variables come in different units: percentages, distances, temperatures and survey scores. We convert them to the same scale so the model can compare them fairly.

Step 02
Build four theme scores

Related variables are combined into four themes. This makes the dashboard easier to read and shows what kind of vulnerability is driving the risk.

Step 03
Combine the themes

The four themes are passed into a ridge regression model. The model learns how strongly each theme is connected to the overall heat-health risk outcome.

Step 04
Show the final score

The model output is converted into a clear 0–100 score, so neighbourhoods can be compared easily on the map.

What happens behind the score variables → four themes → model prediction → dashboard score
z-score = (valuedataset mean) / standard deviation
Theme score = weighted factor score based on related z-scored variables
Model prediction = β₀ + β₁Theme₁ + β₂Theme₂ + β₃Theme₃ + β₄Theme₄
Dashboard score = rescaled model prediction on a 0–100 range
Standardisation

Puts all inputs on a fair scale before they are combined.

Factor analysis

Finds patterns between related indicators and turns them into theme scores.

Ridge regression

Combines the themes in a more stable way when predictors overlap.

03

The four themes

The dashboard groups many neighbourhood indicators into four themes. Each theme captures one side of heat-health vulnerability.

How a theme score is made

For each theme, the variables are first put on the same scale. Protective variables are reversed, so that the meaning stays consistent: higher always means more vulnerable. A factor analysis model then finds the common pattern shared by the variables in that theme. That common pattern becomes the theme score.

The theme score is finally converted to a 0–100 scale, so it is easy to compare across neighbourhoods and across themes.

04

What the model actually does

The model is the step that turns the four separate theme scores into one overall heat-health risk score. The theme scores tell us what kind of vulnerability a neighbourhood has. The model decides how strongly those themes should count together.

Model logic four theme scores → one risk estimate
overall risk estimate = weighted mix of the 4 theme scores
dashboard score = overall risk estimate rescaled for comparison
Input

The model receives the four theme scores, not every raw variable separately.

Learning step

It estimates how much each theme contributes to the overall risk pattern.

Output

It produces one score that can be compared across neighbourhoods.

Why not just average the themes?

A simple average would assume that every theme matters equally. That is too simplistic and does not capture real life events. Heat exposure, population sensitivity, socioeconomic conditions and the built environment do not contribute to risk in exactly the same way. The model lets the data decide the balance between them.

Why ridge regression?

Many neighbourhood characteristics overlap. Places with dense housing may also have less green space or more socioeconomic pressure. Ridge regression is useful here because it keeps the model stable when predictors are related. It reduces extreme weights and makes the final score less dependent on one noisy pattern.

How to read the model result

The model result should be read as a risk index. It is not saying exactly how many people will become ill during a heatwave. It shows where several risk drivers come together more strongly than in other neighbourhoods.

This is why the theme scores remain important. The overall score tells us where risk is high. The themes explain why it is high.

05

How future risk is calculated

The future score is a scenario calculation. We do not train a new model for the future. Instead, we start with the current neighbourhood data and adjust only the variables for which we have a reasonable assumption from climate scenarios or supporting literature.

The future score is calculated in five steps

  1. Use the present as the baseline: each neighbourhood starts from its current data and current theme scores.
  2. Change selected inputs: climate-related variables are adjusted using future heat and weather assumptions. Population or density variables are only changed where the literature gives a defensible direction for change.
  3. Leave unsupported variables unchanged: when there is no clear evidence for a future change, the variable stays the same. This avoids pretending that we know more than we actually do.
  4. Recalculate the themes: after the selected inputs are updated, the four theme scores are calculated again using the same theme structure as the present-day model.
  5. Apply the same ridge model: the future theme scores go through the same model logic, so the future score can be compared with the present score.

What the future score means

If a neighbourhood has a higher future score, it means that the scenario assumptions make its heat-health vulnerability worse. The increase comes from changed inputs, such as stronger heat exposure or population pressure, not from changing the scoring method.

Why scores can go above 100

The baseline score is scaled from 0 to 100 using the present-day range. Future scores are compared against that same baseline range. If a future score is above 100, it means the projected risk is higher than the highest risk observed in the current data.

Future scenario logic same model, changed evidence-based inputs
current neighbourhood data + literature-based assumptions
→ updated future inputs
→ recalculated theme scores
→ same ridge model
→ future risk score compared with the present baseline

The future results should therefore be read as a structured “what if” scenario. They do not claim to predict the exact future. They show how neighbourhood vulnerability could shift if the assumptions used in the scenario become reality.

06

Limits of the score

01
It is a relative score, not a medical forecast. A score of 80 means higher vulnerability compared with other neighbourhoods. It does not mean that 80% of residents will become ill.
02
We assume no behavioural adaptation. In reality, residents may buy air conditioning, relocate, renovate their homes or build mutual support networks. The long-term projections therefore show risk under the model assumptions, not a guaranteed future.
03
Health data gaps exist. 11,554 of 14,668 neighbourhoods have complete health data. Any neighbourhood lacking sufficient CBS or GGD polling data is excluded from the ranking calculations to preserve statistical reliability.
04
It works at neighbourhood level. The score can hide differences between streets, buildings or groups of residents inside the same neighbourhood.
05
Future scores depend on assumptions. They show a scenario, not a certainty. Different climate, population or planning assumptions would produce different future scores.
06
The themes explain risk, but not causality. A high theme score shows an association with vulnerability. It should be used to guide planning, not to prove that one single factor directly caused the risk.