Methodology & Science
This page explains how neighbourhood data is transformed into theme scores, how the model combines them, and how present and future heat-health vulnerability should be interpreted.
It tells us how vulnerable one neighbourhood looks compared with other Dutch neighbourhoods when heat exposure, population sensitivity, social conditions and the built environment are considered together.
A higher score means that more heat-related risk factors are present in the same place. It does not mean that a certain percentage of residents will become ill.
The overall score is the final number shown on the map. It combines the four theme scores into one heat-health vulnerability estimate.
Theme scores explain what is driving the overall score. They show whether risk comes more from heat exposure, population sensitivity, socioeconomic conditions or the built environment.
The score is built in four steps. First, the different variables are made comparable. Then they are grouped into four broader vulnerability themes. The model uses these themes to estimate risk and the final result is converted into the 0–100 score shown on the map.
The variables come in different units: percentages, distances, temperatures and survey scores. We convert them to the same scale so the model can compare them fairly.
Related variables are combined into four themes. This makes the dashboard easier to read and shows what kind of vulnerability is driving the risk.
The four themes are passed into a ridge regression model. The model learns how strongly each theme is connected to the overall heat-health risk outcome.
The model output is converted into a clear 0–100 score, so neighbourhoods can be compared easily on the map.
Puts all inputs on a fair scale before they are combined.
Finds patterns between related indicators and turns them into theme scores.
Combines the themes in a more stable way when predictors overlap.
The dashboard groups many neighbourhood indicators into four themes. Each theme captures one side of heat-health vulnerability.
For each theme, the variables are first put on the same scale. Protective variables are reversed, so that the meaning stays consistent: higher always means more vulnerable. A factor analysis model then finds the common pattern shared by the variables in that theme. That common pattern becomes the theme score.
The theme score is finally converted to a 0–100 scale, so it is easy to compare across neighbourhoods and across themes.
The model is the step that turns the four separate theme scores into one overall heat-health risk score. The theme scores tell us what kind of vulnerability a neighbourhood has. The model decides how strongly those themes should count together.
The model receives the four theme scores, not every raw variable separately.
It estimates how much each theme contributes to the overall risk pattern.
It produces one score that can be compared across neighbourhoods.
A simple average would assume that every theme matters equally. That is too simplistic and does not capture real life events. Heat exposure, population sensitivity, socioeconomic conditions and the built environment do not contribute to risk in exactly the same way. The model lets the data decide the balance between them.
Many neighbourhood characteristics overlap. Places with dense housing may also have less green space or more socioeconomic pressure. Ridge regression is useful here because it keeps the model stable when predictors are related. It reduces extreme weights and makes the final score less dependent on one noisy pattern.
The model result should be read as a risk index. It is not saying exactly how many people will become ill during a heatwave. It shows where several risk drivers come together more strongly than in other neighbourhoods.
This is why the theme scores remain important. The overall score tells us where risk is high. The themes explain why it is high.
The future score is a scenario calculation. We do not train a new model for the future. Instead, we start with the current neighbourhood data and adjust only the variables for which we have a reasonable assumption from climate scenarios or supporting literature.
If a neighbourhood has a higher future score, it means that the scenario assumptions make its heat-health vulnerability worse. The increase comes from changed inputs, such as stronger heat exposure or population pressure, not from changing the scoring method.
The baseline score is scaled from 0 to 100 using the present-day range. Future scores are compared against that same baseline range. If a future score is above 100, it means the projected risk is higher than the highest risk observed in the current data.
The future results should therefore be read as a structured “what if” scenario. They do not claim to predict the exact future. They show how neighbourhood vulnerability could shift if the assumptions used in the scenario become reality.