Methodology · how we check

How we test claims.

Atlas KaaTai tests sweeping claims against open data — across all 400 German districts, open-ended. We confirm and we debunk. Here is how we calculate, where the method reaches its limits, and how you can verify every result yourself.

To the principles Overview

How our fact-checks come about — step by step · ← Overview

In short: we measure how strongly two variables move together across the districts, test the result against confounders, and label it honestly as true, partly true or not true. Every map is yours to open.
What we do — and what we don't

We take a sweeping claim of the kind you hear every day ("Where there's a lot of X, there's a lot of Y") and test it against data for all 400 German districts.

Our role is that of the referee: we test claims from every political direction, confirming them just as readily as we debunk them. We have no camp — only data.

Importantly, we describe patterns between regions; we do not judge people or parties. "In districts with feature A, feature B tends to be higher" is an observation about regions, not a verdict about individuals.

The metric: correlation (r)

We use Pearson's correlation coefficient r. It measures how strongly two variables rise or fall together across the 400 districts — and in which direction.

  • r = +1 — perfect alignment: where one is high, the other is high.
  • r = 0 — no linear relationship.
  • r = −1 — perfect opposition: where one is high, the other is low.

As a rule of thumb: |r| from about 0.2 is a weak, from 0.4 a moderate, from 0.6 a strong relationship. We state the r value openly in every fact-check — you don't have to take our word for it, you can recompute it.

The verdict — true / partly / not true

Every fact-check ends with one of three verdicts:

trueA strong relationship in the claimed direction — one that holds even after removing obvious confounders.
partlyThere is a relationship, but weaker than claimed or largely explained by another factor.
not trueNo relationship — or even the opposite of the claim.

The verdict always refers to the sweeping wording. "Partly" doesn't mean "who cares"; it means the simple story falls short.

The three limits we always keep in mind
  • District ≠ person (ecological fallacy). A map shows averages across regions, not individual people. A district-level relationship implies nothing about any specific person.
  • Correlation ≠ causation. Two variables occurring together does not mean one causes the other. Cross-sectional data show patterns, not chains of cause and effect.
  • Confounders. Urbanity, age or wealth can pull two variables up at the same time and create a spurious link. Where it matters, we remove such factors using partial correlation and report what remains of the relationship.
The confidence score — how reliable is an indicator?

For every active indicator, Atlas shows a confidence score (5–95%). It makes transparent how reliable the regional statement is — providers like Microm or Sinus don't publish such figures; we do.

The score combines two quantities, computed live across all areas of the active level (cells, municipalities, districts, constituencies or federal states):

  • Sample size n — areas with a value: n ≥ 200 → 50 points · n ≥ 50 → 30 · below → 10
  • Dispersion — coefficient of variation σ/|mean|: ≤ 0.35 → +40 points · ≤ 0.8 → +25 · above → +10

Interpretation: High ≥ 70% · Medium 40–69% · Low < 40% ("caution"). With two combined indicators (bivariate mode) the minimum of the individual scores applies. Also shown: n, σ and the P5–P95 spread.

Exception — absolute counts (population, persons, buildings): their large dispersion is real settlement structure — a village cell next to a city cell — not data uncertainty. For these, σ does not enter the score; it derives from n alone (≥ 1000 → 90 · ≥ 200 → 75 · ≥ 50 → 50 · below → 25), and the box marks the indicator as an "absolute value".

Important: the score measures the reliability of the aggregate, not the hit rate per person. A score of 85% means the regional differences are statistically stable — not that 85% of people there match the profile (ecological fallacy).

How the indicators are named

Every indicator carries the name used by its source. A second-vote share is called a second-vote share, a foreign-national share is called a foreign-national share, an income share is called an income share — including where a softer wording would be more comfortable in a sales conversation.

This is deliberate. The moment a label claims something the statistics do not support — a motive, a purchase intent, a „milieu“ — it stops being a translation and becomes a second, unverified claim. Anyone allocating budget from these maps has to be able to tell what was measured from what someone inferred. A pleasant label blurs exactly that line.

There is one exception, and it changes nothing about the statement: Senioren-Anteil (seniors’ share) denotes the share of the population aged 65 and over (Census 2022) — the same quantity, simply more readable than „65+ %“.

How robust an individual figure is can be found in the confidence box on every map.

The scatter-loss box — what does the selection save?

When a %-indicator is active, the 1km map shows a scatter-loss box: what does the spatial selection gain compared to a naive Germany-wide campaign?

  • Hit rate — population-weighted mean of the indicator over a cell set. “Germany-wide” runs over all ~210,000 cells, “selection” over the selected ones.
  • Selection — the active state or minimum-population filter. Without a filter the box shows a top-20% preview: the best-matching cells, cumulated up to 20% of Germany's population.
  • Budget saving = 1 − (hit rate Germany ÷ hit rate selection). Assumption: media cost proportional to the population reached, compared at an equal number of target persons reached. Example: 8% → 22% hit rate yields a 64% saving.
  • Reach / target persons — population of the selection and the target group within it (population × share).

Computed exactly at cell level (true joint distribution, no share multiplication) — using the same values that colour the map. If the selection is below the German average, the box says so honestly: “no advantage”. As with the confidence score: aggregate data describes areas, not persons (ecological fallacy).

How we test relationships — and what an r does not tell you

The panel “What is related to this?” is not a ranking of large numbers but the result of four tests. Each one removes findings that look convincing and say nothing.

  • Within one level only. Atlas indicators sit on different geometries: 53 on the 1 km cell (210,556 areas), 8 at municipal level (10,886), 16 at district level (400). District and municipal values are painted onto cells on the map — they look cell-level but are not. We therefore never compute across a level boundary. Aggregating upwards is allowed and always labelled; inheriting downwards never happens.
  • r² instead of r. We show explained variance. “Explains 21 %” is honest; “r = 0.46” sounds like more than it is. The r value sits next to it in small type. We deliberately omit significance (p): with 10,000 municipalities almost any relationship becomes significant, so the measure no longer separates anything.
  • East/West test as a traffic light. Every finding is recomputed separately for West and East. Green means it holds in both halves. Red means it falls apart when split — the relationship then mainly measures the difference between East and West. We do not hide those findings, we state the reason underneath. Example: buildings from before 1949 against the AfD vote is +0.61 nationwide and +0.12 in the West alone. At municipal level roughly one third of all relationships fail this test.
  • Trivia removed. Indicators from the same survey question are related by definition — shares that add up to 100 % must move in opposite directions. Also filtered out: quantities that are really the population count (“more people, more pharmacies”). That is arithmetic, not insight.

Two numbers, two questions. Every relationship can be computed per inhabitant or per area. Weighted by population, a city counts for more than a village — that is the question a site search asks, which is why it is our default. Unweighted, every municipality counts the same. The difference is substantial: share of foreign residents and rent are related at 13 % per municipality and 43 % per inhabitant. Both figures are correct. That is why the weighting is always stated next to the number, and the switch is one click away.

The limit that remains. These are properties of areas, not of people. On the map a higher average age is weakly positively related to the AfD vote; in election surveys the over-70s are the weakest AfD group. Both are true at once — inferring from the area to the person is a mistake. And a relationship is not a cause: oil heating does not vote.

How the discovery tour picks its records

The guided tour on the 1 km map shows 17 cells with striking values, including rankings: the densest, the youngest, the oldest square kilometre. Only cells with 300 inhabitants or more qualify for such a ranking.

The reason is arithmetic, not caution: a percentage in a cell of twelve residents is decided by individual cases. Below 300 inhabitants, net cold rents range from €1.00 to €39.35 per m²; above it, from €1.90 to €41.29 — the extremes of small cells are not particularly expensive or cheap locations, they are thin data.

Wherever the tour claims a ranking, a checker verifies it against the shipped dataset on every data release. It also fires when a heading carries a superlative without a declared ranking — which is exactly how an incorrect claim about the most expensive location went unnoticed until August 2026, while every individual value was correct.

Two limits remain. The threshold of 300 is our choice, not an official standard. And a single outlier can sit above it: the census reports Germany's highest rent as €41.29 per m² in Recklinghausen — six times the local rent index. We can neither refute nor confirm that figure, so we do not present it as a record.

Where the data comes from

We use exclusively open, freely licensed official data (DL-DE 2.0 or ODbL):

  • Census 2022 (Federal Statistical Offices) — population, age, origin, housing
  • INKAR / BBSR — over 500 regional-statistics indicators at district level
  • destatis — e.g. insolvencies, tourism, vehicle stock
  • Federal Returning Officer — federal election results (most recently 2025)
  • KBA — Federal Motor Transport Authority (vehicle stock, EV share)
  • BKA PKS — police crime statistics
  • OpenStreetMap & BKG — geometries and accessibility

Source and data year appear in every fact-check. Nothing secret, nothing bought.

Check it yourself

This is the heart of it: you don't have to trust us. Every fact-check links the interactive map in the Atlas — the very dataset we used. A direct link opens exactly the indicator (or the combination of two indicators) the video was about.

We don't hide the counter-correlations either: if a different variable explains the pattern better, you can see that for yourself in the same map.

To the interactive map

Neutrality & corrections

Credibility is our most important asset. That's why, across a season, we deliberately mix confirmed and debunked claims from different directions — checking only one side isn't neutral.

We phrase things descriptively, not judgmentally, and name uncertainties openly. Spotted an error or a questionable call? Write to beratung@kaatai.de — we correct transparently and visibly.

As of 2026-08. This methodology grows with the data and the episodes — suggestions welcome.