Content Governance in Healthcare

German Statutory Health Insurance Websites: What Do You Review When You Cannot Review Everything?

No editorial team reviews a mature web estate in full once a year. The question is therefore not whether to prioritize, but whether it happens deliberately or by chance.

A sentence about a benefit appears on a health insurer's website. The sentence is correct. It was already there four years ago, but back then it included the condition that is now missing.

You do not find sentences like that at a glance. You find them only by going through the entire estate, and no estate is small enough for that. I collected the public pages of 84 websites, portals, and subsidiary web properties run by German statutory health insurers and had them reviewed by machine: 56,198 pages, an average of roughly 670 per property. No editorial team in the world can review all of them professionally once a year.

Macro shot of a thick stack of loose sheets with a few slim paper tabs wedged between them; the nearest one is dark green
A few marked places. One of them is where you start.

This Is Not a Resource Problem

The reflex is “more people.” The arithmetic does not work. Reviewing 670 pages professionally once a year requires capacity that does not exist alongside the day job, and the same task returns the following year. What does work is something else: a justified order.

Prioritization happens anyway. The only question is whether it follows a plan or whatever happens to catch the eye. Pressure on that order is now growing from a direction that did not matter three years ago: a literature review describes growing interest in having health questions answered by AI systems.2 These systems summarize public web content and detach it from the context of the source page. A benefit sentence without its condition used to be imprecise. Today it becomes the template for an answer produced elsewhere that nobody can proofread.

A group from the University of Zurich and Northeastern University measured how this plays out in November 2025 using 1,508 real search queries about pregnancy and baby care.3 Two numbers stand out: the AI Overview and the featured snippet on the same Google results page contradicted each other in one third of cases. Advice to seek medical assessment appeared in eleven percent of the AI Overviews. Whether a missing condition resurfaces somewhere later is not something the source page controls.

How I Approached It

There had been no study of this design for German statutory health insurance. What I counted were web properties, not insurers: the 84 units are websites, portals, and subsidiary properties, and one insurer can run several of them. For a sense of how much of the sector that covers, the National Association of Statutory Health Insurance Funds lists 93 insurers (as of 16 July 2026).

The review looked in five directions at once: transparency, legal framing, medical communication, contradictions within an insurer's own text, and signals pointing to AI authorship.

What the study excludes is part of the same claim. Accessibility and UX writing are outside the review framework. Anyone expecting a study of every quality dimension will not find one here.

Until the end of 2024, I was responsible for member communications at a health insurer. I know the editor's side of the work: the estate that has grown over time, the legacy content from three relaunches, the page nobody touches because it is unclear who has subject-matter responsibility for it.

The review runs in three stages. First comes a rule-based pre-screen for defined signal patterns. Then an AI-assisted triage passes conspicuous pages onward and moves the rest into a shallower review stage: lower priority, not clearance. Finally, an in-depth AI review produces structured review records, each with a verbatim passage, an assessment of what is at stake, and an explicit counterargument.

Three safeguards separate this method from simply “running a language model over the site once.”

Evidence requirement. Every record must include a verbatim passage from the page under review. A deterministic comparison then checks whether the quoted passage can actually be found in the captured page text. Of 35,998 review records, this succeeded for 31,347. Another 4,631 quotations could not be located, and 20 records had no usable status. Those stay flagged as needing review rather than going out as substantiated findings.

Two clarifications matter because this number is easy to misread. It counts review records, not pages and not errors: one page can carry several, and none of them is an established defect. The comparison also proves that the passage is present, not that its claim is true.

Second assessment. Two independently configured models assessed the same material in 182 paired cases. They agreed in 75.8 percent of cases. That sounds like a lot, but it is not: once you subtract how often two assessors would agree by chance alone, what remains is moderate agreement (Cohen's kappa 0.532; 95% confidence interval 0.415–0.649). Forty-four cases went on to downstream professional clarification. That is not a flaw in the method but a design principle: two language models are not independent reviewers, and shared training data can produce shared mistakes. Agreement is therefore a triage signal, not proof of truth.

Currency safeguard. The most awkward failure mode occurs when the page under review and the reviewing model share the same outdated knowledge. The model then confirms a superseded statement as correct. Curated currency checks cover topics prone to change.

What comes out at the end is a prioritized worklist. Not findings. AI handles the persistent first pass; the assessment stays with people.

What Recurs Across Insurers

Across the properties, 290 review patterns recur, 42 of them at the highest priority level. The table shows seven of those 42. The patterns overlap, and none of them is an established finding.

The most important point is not any single number but the spread. The need for review is structural, not confined to a few conspicuous insurers.

Review pattern In how many properties
Evidence and advertising-law admissibility of health-related claims 67 of 84
Legal framing and sources behind a statement 50
Conditions, caps, and bylaw basis of benefit promises 47
Scope and limits of benefit communication 45
Currency of content 42
Communication of bylaw-based benefits 40
Consistency of medical statements 24

The first pattern has the widest spread: it occurs in 67 of the 84 properties. Another is more consequential where it occurs: consistency of medical statements affects only 24 properties, but the pattern is especially concentrated within them.

Both readings are useful, and they mean different things. Breadth shows where a cross-insurer review route is worthwhile. Density shows where most of the work sits within an individual insurer.

One qualification should not be skipped: these are review candidates, not established violations. A number in this table means that a person with domain expertise should take a look, not that something is wrong. The study deliberately names no insurers and produces no ranking. A ranking would give the notes a finality that the method does not claim.

The study does not measure error rates. It collects review candidates, and the questions behind them are always the same. Does a statement carry its conditions? Does it still hold? What supports it? Is it medically coherent, and what legal framework does it sit in?

Governance Remains Invisible from the Outside

An additional module checked whether basic governance signals are discoverable along a defined public search path. In other words, can an outside observer tell how an organization maintains its health content?

Of 84 properties, eleven name a clear editorial responsibility for content. Six show a visible route for reporting a content error. Four say how often they review their content. Three say anything about how they handle AI-assisted text production.

These numbers measure public discoverability only. A missing signal does not mean that internal processes are missing, only that they are not visible to members, researchers, and regulators. For organizations expected to promote digital health literacy, that is itself an issue.

Something similar has been observed before. Viviane Scherenberg and Melanie Preuß manually reviewed the digital-health-literacy offerings of 97 insurers in 2023; 16 of them were directly discoverable on the websites.1 Different scope, different measurement. But they got there three years earlier.

What Editorial Teams Can Do with This

The method shifts the question from “can we review everything?” to “where do we direct limited review capacity first?” Four questions are enough as a framework. Every health and benefit page should be able to answer them:

  1. Who has subject-matter responsibility for this statement?
  2. Which reference governs it?
  3. When was it last reviewed?
  4. What happens when the bylaw, the evidence, or the legal framework changes?

Anyone who can answer these four questions for their top pages already has most of the review need under control. Anyone who cannot at least knows where to start.

In practice, this becomes a prioritized review worklist instead of a full sweep. Fixed review cycles for topics prone to change. Clear subject-matter ownership for health and benefit information. Update triggers for changes to guidelines, law, or benefit status. And visible governance signals, so outsiders can see that the content is maintained.

None of this requires an external audit or new technology. Most of it is organization, not tooling. The full method and the bounded work sample are on the AußenBlick GKV project page.

And the Benefit for Members?

It is indirect, but it is the point. Where insurers work through the prioritized patterns, the stock of unreviewed legacy content shrinks, benefit information carries its conditions, and sources and currency become traceable. That works twice over: for members who read the page directly and for the growing group who receive the same information through an AI system, in compressed form and without the context of the source page.

I did not measure that benefit. The study measures review need, not effect. But the direction is plausible: the more current and consistent the source text, the more reliable the outputs that machines derive from it are likely to be.

Sources & References


Method, data cutoff, and limits are documented in the preprint. The full text is available at arXiv:2608.03500 and as a Zenodo preprint. A separate, edited snapshot package is published as a dataset. The raw crawl, full page texts, production code, prompts, and thresholds are not public. Data as of May 2026.

One question to close on, and I am genuinely curious: which review step would you add from your own practice, and which of the four governance questions is hardest to answer in your organization?

🌐