Content Governance in Healthcare

What do you review when you cannot review everything?

No editorial team gets through a grown web estate once a year. So the question is not whether you prioritise, but whether you do it deliberately or by accident.

A sentence about a benefit sits on a health insurer's website. The sentence is correct. It was already there four years ago — with the condition that is missing now.

You do not find sentences like that at a glance. You find them by going through the whole estate, and no estate is small enough for that. I collected the public pages of 84 web properties belonging to German statutory health insurers and had them reviewed by machine: 56,198 pages, roughly 670 per property. No editorial team on earth works through that once a year.

Macro shot of a thick stack of loose sheets with a few slim paper tabs wedged between them; the nearest one is dark green
A few marked places. One of them is where you start.

This is not a resource problem

The reflex is "more people". The arithmetic does not support it. Reviewing 670 pages once a year requires capacity that does not exist alongside the day job — and the same task returns the following year. What does work is something else: an order you can justify.

Because prioritizing happens anyway. The only question is whether it follows a plan or whatever happens to catch the eye. And that order is now under growing pressure from a source that did not matter three years ago: interest in having health questions answered by AI systems is rising2. Those systems summarize public web content and lift it out of the context of the page it came from. A benefit sentence without its condition used to be imprecise. Today it is the raw material for an answer produced somewhere else, which nobody gets to proofread.

How I went about it

No study of this kind existed for German statutory health insurance before. The 84 web properties belong to 83 independent insurers; the 84th is the shared AOK site. With roughly 90 to 95 insurers in Germany, the sample is close to the whole sector and was reviewed across four dimensions at once: medical, legal, editorial, and signals about how a text was produced.

What the study leaves out is part of the same claim. Accessibility and UX writing are outside the review framework. Anyone expecting a study across every quality dimension will not find one here.

Until the end of 2024 I was responsible for member communications at a health insurer. So I know what the work looks like from the editor's side: the estate that has grown over time, the legacy content from three relaunches, the page nobody touches because it is unclear who has subject-matter responsibility for it.

The review runs in three stages. First a rule-based pre-screen for defined signal patterns. Then an AI-assisted triage that closes unremarkable pages and passes on the rest. Finally, an in-depth AI review that produces structured review hints: each with a verbatim quote from the page, an assessment of what is at stake, and an explicit counterargument.

Three safeguards separate this from running a language model over the site once.

Evidence requirement. Every hint has to carry a verbatim quote from the page it came from. A deterministic check then tests whether that quoted passage can actually be located in the captured page text. Of 35,998 review hints, that worked for 31,347. The rest stays flagged as needing review rather than being included in the analysis.

Two clarifications, because this number is easy to misread. It counts review hints, not pages and not errors — one page can carry several, and none of them is an established defect. And the check confirms that the passage is there, not that its claim is true.

Second opinion. In 182 paired cases, two independently configured models assessed the same material. They agreed in 75.8 percent of cases. So roughly every fourth case was classified differently and needed a later expert decision. That is not a flaw in the method; it is its design principle: two language models are not independent assessors, and shared training data can produce shared mistakes. Agreement is a triage signal, not proof of truth.

Currency safeguard. The most awkward failure mode occurs when the page under review and the reviewing model share the same outdated knowledge — the model then confirms a superseded statement as correct. For topics prone to change, curated currency checks test for that problem.

What comes out at the end is a prioritized worklist. Not findings. The AI handles the first pass; the judgment stays with people.

What recurs across insurers

Across the properties, 290 review patterns repeat, 42 of them in the highest priority tier. The most important result is not in any single number but in the spread. The need for review is structural, not confined to a few conspicuous insurers.

Review pattern In how many properties
Evidence and advertising-law admissibility of health-related claims 67 of 84
Legal framing and sources behind a statement 50
Conditions, caps and bylaw basis of benefit promises 47
Scope and limits of benefit communication 45
Currency of content 42
Communication of bylaw-based benefits 40
Consistency of medical statements 24

The most widely distributed pattern is the first: it occurs in 67 of the 84 properties. Another pattern is denser where it appears: consistency of medical statements affects only 24 properties but recurs frequently within them.

Both readings are useful and they mean different things. Breadth shows where a cross-insurer review route pays off. Density shows where the work is concentrated within an individual insurer.

One qualification I would rather not skip: these are review candidates, not established violations. A number in this table means somebody with domain expertise should look — not that something is wrong. The study deliberately names no insurers and produces no ranking. A ranking would give the hints a finality that the method does not claim.

What stands out is where the emphasis falls. Very little of it is about wrong facts. It is about conditionality, currency, sourcing and legal framing — whether a statement carries its preconditions, whether it still holds, what it rests on, and which legal framework applies.

The side finding: governance is invisible from outside

An additional module coded whether basic governance signals are discoverable along a defined public search path. In other words, can an outside observer tell how an insurer maintains its health content?

Of 84 properties, eleven name a clear editorial responsibility for content. Six show a visible route for reporting a content error. Four say how often they review their content. Three say anything about how they handle AI-assisted text production.

These numbers measure public discoverability and nothing beyond it. A missing public signal does not mean internal processes are missing — they almost certainly exist. It only means they are not visible to members, researchers or supervisory bodies. For organizations charged with promoting digital health literacy, that is a topic in itself.

The observation is not new; only the scale is. Viviane Scherenberg and Melanie Preuß reviewed the digital-health-literacy offerings of 97 insurers by hand in 2023 and arrived at a similar observation about limited central discoverability.1 Their inventory was considerably smaller, but it came three years earlier.

What editorial teams can do with this

The method shifts the question from "can we review everything?" to "where do we point limited review capacity first?" Four questions are enough to structure the review. Every health and benefit page should be able to answer them:

  1. Who has subject-matter responsibility for this statement?
  2. Which reference governs it?
  3. When was it last reviewed?
  4. What happens when the bylaw, the evidence or the legal frame changes?

Anyone who can answer these four for their top pages already has most of the review need under control. Anyone who cannot at least knows where to start.

In practice, that means a prioritized review worklist instead of a full sweep. Fixed review cycles for topics prone to change. Clear professional ownership for health and benefit information. Update triggers for changes to guidelines, law or benefit status. And visible governance signals, so outsiders can see that the content is maintained.

None of this requires an external audit or new technology. Most of it is organization, not tooling.

And the benefit for members?

It is indirect, but it is the point. Where insurers work through the prioritized patterns, the stock of unreviewed legacy content shrinks, benefit information carries its conditions, and sources and currency become traceable. That works twice over: directly for members who read the page, and indirectly for the growing group who get the same information through an AI system, in compressed form and without the context of the original page.

I did not measure that benefit. The study measures review need, not effect. But the direction is plausible: the more current and consistent the source text, the more dependable whatever machines make of it.

Sources & references


Method, data cutoff and limits are documented in full. The full text and the anonymized dataset are openly available: 10.5281/zenodo.20591063. Data as of May 2026.

One question to close on, and I am genuinely curious: which review step would you add from your own practice — and which of the four governance questions is hardest to answer in your organization?

🌐