Content Testing: Test Clarity, Measure Impact
A lean content-testing process for clearer interfaces, supportable conversion hypotheses and more reliable AI conversations.
Some interface copy immediately reveals that nobody has read it in context. On 16 September 2025, I documented one such case on LinkedIn: a button in the “AOK Mein Leben” app said “AOK Mein Leben beenden”—literally, “End AOK My Life.”
The technical action was unremarkable. It was supposed to close the app. In combination with the product name, however, the system text sounded macabre. Two days later, AOK connect confirmed in a comment on the post that the button label had been changed.
The case does not prove a loss in revenue. It demonstrates something smaller but more useful: copy that is clear in isolation can take on a very different meaning inside an interface. For statutory health insurance websites, the next step is therefore to review content systematically in context before an isolated sentence is held up publicly as an example of poor UX.
What content testing can do
Content is not filler for a slick interface. Content is the interface: there is no salesperson standing next to it to answer questions, so the copy alone has to do the persuading.
Content testing examines whether people can find, understand and use content for their task. Unclear copy can therefore contribute to abandonment, support requests or poor decisions: people are less likely to choose something they do not understand. Whether that produces a measurable commercial effect in a given case is something the test first has to establish.
The method replaces neither analytics nor an A/B test. It answers different questions:
- Qualitative tests show why a phrase causes confusion or a task fails.
- Analytics show where and how often people abandon, go back or seek help.
- A sufficiently powered A/B test can determine whether a variation changes a predefined behaviour.
“Easier to read” is not an outcome by itself. A randomised study with 2,235 participants found no automatic improvement in knowledge, acceptance or trust for health information merely because the reading level was lowered.1 Good language has to fit the task and be measured against an appropriate outcome.
A small test that leads to a decision
Five points should be settled before the first session:
- Research question: What should be clearer after the test?
- Audience and task: Who uses the content, and in what situation?
- Success criterion: How will we know that the task succeeds?
- Guardrail: What must not get worse, such as factual accuracy, abandonment or demand for support?
- Decision: Which finding triggers which change?
For the AOK case, the question could be: Do participants understand the button as closing the app? The primary criterion would be the correct interpretation without assistance. The guardrail would be that the alternative remains short and unambiguous in the existing interface. Only after this qualitative round would it make sense to test whether the change affects abandonment or support requests.
One quick method is content highlighter testing.2 Participants mark useful and confusing passages, then explain their choices. Colour alone is not enough: “helpful” and “unclear” also need text labels or symbols, sufficient contrast and an alternative for people who cannot mark content visually.3
The number of participants depends on the question and method. The GOV.UK Service Manual often suggests four to eight participants for qualitative rounds.4 That is planning guidance, not a guarantee. Quantitative impact questions require their own sample-size and test-duration planning.
Testing chatbot reliability
For AI conversations, content testing becomes a test of the answer, the failure mode and the next step: does the person reach their goal, does the answer match a predefined reference, and does an escalation actually reach the right team? A friendly tone cannot rescue a false answer. The Microsoft guidelines for human-AI interaction offer a compact framework for this.
Data can make discussions more objective
In many organizations, wording is a political minefield: marketing wants it emotional, legal wants it safe, product wants it precise. One familiar pattern behind this is the “HiPPO effect” (Highest Paid Person's Opinion): the opinion of the highest-paid person in the room wins. That is not a universal diagnosis of every organization, just a decision dynamic you can observe on many teams.
Content testing does not resolve this power dynamic on its own. But a test can change a discussion when the team, subject-matter experts and legal reviewers agree beforehand on the same hypothesis, criteria and next decision. The question is no longer “I think this sounds better,” but “The data shows that variant B gets the job done more reliably.”
According to a survey of 305 medium-sized and large companies, management support and perceived data quality encourage the use of data in decision-making. This does not mean that one content test changes an organisation's culture. It can, however, create a manageable setting in which a decision is justified transparently instead of hierarchically.
From finding to impact
A test does not end with a list of highlighted sentences. It ends with a documented decision, a new variant and a measurement point. Qualitative research explains the finding; analytics or a properly planned A/B test then determine whether the change works in production.
That keeps the commercial claim honest, too: clarity can influence conversion, abandonment and demand for support. How much it does so remains unknown until it is measured.
It's supposed to hurt
Content testing can hurt. It shows that our internal priorities are often beside the point to the people using the product, and that a brand name like “Mein Leben” combined with a system label like “beenden” can backfire badly.
But that discomfort is exactly the point. Teams that take content testing seriously shift their focus from output (“we wrote ten pages of copy”) to outcome (“we checked whether the audience understands the task”). Whether that saves budget or grows revenue is a hypothesis only measurement can confirm, not a promise. But the odds get better: fewer punchlines on LinkedIn, more satisfied users.

Sources & References