3 The bundling problem
Here is the sentence this book is about. It is in chapter 13 of Delivering Better Oral Health, in the introduction, and it is not hidden:
When a recommendation covers multiple components, each GDG considered the underlying evidence for each component and based the strength of recommendation on the main component.
— Delivering Better Oral Health, chapter 13 (1)
Read it twice. It says that when one recommendation contains several separate instructions, the panel looked at the evidence for each, and then set the strength of the whole recommendation by reference to one of them.
Everything else in this book follows from that sentence. So does a caveat I want to put in early, because the argument is easy to overstate and I overstated it in my first draft. DBOH does not say that the strength label applies to only one component, and it does not say which component was the main one in any given row. It says the main component determined the strength. What follows is therefore an argument about what a reader can recover from the published tables, not a claim about what the panel believed.
What it looks like in practice
The clearest example is the brushing advice for children aged three to six. In the summary table, which DBOH addresses to dental teams rather than to the public, it appears once, with one label:
Teeth should be brushed by a parent or carer. As the child gets older, a parent or carer should assist them to brush their own teeth:
- on all tooth surfaces
- at least twice a day
- last thing at night (or before bedtime) and on at least one other occasion
- with toothpaste containing at least 1,000ppm fluoride
- using a pea-sized amount of the toothpaste
- spitting out after brushing rather than rinsing, to avoid diluting the fluoride concentration
Strength of recommendation: Strong
— Delivering Better Oral Health, chapter 2, table 1b (1)
Six instructions. One word.
Now here is what chapter 13 says about the same recommendation. I have added the emphasis; everything else is verbatim.
Recommendation based on moderate certainty for toothbrushing with fluoride toothpaste, for fluoride concentration for permanent teeth (evidence around primary teeth less clear) and spitting versus rinsing. Evidence for frequency or timing is low certainty. Advice to use pea-sized amount based on possible fluorosis risk.
— Delivering Better Oral Health, chapter 13, table 2 (1)
So, unbundled:
| Instruction | Certainty, in the guideline’s own words |
|---|---|
| Brush with fluoride toothpaste | moderate |
| Toothpaste of at least 1,000ppm | moderate for permanent teeth, “less clear” for primary teeth, which are the teeth a three-year-old has |
| Spit rather than rinse | moderate |
| At least twice a day | low |
| Last thing at night and once more | low |
| A pea-sized amount | not a caries judgment at all: based on possible fluorosis risk |
The three sources cited for the whole thing are the Cochrane review of fluoride toothpaste concentrations (2), SIGN 138 (3), and the Cochrane review of topical fluoride as a cause of fluorosis (4). Not one of them studied what time of day to brush. The timing bullet’s evidence, traced through SIGN 138, is the ORCA abstract from Chapter 1.
Which of the six is the “main component”? DBOH does not say, here or anywhere. My inference is that it is brushing with fluoride toothpaste: it is the only component with moderate certainty from a large body of randomized evidence, it is the one chapter 8 of the guideline describes as the key message, and a Strong label would be hard to justify on any of the others. Chapter 8 sets out that evidence and concludes the label is thoroughly deserved.
But that is an inference, and the fact that a reader has to make it is the problem. The label sits at the end of all six bullets. Nothing in the table records that two of them are low certainty, that one is a judgment about fluorosis risk rather than about tooth decay, or which of them the panel had in mind when it wrote Strong.
This is not one bad row
Of the 91 recommendations in the 2025 edition, 16 bundle two or more separable instructions, and 7 of those carry a Strong label. The largest bundle under a single Strong label has 6 components.
Every recommendation in chapter 2 was extracted with its strength label, joined to its chapter 13 evidence statement, and its components counted from the guideline’s own bullets. The code is scripts/extract_dboh.py and the output is appraisals/dboh-2025.csv. None of the counts in this book are typed by hand; they are read from that file at build time.
The count is a floor, and a conservative one. It counts bullets, not instructions, and it undercounts in two ways.
It ignores imperatives in the stem above the bullets. The recommendation quoted above opens “Teeth should be brushed by a parent or carer. As the child gets older, a parent or carer should assist them to brush their own teeth,” which is itself two instructions about supervision, counted here as none.
And a recommendation written as flowing prose can bundle just as much while counting as one. “Minimise the amount and frequency of consumption of sugar-containing food and drinks” is two instructions, amount and frequency, on a single line with no bullets.
I have kept the bullet rule because it is mechanical and reproducible, and because erring low is the right direction for a claim of this kind. Every figure in this chapter would be larger under a more generous rule.
The same pattern recurs across the age groups, though not in the same size. The recommendation for children under three has five bullets, the one for children aged 7 to 18 has three, and the adult one has three.
The adult version is worth its own look. Its evidence statement reads:
Recommendation based on moderate certainty evidence for value of toothbrushing with fluoride toothpaste. Moderate certainty from studies with children and adolescents for spitting versus rinsing. Low-certainty evidence from children and adolescents for frequency and timing. Evidence for the concentration is based on studies on immature permanent dentition in children and adolescents.
— Delivering Better Oral Health, chapter 13, table 6 (1)
Three of the four components of the advice given to adults are explicitly sourced to evidence from children and adolescents. The guideline says so plainly, three times, in one paragraph. (The fourth, the value of brushing with fluoride toothpaste, carries no such qualifier here, though the underlying review’s adult evidence is three studies against 85 in children (2).)
That is potential indirectness, GRADE’s own term for evidence that does not match the question on population, intervention, comparator or outcome. Whether it warrants rating certainty down depends on whether the mismatch plausibly changes the effect, which is a judgment rather than an automatic consequence, and Chapter 4 describes how I made it. What is not in doubt is that the judgment has been disclosed in a chapter most readers will never open, under a table that says Strong.
Why this is different from the usual complaint
Chapter 2 granted that a Strong recommendation on low-certainty evidence can be perfectly legitimate. Toothbrushing is the textbook case: it is cheap, it is safe, people are willing to do it, and the downside of being wrong about the exact effect size is close to nothing. If DBOH had written “we are confident you should brush your child’s teeth, and we are not certain about the details,” nobody could object, least of all me.
Bundling is a different failure, and it survives that defense.
The problem is not that the panel was too confident. It is that the label cannot be resolved to its parts. A Strong label on a single recommendation carries a clear message: the panel thinks nearly everyone should do this thing. A Strong label on a bundle of six carries a message the guideline’s own methods tell us was set by one component, without telling us which. The reader is left to guess, and the guess is not recoverable from the table.
That has practical consequences, and they are not symmetric:
For the clinician the table is written for. A Strong recommendation is what you cite when a patient pushes back. If a patient asks “does the timing really matter?”, the honest answer is “probably, but the evidence for that specific part is low certainty, and it is the fluoride toothpaste that carries the weight.” The table does not let you give that answer without going to chapter 13, and chapter 13 is a separate document on a separate web page.
For the person who eventually receives the advice. Attention is finite. It would help to know which instruction matters most. That information exists in the guideline and does not survive the journey to the summary table, still less to the leaflet after that. A parent who buys 1,000ppm toothpaste and cannot reliably get a toddler to brush at bedtime has probably done the part that matters, and has no way to know it.
For anyone who later doubts one component. This is the one I care about most. If a reader discovers on their own that the bedtime-brushing bullet rests on a conference abstract, the conclusion they are likely to draw is not “that bullet was weak.” It is that the Strong label was oversold, and to discount the whole recommendation, fluoride toothpaste included. A bundle transmits doubt as efficiently as it transmits authority. That is how good advice can be damaged by its own packaging, and it is the main reason I think this is worth fixing.
What would fix it
Less than you might think, which is the frustrating part.
A recommendation with six components could carry six labels, or one label plus a line naming the component it was set by. Much of what that would need already exists: chapter 13 publishes component-level certainty in plain language, which is more than most guidelines do anywhere.
What does not yet exist is the strength judgment at component level. Chapter 13 tells us how certain the evidence is for each part; it does not tell us which part the panel treated as main, or what strength each part would carry on its own. Those are decisions only the panel can publish, and the ask of this book is that it should.
The 2025 edition is already most of the way there. My argument is not that DBOH is careless. It is that DBOH did the hard part, wrote it down, and then stopped one step short of the table people read.
Where this leaves the audits
For each recommendation from here on, I unbundle first, and then ask of each component separately:
- What exactly is this instruction?
- What evidence does the guideline cite for this component?
- Does that evidence answer this question?
- What certainty does it warrant, on my own reading?
- Given all of that, is a Strong label defensible for this component?
Sometimes the answer to (5) is yes. Fluoride toothpaste is a strong recommendation by any reasonable standard, and I will say so at length in Chapter 8. Sometimes it is no. The point of doing it component by component is that both answers become sayable, which under the current format they are not.
The next chapter sets out the rest of the method, and what would make me change my mind about any of this.