2 Certainty is not strength
When I wrote the original version of this material as a blog post, I ended it with a sentence I now think is the most instructive mistake in the whole series. Having gone through the evidence for “spit, don’t rinse,” I concluded:
Therefore, the strength of the recommendation is more likely to be “very low” and not “strong.”
That sentence contains a category error, and if you can see what it is, you can skip the rest of this chapter.
Very low is not a strength. It is a certainty. Strength comes in exactly two flavors in the GRADE system, Strong and Conditional, and it describes a recommendation. Certainty comes in four, high through very low, and it describes a body of evidence about an outcome. I had mashed the two scales together, which is precisely the confusion this book exists to point out in other people. It seemed worth opening with.
The two ideas are worth keeping apart because the whole argument of this book depends on the distinction, and because the sloppy version of the argument, the one that says low certainty means it cannot be Strong, is wrong.
What certainty means
Certainty is how much we should trust an estimate of an effect.
It starts from study design. A body of randomized trials starts at high, because randomization is the one thing that reliably makes the two groups comparable in respects nobody thought to measure. A body of observational studies starts at low, because it does not.
From there it moves. GRADE has five reasons to rate certainty down: risk of bias in the studies, inconsistency between them, indirectness (the studies do not quite match the question), imprecision (the confidence interval is wide enough to include importantly different conclusions), and publication bias. It also has three reasons to rate observational evidence up: a large effect, a dose-response gradient, and the situation where all the plausible confounding would push the result the other way, so the effect survives despite the biases rather than because of them.
Those three upgrade criteria are easy to forget. I forgot them. In my post on brushing twice daily I noted that the relevant review pools observational studies, therefore started at low certainty, therefore ended at low certainty, and I never asked whether anything in the data justified moving it back up. That was a genuine misapplication, not a stylistic quibble, and Chapter 5 redoes the assessment properly. Doing it properly turns out not to produce an upgrade, for reasons that chapter sets out; the error was failing to ask, not failing to find.
What strength means
Strength is a different animal. It is the answer to a different question: not how sure are we about the size of the effect, but how sure are we that doing this is, on balance, better than not doing it, for nearly everybody this advice is aimed at.
GRADE’s own account of it makes the difference explicit. The direction and strength of a recommendation depend on four things (1):
- the estimated size of the desirable effects, and of the undesirable ones;
- confidence in those estimates, which is certainty, one input of four;
- patients’ values and preferences;
- resource use.
Certainty is in that list. It is not the list.
This is not a loophole. It follows from what a Strong recommendation is for. Labeling something Strong means: most well-informed people would choose this, so a clinician can simply advise it, and a health system can build policy on it. Labeling something Conditional means: reasonable well-informed people would differ, so the clinician’s job is to help this particular person decide.
You can be confident about that balance while being unsure about the numbers. If an intervention costs nothing, has no plausible harm, and the evidence points weakly towards benefit, then even a low-certainty estimate leaves the balance lopsided. Nobody is harmed by the low certainty, because the downside of acting on a wrong estimate is approximately zero.
Toothbrushing is close to that shape. This is the strongest argument against the crude version of my own thesis, and any critic reading this book will reach for it, so I want to hand it over rather than have it thrown.
GRADE says this is allowed, and rare
GRADE recognizes the general point. Its guidance names five specific situations in which a Strong recommendation can be justified despite low or very low certainty (1). It also says such recommendations should be uncommon, and it calls them discordant: the strength and the certainty point different ways, and the panel owes the reader an explanation.
One caution, against myself. Those five named paradigms are narrower than the argument I just made for toothbrushing. They cover situations such as a potentially life-threatening benefit on low-certainty evidence, high-certainty harm or cost against uncertain benefit, and low-certainty equivalence where one option carries high-certainty lesser harm. “Cheap, safe, and the estimate points weakly towards benefit” is not on the list. So a critic who reaches for “GRADE allows this” on behalf of brushing frequency will find Andrews’ table does not quite reach it either.
That does not rescue my crude version, because the five paradigms are examples rather than an exhaustive gate, and the real route to a Strong recommendation on weak evidence is a full evidence-to-decision assessment weighing all four inputs. The honest position is the narrow one: this book cannot run that assessment (Chapter 4), so it cannot say a Strong label is impossible. What it can say is whether the published reasoning gets there, and whether the guideline followed its own stated rule.
So the honest question is never “is the certainty low?” It is: is this one of the situations where discordance is justified, and did anyone say so?
That question turns out to have been asked before, at scale, by the people who built GRADE.
What happened when someone actually checked
In 2014, Alexander and colleagues, with Gordon Guyatt among them, went through every WHO guideline from 2007 to 2012 that had used GRADE and graded both strength and certainty. They found 456 recommendations, of which 289 were Strong. Of those 289, 95 rested on low-certainty evidence and 65 on very low-certainty evidence, which the paper reports as 55.5% of all the Strong recommendations (2). (95 and 65 make 160, and 160 of 289 is 55.4%. I give the paper’s own figure and note the rounding rather than silently correcting someone else’s arithmetic.)
Then, in a second paper, they did the thing this book does. They took the 160 discordant recommendations and asked, one at a time, whether each fitted one of GRADE’s five legitimate paradigms. The result (3):
| Judgment on the 160 discordant recommendations | n | % |
|---|---|---|
| Fitted one of the five legitimate paradigms | 25 | 15.6 |
| Evidence actually warranted moderate or high certainty, so the recommendation was not discordant after all | 33 | 21 |
| Were good practice statements, which GRADE says should not be graded at all | 29 | 18 |
| Should have been Conditional, not Strong | 73 | 46 |
Read the rows carefully, because two of them exonerate the guideline and two do not. Twenty-five were legitimately discordant. A further 33 turned out not to be discordant at all, because the evidence had been under-rated; those recommendations remain correctly Strong. That leaves the bottom two rows, 102 of 160, where the Strong label does not survive scrutiny, and the largest single category is simply that a Conditional recommendation was labeled Strong.
The authors’ own conclusion is that this pattern is “possibly threatening the integrity of the process.”
I quote this at length for two reasons. First, it settles the question of whether the exercise in this book is fair game: it is the exercise GRADE’s authors performed on the world’s largest health agency, and published. Second, it sets the standard I have to meet. It is not enough for me to find that a recommendation is Strong on low certainty. That finding is common, and sometimes fine. I have to say which of the four rows above I think it belongs in, and why.
What DBOH says its own labels mean
Everything above is GRADE’s framework. Before going further it is worth reading what Delivering Better Oral Health says its own labels mean, because it turns out to be stricter than GRADE, and because a great deal of this book can be argued on the guideline’s own terms rather than on mine.
Chapter 2 defines all three, in one passage:
strong recommendations. The GDG is highly confident that desirable consequences outweigh undesirable or undesirable consequences outweigh desirable, typically based on high or moderate certainty evidence
conditional recommendations. The GDG is less confident of the effectiveness of an intervention (low or very low certainty evidence) or the balance between benefits and harms is unclear
— Delivering Better Oral Health, chapter 2 (4)
Read the parentheses.
GRADE, as this chapter has just spent several pages establishing, permits a Strong recommendation on low-certainty evidence in defined circumstances, and treats it as discordant and uncommon. DBOH goes further than that. It writes the certainty into the definitions themselves: Strong is typically high or moderate, and low or very low certainty is what Conditional is for.
That has a consequence I did not appreciate when I started, and it changes the shape of the argument in the rest of this book. When a later chapter finds a Strong label sitting over a component that chapter 13 itself rates low certainty, the question is no longer only “would GRADE allow this?” It is also “does this match what the guideline said it would do?” The second question is much harder to wave away, because the standard being applied is the guideline’s own.
I should be fair about the word typically. It is doing real work: it signals that the rule admits exceptions, and a panel is entitled to take one. But an exception is something you notice and explain. Where I find one in the chapters that follow, what I am looking for is not a rule violation. It is whether the exception was declared.
The third category
There is a category above that neither Strong nor Conditional covers, and Delivering Better Oral Health uses it heavily: Good practice. Its own definition, from the same passage, is candid:
Clinical opinion suggests this advice is well established or supported. No robust underpinning research evidence exists. Good practice points are primarily based on extrapolation from research on related topics and/or clinical consensus, expert opinion and precedent, and not on research appropriate for rating the certainty or quality of the evidence.
— Delivering Better Oral Health, chapter 2 (4)
That is an honest label, and I have no quarrel with it. Note what it does and does not say: not that no research exists, but that no research appropriate for rating certainty exists, and that the advice rests on extrapolation from related topics plus clinical consensus. GRADE’s position is that panels should not attach certainty grades to such statements (5), and DBOH complies.
Of the 34 Good practice points in the 2025 edition, 15 carry some evidence statement anyway. That is not a contradiction, since the definition explicitly permits extrapolation from related research, and pointing the reader at that research is better than not. It does mean a reader has to look carefully to tell the difference between a citation that supports the advice directly and one that supports something adjacent to it. Telling those apart is most of what this book does.
What I will and will not conclude
For the rest of this book, when I find a Strong recommendation resting on weak evidence, I will not say the recommendation is wrong. I will ask the four questions Alexander’s team asked:
- Is the certainty actually higher than stated, and the guideline is being modest?
- Is this a legitimate discordant recommendation, where the balance of effects is clear even though the estimate is not?
- Is this really a good practice statement wearing a Strong badge?
- Or should it simply have been Conditional?
That is the honest version of the argument. The dishonest version, the one that counts low-certainty citations and declares the guideline discredited, is easier to write and easier to dismiss, and it would deserve dismissing.
But there is a fifth possibility that Alexander’s categories do not cover, because WHO’s recommendations are mostly single interventions and DBOH’s often are not. It is the one this book is actually about, and it is the subject of the next chapter.