7 Spit, don’t rinse
Of the three brushing recommendations I audited as blog posts, this is the one where the guideline comes off best and where I come off worst. DBOH rates the evidence for it moderate, which is higher than for frequency or timing, and when I went back to the primary trial to check that rating, I found it was not the trial I had described.
The advice
- spitting out after brushing rather than rinsing with water, to avoid diluting the fluoride concentration
Strength of recommendation: Strong
— Delivering Better Oral Health, chapter 2, tables 1d and 1f (1)
Table 1b, for children aged three to six, carries the same instruction without the words “with water”.
What the guideline says its evidence is
For children aged 3 to 6:
Recommendation based on moderate certainty for toothbrushing with fluoride toothpaste, for fluoride concentration for permanent teeth (evidence around primary teeth less clear) and spitting versus rinsing.
— Delivering Better Oral Health, chapter 13, table 2 (1)
For adults:
Moderate certainty from studies with children and adolescents for spitting versus rinsing.
— Delivering Better Oral Health, chapter 13, table 6 (1)
Moderate is a serious rating. In GRADE it means we are reasonably confident the true effect is close to the estimate, and that further research is more likely to refine it than reverse it. It is the same rating DBOH gives fluoride toothpaste itself.
The citation is SIGN 138, whose §5.7.1 grades this A and rests on two studies: a Swedish randomized trial (2) and a Scottish analysis (3).
The Swedish trial, read properly
Here is what I wrote about it in 2023:
In the Swedish one, 131 4-year-old children were instructed to rinse the remaining toothpaste in their mouth with a sip of water for a minute. Compared to no instruction, on average, 0.41 less decayed or filled tooth surface was seen after 3 years.
And elsewhere in the same series I gave the number as 369. Both numbers are in the literature, and I had not worked out why. The answer matters, so here it is.
There are two Sjögren publications from 1995. One is a doctoral thesis published as a supplement to the Swedish Dental Journal, which reports several sub-studies together (4). The other is the peer-reviewed report of the trial itself, in Caries Research (2), and it is the one SIGN cites. In that trial:
- 369 four-year-olds were randomized, to four groups;
- 281 (76%) completed at age seven;
- of the completers, 131 were in the two test groups and 150 in the two control groups.
So 369 is the randomized total and 131 is the number of completers in the test arms. Neither figure is wrong; I had used each without the other.
The 24% attrition is worth pausing on. Reporting group sizes as the number who finished, rather than the number randomized, means the published analysis is not by intention to treat, and 88 children left the study without our knowing which way they went.
But the far more important thing is what the intervention actually was.
The trial report (2) gives the test groups four instructions. I have kept the wording close and expanded the abbreviated units, so read this as a close paraphrase rather than a verbatim quotation:
- to spread the paste evenly on the teeth prior to brushing;
- not to expectorate more than necessary during brushing;
- to filter the remaining dentifrice foam in the dentition, together with a sip of water, by active cheek movements for 1 minute before expectorating;
- not to carry out any further water rinsings afterwards, and not to eat or drink for 2 hours after brushing.
Read item 3 again. The children in the test group were told to take a sip of water and swill the toothpaste foam around their mouth for a full minute.
They were told to rinse. With toothpaste slurry, deliberately, as a mouthwash.
So the trial DBOH cites as moderate-certainty evidence for “spit, don’t rinse” is a trial of a four-part technique in which one part is a form of rinsing, one part is not eating or drinking for two hours afterwards, and one part is how you apply the paste in the first place. Only the first clause of item 4 corresponds to what a modern reader understands by “don’t rinse.”
This is Chapter 3’s problem again, one level further down. The guideline bundles six instructions and lets the best-evidenced one carry the label. The trial underneath bundles four behaviors and reports one result. You cannot get from that result to any single component.
And the “no eating or drinking for two hours” instruction is not incidental. If that alone drove the effect, the correct advice would be different from the advice being given, and nobody could tell from this trial.
The result itself is real: the test groups developed a mean of 1.14 new approximal decayed and filled surfaces over three years against 1.55 in the controls, p<0.05, a relative reduction of about 26%. The 0.41 in my blog post is the difference between those two numbers. The trial report gives no confidence interval for it, which is of its time and is why I cannot give one here either.
Assessed with RoB 2, for the effect of assignment to the intervention.
Randomization process. The report states children were randomly assigned to four groups but gives no detail of sequence generation or allocation concealment. Baseline balance is reported. Some concerns, not high: the absence of reported detail in a 1995 paper is weaker evidence of a problem than it would be today.
Deviations from intended interventions. The trial is described as double-blind, which refers to the two toothpastes; it cannot be blind for the technique instruction, since the instruction is the intervention. For the effect of assignment, though, lack of blinding is only a problem if it produced deviations beyond what would happen in usual practice, and if those were not analyzed appropriately. Adherence to a four-part daily behavior over three years in four-year-olds was surely incomplete, but incomplete adherence is exactly what assignment to this advice means in the real world, and non-adherence is not itself a deviation that biases the assignment effect. What I cannot check is whether the analysis was by assigned group, because group sizes are reported as completers. Some concerns, downgraded from the High I first assigned.
Missing outcome data. 24% did not complete and the analysis is on completers. RoB 2 asks whether missingness could depend on the true outcome. Attrition in a three-year study of four-year-olds is often ordinary loss to follow-up, but here it is substantial, no comparison of attrition by group is given, and no sensitivity analysis is reported. That is enough for some concerns and, given the size, arguably High. I judge High, on the basis that a quarter of a caries trial’s participants is enough to move a 0.41-surface difference, and that nothing in the report allows the possibility to be excluded.
Measurement of the outcome. Approximal lesions scored on bitewing radiographs, restricted to the distal surface of the first and mesial surface of the second primary molars. Radiographic scoring is reasonably objective, but examiner blinding to group is not reported. Some concerns.
Selection of the reported result. No protocol available. Some concerns.
Overall: high risk of bias, driven by the missing outcome data domain alone, with some concerns in the four others. This is a weaker basis than the blog post claimed, where I recorded High in two domains; applying the signalling questions rather than the general impression moved the deviations domain down. The overall judgment is unchanged, and it still sits awkwardly with DBOH’s rating of moderate certainty. The domain-by-domain reasoning is the five paragraphs above; I have not published a separate signalling-question table, and an earlier draft implied I had.
The Scottish study
SIGN’s second source is Chestnutt and colleagues, 1998 (3). As I noted in the blog post, and as its own title says, this is an analysis of “the influence of toothbrushing frequency and post-brushing rinsing on caries experience in a caries clinical trial.”
The trial it sits inside was designed to compare toothpastes with and without zinc citrate. Rinsing behavior was not randomized. It was self-reported at interview, with participants shown a board of sketches and asked which of four methods they used: transferring water with the toothbrush, putting the mouth under the tap, using cupped hands, or using a beaker.
The reported result is that adolescents who said they used a beaker had a caries increment of 6.84 against 5.84 in those who did not, p<0.05: one more decayed, missing or filled surface.
Three problems, in increasing order of seriousness.
It is observational. Whatever the trial’s randomization achieved, it did not achieve balance on rinsing habit. Comparing beaker users with non-users is comparing groups that chose their own exposure, in a population where oral hygiene habits track socioeconomic position closely.
The four categories became two. The authors collapsed the four rinsing methods into beaker versus everything else, giving as their reason that three of them produced similar results.
In the blog post I called this “p-hacking.” I have withdrawn that. P-hacking means running analyses until one crosses a significance threshold, and I have no evidence of it. What I can say is weaker and still worth saying: the grouping is data-driven on the authors’ own account, and with no protocol available there is no way to know whether it was specified in advance. Collapsing categories because they look alike is a legitimate exploratory move. The problem is what happens two citations later, when an exploratory contrast within an observational sub-analysis is cited by a national guideline as evidence for a Strong recommendation.
Confounding. The four groups differ at baseline in caries experience and in sex distribution. If you were designing this analysis today you would specify a causal model and adjust for the common causes of both rinsing habit and caries.
In the blog version I drew a directed acyclic graph and read a “minimal adjustment set” off it: socioeconomic position, age, sex and parental education. I have withdrawn that phrase, because a minimal adjustment set is only minimal relative to a stated graph, and the graph was mine rather than anything the data supported.
What I can offer instead is a list of candidate confounders: variables that plausibly influence both which rinsing method a twelve-year-old uses and how much caries they develop. Socioeconomic position, parental education, sex, diet, baseline caries experience, dental attendance and fluoride exposure all qualify, and the trial report shows the four rinsing groups differing at baseline in caries experience and in sex distribution. Which of these actually need adjusting depends on assumptions I cannot verify from a published table.
Two cautions worth keeping whatever graph you draw. Adjusting for a variable that sits between rinsing and caries, such as fluoride retained in the mouth, would block the very effect being estimated. And adjusting for a common effect of both, such as an overall oral hygiene score, can open a path rather than close one. “Control for everything you measured” is not a method.
None of this can be settled without the individual data, so this remains a statement about what the study cannot tell us, not a re-analysis.
Where does “moderate” come from?
This is the honest difficulty of the chapter.
I can see how one gets to low certainty: one randomized trial at high risk of bias, testing a bundle, plus one observational analysis with post-hoc grouping. I cannot reconstruct a route to moderate from these two studies, and DBOH does not show its working beyond the one-line statement.
There are two possibilities and I cannot distinguish between them from the published document. Either the panel considered evidence not listed in chapter 13’s reference for this row, or moderate is the rating for the bundle as a whole, carried by fluoride toothpaste, and the phrase “and spitting versus rinsing” extends it further than the underlying evidence reaches.
I lean towards the second, because the same sentence rates three different things at once. But I want to be clear that this is an inference about a document, not a finding about the evidence, and I would withdraw it if the panel published its reasoning.
PubMed, 15 August 2026, for randomized trials of post-brushing water rinsing. Four records, all screened in full. Full strategy and screening decisions in appraisals/searches/ch07-post-brushing-rinsing.md.
No randomized trial isolating post-brushing rinsing with a caries outcome was found. Of the four records this search returned, three measure salivary fluoride retention, which is the same surrogate problem as Chapter 6:
- A crossover in 10 adults found higher salivary fluoride for up to 30 minutes without rinsing, and reported no significant difference in plasma or urinary fluoride between methods (5). With ten participants, that is a failure to detect a difference rather than a demonstration of equivalence, and the paper is a phase I study; I note it because it is the only safety data found, not because it settles anything.
- A randomized trial in 120 adults across 12 arms found rinsing significantly affected salivary fluoride retention to 90 minutes (6).
- A crossover in 32 children found the same (7).
These agree with each other and point the right way. They also illustrate the step this book keeps finding. Nazzal and colleagues conclude that their results “support the current recommendation” of spitting without rinsing. Their results are about fluoride in saliva; the recommendation is about tooth decay. That sentence is where a surrogate quietly becomes a clinical claim.
The nearest thing to a caries outcome is an in-situ study using orthodontic bands and quantitative light-induced fluorescence (8). Its test regimen was 5,000ppm fluoride and no rinsing, against a control of 1,450ppm with three rinses. Two things were changed at once, so the result cannot be attributed to the rinsing. It is the confounded-comparison problem of the Swedish trial, repeated fifteen years later.
Verdict
- Certainty of evidence
-
Low. One randomized trial at high risk of bias which tested a four-part bundle rather than rinsing alone, plus one non-randomized within-trial analysis with post-hoc regrouping. This is one rating below DBOH’s stated moderate, and I cannot reconstruct how moderate was reached.
In the blog version I said “very low.” I have moved up, not down: there is a randomized trial here with a caries outcome pointing in the right direction, which is more than either of the two preceding chapters can say.
- Directness to the advice as worded
-
Poor on intervention, and narrower on population than I first said. The Swedish trial has a patient-relevant outcome and a three-year follow-up. Its intervention is a four-part bundle of which not rinsing is one part, so the effect cannot be attributed to the advice as worded: that is indirectness on the intervention, and it is the decisive one here.
On population, the trial enrolled four-year-olds. The recommendation quoted at the top of this chapter comes from tables 1d and 1f, which cover children aged 7 to 18 and adults. So the population matches DBOH’s table 1b and is extrapolated for the rest, which is the same child-to-adult step DBOH itself flags for the other brushing components. The Scottish analysis matches the question reasonably well on population and outcome; its problem is uncontrolled confounding and a data-driven grouping, which are risk of bias, not indirectness. I had these filed under the wrong heading in an earlier draft.
- Is the strength label defensible?
- Arguably yes, unlike the previous two chapters. This is the best candidate in the brushing bundle for a legitimate discordant recommendation: the burden is close to zero, no harm has been identified, the mechanism is clear and consistent with the salivary fluoride studies, and the one randomized trial points the right way. A Strong recommendation on low-certainty evidence is defensible here on those grounds, subject to the same caveat as everywhere in this book that I have not run a full evidence-to-decision assessment. What is not defensible is stating the certainty as moderate without showing how.
- What would change my mind
- A trial randomizing people to rinse or not rinse after brushing, with everything else identical, and caries measured at two years or more. That is a cheap, ethical, entirely feasible trial that these searches did not find in thirty years of literature. Alternatively, publication of the reasoning behind DBOH’s moderate rating would settle whether I have simply missed something the panel saw.
What this means for you
Spit, don’t rinse. Of the three pieces of brushing advice in this part of the book, this is the one I would defend most readily, and I follow it myself.
But I have to state the finding precisely, because stating it loosely would commit the exact error this book is about. A randomized trial found that a four-part toothpaste technique reduced approximal decay by about 26% over three years. The effect of the one component that corresponds to “don’t rinse” is not identifiable from it. The trial also told children not to eat or drink for two hours afterwards, and to swill toothpaste foam around the mouth for a minute first. Any of those could be carrying the benefit.
So: a sensible practice, a plausible mechanism, one supportive randomized trial of a bundle that includes it, and no direct evidence for the component itself. That is a good enough reason to keep doing it. It is not the same as “this has been shown to work,” and the distinction is the whole point of the book.