4  How I did this

This chapter is the protocol. It exists so that you can tell the difference between a finding and an opinion, and so that someone who disagrees with me has something concrete to attack.

I am aware of the irony of a book that complains about untraceable evidence having an untraceable method, so I have tried to write this the way I would write the methods section of a paper I expected a reviewer to be unkind about.

What kind of study this is

The design has a name, even if it is not a common one: evidence-to-recommendation traceability. You take a set of published recommendations, follow each to the evidence its own authors say it rests on, and assess whether that evidence answers the question the recommendation asks.

Two things it is not.

It is not a systematic review, and the reason is not that it contains no meta-analysis, since plenty of systematic reviews do not. It is that it lacks the features that make a review systematic: the searches are single-database, screening was partial, there was one screener with no independent duplication, and there is no registered protocol. Where a good systematic review of a question exists I use it, and my job is to check whether the guideline used it correctly.

It is not an appraisal of the guideline development process. I was not in the room. I cannot know what a panel discussed and decided not to write down. Everything here is based on the published document, which is also all any reader has.

The unit of analysis

The unit is the component, not the recommendation.

Chapter 3 explains why, and the claim needs stating precisely because a looser version of it appears in earlier drafts of several chapters.

DBOH says the strength was set on the main component. It does not say the label applies only to that component, and I am not entitled to say so either. A recommendation can perfectly well be a package: “brush twice daily with fluoride toothpaste and spit rather than rinse” is a routine, the trials mostly tested routines, and recommending the package is a legitimate decision even where the marginal contribution of each part is unknown.

The narrower claim, which is the one this book can defend, is about recoverability: when six instructions sit under one label, a reader cannot tell from the published summary which of them the evidence supports, how strongly, or what reasoning produced the label. That is a criticism of what the guidance makes visible, not a proof that the recommendation is wrong.

So each recommendation is first split into its separable instructions, following the guideline’s own bullets where it has them and its own sentence structure where it does not, and each component is carried through the steps below on its own. Where a chapter concludes that a component would merit a different label, that is my reading of the evidence for that component, not a completed evidence-to-decision assessment of the package.

The seven steps

Every audit chapter does these, in this order, under these headings.

1. The advice. Quoted verbatim from chapter 2 of Delivering Better Oral Health (1), with its strength label and the table it appears in. Not paraphrased. If a recommendation sounds worse when quoted exactly than when summarized, that is information, and paraphrasing would destroy it.

2. What the guideline says its evidence is. Quoted verbatim from chapter 13, with its certainty rating and its cited references. Again not paraphrased.

3. Following the citation. I obtain and read the cited source, and compare it to the advice on four axes:

  • Population: are these the people the advice is for?
  • Intervention: is this the thing the advice tells you to do?
  • Comparator: is this the alternative the advice implies?
  • Outcome: is this something a patient would notice or care about?

A mismatch on any axis raises indirectness in GRADE’s vocabulary. Two steps, not one: finding a mismatch identifies potential indirectness, and a separate judgment then asks whether the mismatch is serious enough to change the answer. Evidence in twelve-year-olds applied to fourteen-year-olds is a mismatch that changes nothing. Evidence about salivary fluoride applied to a claim about tooth decay is a mismatch that changes everything. I try to keep the two steps visibly apart, and where I rate certainty down for indirectness I say what about the mismatch I think matters.

This is the single most common problem I found, and it is invisible unless you go and look, because a citation looks the same whether or not it is on-topic.

4. Appraising it. With the appropriate tool, and with the judgments shown rather than summarized:

What I am appraising Tool
An individual randomized trial RoB 2
An individual non-randomized study of an intervention ROBINS-I
Whether a systematic review was well conducted AMSTAR 2
Whether a systematic review is at risk of bias ROBIS
A body of evidence for one outcome GRADE

These do different jobs and it is worth saying so, because in the original blog series I ran GRADE, AMSTAR 2 and ROBIS over the same review and reported three separate damning verdicts, which reads as piling on. AMSTAR 2 and ROBIS overlap heavily; both are about the review. GRADE is about the evidence, and would apply even if the review were flawless. I now use one review-level tool per review, say which and why, and reserve GRADE for the body of evidence.

Where a tool has critical domains, as AMSTAR 2 does, I name which specific domains failed rather than reporting a count. “Flaws in five domains” is not a finding; AMSTAR 2’s overall rating depends on which of its seven critical domains were among them.

5. Is there better evidence? A search, recorded so it can be repeated: database, full strategy, date run, number of hits, and what was screened in. “I searched PubMed” is not a method. The strategies are in appraisals/searches/ in the repository, one file per chapter that reaches a verdict.

Searches were run to a stated cut-off date, given per chapter. Anything published after that date is not in this book, and the honest thing is to say so rather than to imply currency the book does not have.

6. The verdict. A fixed box with four fields, the same four every time:

  • Certainty of the evidence, for this component, on my reading, with reasons.

  • Directness to the advice as worded, from step 3.

  • Is the strength label defensible?, answered using the four categories from Chapter 2: higher certainty than stated, a legitimate discordant recommendation, a good practice statement in disguise, or one that should have been Conditional.

  • What would change my mind. Discussed below.

    WarningWhat that third field can and cannot claim

    Strength, as Chapter 2 insists, depends on four things: the size of desirable and undesirable effects, certainty in those estimates, patients’ values and preferences, and resource use. A full evidence-to-decision assessment weighs all of them, plus acceptability, feasibility and equity.

    I have not run one. I have no data on how patients value the burden of brushing at a particular hour, no cost analysis, and no survey of acceptability. What I can assess is the evidential half: what the certainty is, whether the cited evidence is direct, and whether the published reasoning supports the published label.

    So when a verdict says a Strong label is not defensible, it means: on the evidence and the reasoning the guideline has published, I cannot get there. It does not mean the panel, with the values and resource judgments it holds and I do not, could not. Where I think a discordant Strong recommendation is defensible anyway, as for spitting rather than rinsing, I say so and give the reason.

7. What this means for you. One paragraph, plain language, no hedging beyond what is honest.

Effect sizes

Every effect I report gets a point estimate, a confidence interval, and the outcome it is measured on. Not one of these three is optional, and in the original blog series all three were sometimes missing. “14% fewer” and “0.41 fewer surfaces” and “0.28 standardised mean difference” all appeared without intervals, which invites the reader to treat a noisy estimate as a fact.

Standardized mean differences are converted back to natural units wherever the review reports enough to do it, because “0.28 of a standard deviation” is not a quantity anyone has intuitions about, and judging it with adjectives is worse than not judging it at all.

Where an outcome is common, as dental caries is, an odds ratio is not a risk ratio and I say so. An odds ratio of 1.5 for an outcome affecting a third of children does not mean a 50% higher risk.

What would change my mind

Every verdict box ends with this field, and it is not decoration.

A criticism that no evidence could answer is not a criticism, it is a posture. If I say the evidence for brushing at bedtime is indirect, I should be able to state what direct evidence would look like: a trial randomizing people to brush last thing at night versus two hours earlier, everything else held constant, with caries or pain measured at three years. If such a trial exists and I missed it, that field tells you exactly what to send me.

It is also a discipline against a specific failure mode. It is easy, once you have found three weak citations, to start reading every new citation as weak. Writing down in advance what would satisfy me is the only guard I have against that.

What I got wrong before

This book began as four blog posts, and I have corrected several things in moving them here. They are listed rather than quietly fixed, because a book about other people’s citation practice should show its own errata:

  • I described the strength of a recommendation as “very low,” which confuses strength with certainty. Chapter 2.
  • I rated a body of observational evidence as low certainty without considering GRADE’s three criteria for rating observational evidence up. Considering them properly does not change the answer, because the review reports two overlapping dichotomies rather than an ordered gradient, but failing to ask was a genuine misapplication. Chapter 5.
  • I gave the Swedish rinsing trial as having 131 participants in one post and 369 in another. Both numbers appear in the literature because they refer to different things: 369 were randomized and 281 completed, of whom 131 were in the test groups (2). There is also a doctoral thesis reporting the same trial alongside other sub-studies (3), and the two are routinely conflated. Chapter 7.
  • I said there was “no evidence” for the two-minute brushing rule. There is no evidence on caries. There is evidence on plaque removal. Stating absence of evidence loosely is the error I accuse the guideline of. Chapter 18.
  • I said the 2019 Cochrane review showed “we still don’t have any trial evaluating the impact of flossing plus toothbrushing versus toothbrushing alone.” It includes fifteen such trials. What it has none of is a trial measuring interproximal caries. I had written that from the abstract and from memory of it, and reading the full review is what caught it. It is the largest factual error in the original series. Chapter 17.
  • I attacked a twelve-person cross-over study for being underpowered. For a salivary pharmacokinetic cross-over, twelve may be perfectly adequate; the fatal problem is indirectness, not power. Leading with the weaker argument made the stronger one look like piling on. Chapter 6.
  • I called a post-hoc regrouping of four categories into two “p-hacking.” It is selective analysis, or analysis decided after seeing the data. P-hacking means something more specific. Chapter 7.
  • I attributed Delivering Better Oral Health to Public Health England, which has not owned it since 2021. The edition the blog series audited (4) is also not the edition this book audits (1).

What I got wrong in the drafts of this book

The chapters above were reviewed adversarially before publication, by a model instructed to find factual errors, GRADE misapplications, overstatement and unfairness, with the archived guideline and the appraisal data available to check against. It found a great deal. The substantive corrections:

  • I described SIGN 138 as attaching no grade to the bedtime-brushing recommendation, and as citing a single source for it. Both wrong. SIGN cites two sources, and marks the recommendation with the symbol its own key defines as a Good Practice Point, “recommended best practice based on the clinical experience of the guideline development group.” That is a better fact than the one I had, and it makes the argument sharper rather than weaker.
  • I read Kumar’s two brushing-frequency thresholds as the marginal effects of a first and a second daily brushing. They are two different dichotomies of the same variable, from overlapping but non-identical study sets, with no test for trend and no interaction test. The inference does not follow. Chapter 5.
  • I wrote “I did not rate down further, because low is already low.” GRADE runs to very low. That is a straightforward error and it is corrected in place.
  • I said the guideline “cannot be interpreted” and that a Strong label on a bundle communicates confidence about “at least one” component. DBOH says the main component determined the strength; it does not define the label’s scope that way, and it never identifies which component is main. My identification of fluoride toothpaste as the main component is an inference and is now marked as one. Chapter 3.
  • I quoted recommendations as verbatim while merging wordings that differ between tables: table 1a says “on one other occasion,” not “on at least one other occasion,” and table 1b omits “with water” from the rinsing bullet. In a book whose method is verbatim quotation, that is not a small thing.
  • I called Kumar 2016 direct evidence that answers the question while dismissing SIGN’s cross-sectional studies. Both bodies are observational. The difference is real but narrower than I made it.
  • I assigned RoB 2 judgments of High to two domains of the Swedish trial on general impression rather than by working through the signalling questions. One moved to Some concerns on doing it properly. The overall judgment did not change, but the basis is now stated.
  • I filed the Scottish study’s uncontrolled confounding under indirectness. Confounding is risk of bias. Indirectness is about population, intervention, comparator and outcome.
  • I claimed a “minimal adjustment set” from a DAG I did not publish. A minimal adjustment set is minimal only relative to a stated graph. It is now a list of candidate confounders.
  • I concluded chapter 7 by saying “spit, don’t rinse” has a randomized trial behind it that found less decay, four pages after explaining that the trial tested a four-part bundle whose components cannot be separated. That is the same attribution error I accuse the guideline of, in my own conclusion.
  • Several absence claims exceeded the searches that supported them: “nobody has measured,” “never been tested,” “everything published since 2000.” A PubMed-only search with one screener supports “I did not find one.”
  • I described a null result for systemic fluoride absorption in ten people as “reassuring on safety.” Failure to detect a difference is not evidence of equivalence.
  • I discussed a cluster-randomized trial’s results without citing it, from memory of the blog post. The passage is removed rather than patched.
  • I said the six-bullet brushing structure recurs across age groups. The bundles are five, six, three and three.
  • I said Good practice means “no research behind it at all.” DBOH’s definition is narrower and explicitly permits extrapolation from related research.

The second pass, and the errors it found

The finished manuscript was then read again, chapter by chapter, by two independent models working from the same brief and with the archived guideline, the appraisal data and the supplied full texts available to check against. They were told to verify every quotation and every number against its source rather than against my summary of it. They agreed on several findings and each caught things the other missed. The substantive corrections:

Searches that were not searches. Three of my update searches queried a Cochrane review by its identifier and returned nothing, and I recorded the nothing as meaning no update existed. Two updates existed. Wong and colleagues updated the fluoride-and-fluorosis review in June 2024 (5), and Walsh and colleagues updated the oral-cancer diagnostic review in July 2021 (6). Both were published before the edition of DBOH this book audits, and both change what the relevant chapter can say. My brushing-frequency search was worse: it was a known-item query capped at 2018, and could not have found the 2025 review it missed (7). All three failures are recorded in appraisals/searches/ rather than quietly re-run.

Nineteen search records that did not exist. Chapters pointed the reader at appraisals/searches/chNN-*.md files for nineteen chapters where no such file had been written. That is the failure this book exists to object to, committed by the book, and scripts/check_citations.py now fails the build if a chapter names a record that is not there.

A fourth pass, and the chapter it demolished. A further independent review read the corrected manuscript and went after the corrections themselves. Its most important finding was that Chapter 32, which I had added in the third pass, rested on a misread date. Every DBOH page displays “Updated 10 September 2025”; I treated that as the date the evidence was reviewed. The publication’s change log says the 2025 change added case studies and improved the formatting, and that the last full evidence review was 21 September 2021.

Measured against the right date, the chapter’s headline statistic reversed. I had said the median recommendation rested on seven-year-old evidence with sixteen rows over a decade old. Against the 2021 review it is three years, and no row exceeds ten. The chapter now reports that reversal rather than the original claim, and the extraction script measures against 2021.

Three further claims in that chapter went out with it: that varnish trials had stopped for want of funding, which I never investigated and which a 2020 randomized trial contradicts; that the guideline and I had failed “by the same method”, when DBOH’s per-row search procedures are not published; and a prediction that this book would stay current for about seven years, which was a guess formatted to look like a calculation.

Fourth pass: absence claims that were about a review’s scope, not the literature. Chapter 24 said there is no adult fluoridation evidence. The Cochrane review asks what happens when fluoridation starts or stops, so a study of continuous exposure cannot qualify; LOTUS matched 6.4 million NHS dental patients and found a very small effect its own authors report as smaller than most of their stakeholders considered meaningful. Chapter 17 said the decay outcome had never been measured. Hujoel and colleagues pooled six trials in 808 children in 2006; professionally delivered flossing cut interproximal caries by about 40%, self-performed flossing in adolescents did nothing, and the review states that no adult or unsupervised trials exist. Both chapters now describe the actual gap.

Fourth pass: a guideline statement cut before its exception. DBOH’s evidence for prescription-strength toothpaste ends “Moderate-certainty evidence for effectiveness of 5,000ppm fluoride for root caries.” Chapters 8 and 29 quoted the first half and concluded the trials were not there. The root-caries exception is restored in both.

Fourth pass: a prevalence presented as an attributable harm. Chapter 25 set roughly 12% fluorosis of aesthetic concern beside 0.24 of a baby tooth as though the pair were a balance sheet. A prevalence under an exposure is not the excess caused by that exposure, and a proportion of people and a count of teeth are not commensurable anyway.

Fourth pass: non-significance still reported as absence. Several chapters explained the distinction and then went on making the old claim in a verdict or a closing paragraph. Chapter 16’s oscillating-rotating subgroup is the clearest: SMD 0.07 (95% CI −0.20 to 0.33) is imprecision, not proof of no effect.

Fourth pass: my own GRADE reasoning, in the chapter that complains about GRADE. Chapter 5 declined to downgrade for risk of bias on the grounds that doing so would double-count the confounding already considered under the upgrade criteria. Declining an upgrade is not a downgrade already taken, and confounding and exposure misclassification are separate domains. The rating moves to very low.

Fourth pass: an odds ratio reported backwards. Chapter 10 gave Weintraub’s comparisons as though varnish arms had more decay. The published comparisons take counseling alone as the exposure. The direction is now stated explicitly and the reciprocals are shown.

Fourth pass: an argument that shared bias is harmless. I wrote in Chapter 11 that a bias every trial shares cannot explain a difference between arms. That confuses a bias shared across studies with one that cancels within them, and pooling does not remove a common direction of error. Retracted in place.

Fourth pass: grading a test result rather than an effect. Chapter 26’s verdict rated certainty in “no difference was detected”, which describes an analysis rather than an estimate. It now states the risk difference and interval and rates certainty in that. The claim that equal attrition protects a comparison is also gone: equal rates are not equal missing information.

A third pass, by a fourth reader. After the second pass the manuscript was read again, in full, by an independent model working from a prepublication-audit brief rather than the fact-checking brief given to the first two. It found things both earlier passes had missed, and the corrections below marked as third-pass are its.

Third pass: an evidence table joined to the wrong age group, thirteen times. My extraction script matched chapter 2 to chapter 13 on recommendation text alone. DBOH reuses the same sentence across age groups while giving each age group its own certainty rating, so an exact text match returned whichever population came first in the document and scored it 1.00. Thirteen of the ninety-one rows carried another population’s evidence statement at full confidence. The worst case gave an adult recommendation a child’s very low certainty rating where the guideline says moderate. The matcher now maps chapter 2’s tables to chapter 13’s by population, allows the higher-risk tables to inherit from the base tables as the guideline does, and flags rather than guesses when text alone cannot decide.

Third pass: an absolute risk that was not a transformation of its own odds ratio. Chapter 11 reproduced three absolute-risk illustrations from a Cochrane abstract without checking them against the odds ratio printed in the same sentence. They do not follow from it, and the discrepancy is in the source. The chapter now quotes the review, shows the conversion, and labels the recalculated figures as mine.

Third pass: a retention figure that was not retention. Chapter 26 called the ratio of the analysis set to the randomized total “92.5% retention.” It is the proportion who attended at least one visit. The trial reports 67% attending a final visit. The chapter also failed to say that children presenting with pain or sepsis were excluded from the trial, which is the single most important thing a parent needs to know before reading it.

Third pass: a sugar review update, missed the same way as the others. Chapter 13 recorded that no review superseded Moynihan and Kelly 2014. A ten-year update was published in 2022 and had raised the evidence for the 5% threshold from very low to low.

A recommendation left unaudited. Chapter 21 stated that tobacco carries five Strong recommendations across two tables and then audited three rows from one of them. The true count is seven rows: six Strong and one Conditional. The Conditional row, DBOH-074 on e-cigarettes, was neither quoted nor searched. It turned out to be the most consequential row in the guideline for this book’s argument, because its cited review had been superseded seven times and had moved from low to high certainty before DBOH went to press. That finding is now Chapter 32, which is new in this pass.

A quotation cut before the clause that complicated it. Chapter 25 quoted the Key Points of the fluoride and IQ meta-analysis and stopped at a semicolon, omitting the clause reporting that among low risk-of-bias studies the inverse association held below 1.5 mg/L for water as well as urine. The chapter then concluded that at fluoridation concentrations “the signal is not there.” Both the quotation and the conclusion are corrected.

An evidence statement attached to the wrong recommendation. Chapter 28 quoted the adult recall recommendation and then quoted the children’s evidence statement against it, which made the guideline look far more pessimistic than it is. DBOH rates the adult row moderate and states the INTERVAL finding. The chapter’s argument changed accordingly.

A GRADE domain the wrong way round, again. I wrote that the moderate rating for fissure sealants “already reflects a downgrade” for unblindable detection bias. The review downgraded for indirectness, and says explicitly that it did not downgrade for risk of bias. Chapter 11.

An absence that was not one. I wrote that no trial has compared fluoride varnish application frequencies. At least two trials inside the very review I was citing randomized frequency, one of them reporting a dose-response (8). Chapter 10.

A failed join that became a factual claim. Chapter 23 said nothing in the guideline discloses that identical brushing advice carries different labels for different diseases. Chapter 13 discloses exactly that, in one sentence. My extraction script had failed to match the row and I wrote the chapter from my own spreadsheet instead of from the source.

Brand subgroups reported as headline results. The two plaque figures in Chapter 16 were a Procter and Gamble subanalysis and a Colgate subanalysis, not the review’s overall estimates, in a chapter whose main complaint is industry funding.

Silent edits inside quotation marks. A misquoted CMO drinking guideline that turned a risk-minimising statement into a flat limit; a guideline typo I had tidied up; several omissions marked with no ellipsis; and a shortened cancer referral list. Each is corrected in place and the method chapter’s promise of verbatim quotation now holds.

Medical advice beyond the evidence. A suggestion that worried parents might switch to filtered or bottled water, in a chapter whose own verdict is that the question is unresolved; an unqualified instruction to accept fluoride varnish, for a product with listed contraindications; and a claim that people whose gums bleed were not in the scale-and-polish trials, when the larger trial enrolled BPE 0 to 3 and balanced its arms on gingival bleeding. All three are removed or bounded.

Certainty ratings that were mine rather than the sources’. Very low where Ranzan and DBOH both say low; “high to moderate” as though a certainty rating could be a range; low where Cochrane says very low for interdental brushes; and consistency listed as a GRADE upgrade criterion, which it is not.

What that list is for

First, it is long, and it was produced against chapters I had already written carefully and already corrected once. I include it because a book making this argument has no standing to hide its own corrections, and because the pattern in it is the same pattern the book documents in the guideline: not carelessness, but a citation followed one step less far than it should have been.

Second, both passes disputed judgments I have kept, on the grounds that the disputes are about interpretation rather than fact. Where that happened I have said so in the chapter rather than quietly holding my ground.

Third, and least comfortably: the errors above were found by checking my citations against their sources. That is precisely the method of this book, and it took two outside readers to apply it to me. If the argument here is right, nobody should be exempt from it, including its author.

What I cannot do

Some limits worth stating plainly.

I am one person. A proper systematic review of guideline recommendations would have two independent assessors and a documented disagreement procedure. I have neither. Every judgment in this book is one person’s, and the appraisal data is published partly so that a second assessor can disagree with me in public.

I could not obtain every source. The ORCA abstract behind the bedtime-brushing recommendation is a single page in a conference supplement; there is no full report to read, and that is itself part of the finding.

I read English. Where a guideline cites work in another language I have relied on the English abstract, and I flag it where it happens.

And I am not neutral. I think oral health guidance should say what it knows and what it does not, and I started this project already annoyed about a particular recommendation. Declaring that is better than pretending otherwise. The defense against it is not my good intentions, it is the fact that every judgment I made is published in a file you can open.

Reproducing this

git clone https://github.com/choxos/OralHealthBook
cd OralHealthBook
make extract   # rebuild appraisals/ from the archived guideline
make check     # syntax checks: citations resolve, no verdict is blank
make web       # build the book

The archived copies of Delivering Better Oral Health and SIGN 138 used throughout are in sources/, so the extraction can be re-run against exactly the text I read, even after the guideline is next updated.

What make check does and does not do. It is a syntax check. It confirms that every @key in the text resolves to an entry in the bibliography, that no key is defined twice, that every named search record exists, that every cross-reference points at a section that exists, and that no Verdict box has an empty field. It does not confirm that a citation supports the sentence attached to it, that a search was adequate, that a quotation is accurate, or that a number was transcribed correctly. Those failed repeatedly in drafts that passed this check, and every one of them was caught by a person reading the source. A passing build is evidence that the scaffolding is intact, not that the science is right.