Draft methodology document, version 4 — submissions
The 21 submissions received, published in full with declared interests and secretariat responses.
§2Submissions and responses
21 submissions were received. Each is published in full below with its declared interest, the secretariat response and the disposition. The Institute publishes submissions it did not accept in the same form as those it did.
The treatment of sponsor-conducted evidence is unworkable where all the evidence is sponsor-conducted
The respondent submits on methodology document, version 4.
The respondent states that the draft directs assessors to consider sponsor conduct as a risk-of-bias consideration, and that for several compounds every contributing trial shares one sponsor, so the direction produces a uniform downgrade that distinguishes nothing.
The respondent proposes that sponsor concentration be reported as a property of the evidence base rather than used as a downgrade.
The secretariat accepts this submission in part. Sponsor concentration becomes a reported property of the evidence base. It continues to inform the risk-of-bias assessment where a specific feature of a trial, rather than the identity of its sponsor, supports a judgement.
Sponsor concentration is now reported with every certainty rating as the number of independent sponsors contributing evidence, and a downgrade may not be made on sponsor identity alone but must name the feature of the trial on which it rests.
Quantitative claims are reproduced without the method that produced them
The respondent submits on methodology document, version 4, on a matter that is not specific to this draft but is visible in it.
Several figures in the draft are quoted from sources that determined them by different methods. A figure obtained by one determination and a figure obtained by another are not comparable, and the draft places them in the same sentence without distinguishing them.
The respondent, an analytical chemist, proposes that every quantitative claim carry the method that produced it at the point of use rather than in the reference.
The respondent endorses the general approach taken in submission 001 and asks that it be extended to the matter identified here.
The secretariat accepts this submission. Placing two figures side by side is an implicit claim that they are the same kind of quantity, and in the cases identified they were not.
Every quantitative claim now carries the determination that produced it at the point of use, and figures obtained by non-comparable methods are no longer presented in the same row or sentence.
The framework does not say when a concern warrants one level and when it warrants two
Having read the draft of methodology document, version 4, the respondent puts one point to the committee.
Serious and very serious are both available in every domain and the framework gives no criterion for choosing between them. In practice the choice determines the published rating more often than the choice of domain does.
The respondent proposes worked criteria for the two-level downgrade in each domain.
The secretariat accepts the proposal for three domains and declines it for the remainder.
Worked criteria are now published for imprecision, inconsistency and indirectness. For risk of bias and publication bias the judgement is left to the assessors with a requirement to state the reason, because a criterion expressed in advance would be applied to cases it was not written for.
Imprecision is judged against a relative threshold where an absolute one is what matters
This submission concerns methodology document, version 4 and makes one point.
The framework judges imprecision by whether the confidence interval crosses a relative threshold. For an outcome with a low baseline rate, an interval that crosses that threshold may span absolute differences that are all negligible, and the resulting downgrade is spurious.
The respondent proposes that imprecision be judged on the absolute scale wherever a baseline risk can be established.
The secretariat accepts this submission.
Imprecision is now judged on the absolute scale wherever a baseline risk is available from the contributing trials, with the baseline and its source recorded in the domain note.
References should carry a persistent identifier for every cited source
The respondent notes that methodology document, version 4 has to work where the evidence base is very thin as well as where it is deep, and submits with that in view.
Several references in the draft carry a journal, a year and a volume but no persistent identifier. The respondent states that retrieval of such a reference is materially slower and that identifiers should be supplied throughout.
The respondent asks in the alternative that where an identifier exists but is not carried, the omission be explained rather than left as a gap the reader must interpret.
The secretariat accepts the second limb of this submission and declines the first. Identifiers are supplied wherever the Institute holds one. Where the Institute does not hold an identifier it will not supply one, because a reconstructed identifier that resolves to the wrong record is a worse defect than an absent one.
Every reference without a persistent identifier now carries an explicit statement that the identifier is not held by the Institute, so that its absence is a recorded fact rather than an apparent oversight.
The rating conflates absence of evidence with conflicting evidence
The respondent read methodology document, version 4 in draft. The point applies to it and to the series generally.
The framework applies a single very low certainty rating to two situations a reader will act on differently. In the first, several studies exist and disagree. In the second, no study exists at all. The same badge in both places gives the same signal for two different states of knowledge.
The respondent proposes that the rating be accompanied by a standing formulation identifying which of the two applies, identical wherever it appears so that it can be recognised at a glance.
The secretariat accepts this submission in full. The distinction is real, it is decision-relevant, and the draft did not make it.
A standing formulation has been adopted and is applied to every low and very low certainty rating in the document set. The rating scale itself is unchanged, so that Institute ratings remain readable against published assessments using the same scale.
The treatment of imprecision does not work where the event did not occur
This submission concerns methodology document, version 4 and a convention used across the Institute’s output.
The imprecision rules in the draft operate on the width of an interval. For a harm observed zero times, there is no interval of the kind the rules assume, and the draft gives no instruction.
The respondent proposes that the rule for a zero-event outcome be stated explicitly and that it produce a rating reflecting what the exposure could have detected.
The secretariat accepts this submission. The gap would have been filled by ad hoc judgement, which is what a framework exists to prevent.
The framework now states the rule for zero-event outcomes: the rating is determined by the total exposure and the frequency that exposure could have detected, and the resulting statement records that the evidence is uninformative rather than reassuring.
Interoperability with published certainty guidance must be preserved in any revision
This submission concerns the draft of methodology document, version 4. The respondent assesses evidence for a national body and the observation arises from applying comparable guidance.
The respondent, who has published on the implementation of certainty frameworks across health systems, states that the proposed distinction risks producing a five-level scale that no external assessment can be read against.
Worked examples showing how the distinction would be implemented in three assessment bodies accompanied the submission, together with a request that the final framework state its relationship to published guidance at first use.
The secretariat accepts this submission. Interoperability was the constraint that shaped the revision and the draft did not say so.
The framework states at first use that the Institute rating scale is the four-level scale used in published guidance, and the distinction between absent and conflicting evidence is carried in an accompanying formulation rather than as a fifth level.
The framework does not say what new evidence would change a rating
The respondent’s comment on methodology document, version 4 arises from having applied a comparable framework to the same compounds.
A rating is a judgement about a body of evidence at a date. Without a statement of what would change it, a reader cannot tell whether a newly published trial is material.
The respondent proposes that every rating carry a statement of the finding that would alter it.
The secretariat accepts this submission.
Every rating now carries a statement of what would change it, and the search date against which it was made. A trial meeting the stated description triggers reassessment ahead of the ordinary cycle.
The framework should say when a rating must not be issued at all
Having read the draft under consultation, which concerns methodology document, version 4, the respondent submits as follows.
A certainty rating over an empty evidence base is a rating of nothing, and the machinery of domains and downgrades applied to no studies produces a number that looks like an assessment. The framework should forbid this rather than leave it to judgement.
The respondent proposes an explicit rule: where no eligible study exists for an outcome, no rating is issued and the outcome is recorded as unassessed.
The secretariat accepts this submission without qualification. It identifies a defect the Institute had not corrected.
Where no eligible study exists, no certainty rating is issued, no domain assessment is rendered, and the outcome is recorded as not assessed with the reason stated. The rule is applied retrospectively across the series.
The document should describe how an assessment becomes a decision
The respondent submits on methodology document, version 4. The point would apply equally to any document in the series.
The respondent states that the methodology stops at the certainty rating and that the step from a rating to a decision is where most disagreement actually occurs.
The respondent proposes an evidence-to-decision framework.
The secretariat notes this submission and records the point as correct about the field and not about this document.
No amendment arises. The Institute assesses evidence and does not make decisions, because a decision requires values and a resource context the Institute does not hold. The boundary is stated on the methodology page and the submission has prompted it to be stated at the head of the document rather than in the section on scope.
The document set should be published in translation
The respondent has read methodology document, version 4 and submits on a matter of presentation.
The respondent notes that the assessments concern compounds supplied internationally and that publishing only in English restricts access to the assessment to readers who work in it.
The respondent proposes machine translation of the document set as an interim measure, with human review of the certainty language.
The secretariat does not accept this submission, and records that the underlying point is sound and that the proposed remedy is the difficulty.
A translation whose certainty language has drifted is a different assessment carrying the Institute's name, and the Institute cannot review translations it does not have the capacity to review. The documents remain in English. The submission is published in full because the access problem it identifies is real and unresolved.
Nothing states how many assessors rate a body of evidence
The respondent has read methodology document, version 4 in draft and makes one submission.
A rating produced by one person and a rating produced by two independently with adjudication are different objects, and the framework does not distinguish them.
The respondent proposes that the number of assessors and the adjudication method be recorded on every rating.
The secretariat accepts this submission.
Ratings are made independently by two assessors and adjudicated by a third where they differ. The method and the fact of any adjudication are now recorded with the rating.
The risk-of-bias instrument is not named, so a judgement cannot be reproduced
This is a submission on methodology document, version 4, made from a statistical standpoint.
The framework requires risk of bias to be assessed and does not say against what. Two assessors using different instruments will reach different domain judgements on the same trial, and neither could be said to have applied the framework incorrectly.
The respondent proposes that the instrument and its version be named, and recorded with every assessment.
The secretariat accepts this submission.
The instrument and version are now named in the framework and recorded on every assessment, so that a judgement can be checked against the instrument that produced it.
It is not clear which provisions bind the assessment committee
The respondent submits on methodology document, version 4. A framework of this kind is judged by whether two competent assessors applying it to the same evidence reach the same rating.
The respondent states that the document mixes requirements with descriptions of current practice in the same voice, so that a reader cannot tell which departures would be a breach and which would be a change of habit.
The respondent proposes that binding provisions be distinguished typographically and listed.
The secretariat accepts this submission. A rule indistinguishable from a description is not enforceable and does not reassure.
Binding provisions are now stated in a fixed form, are listed together in an annex, and a departure from any of them must be recorded in the document it affects with the reason, while descriptive passages are marked as descriptions of practice.
The conditions for upgrading observational evidence are too permissive
The draft of methodology document, version 4 was read by a respondent whose concern is its interoperability with published certainty guidance.
The respondent states that the draft permits an upgrade for a large effect without requiring that confounding of the magnitude needed to produce it be shown to be implausible.
The respondent proposes that upgrading be removed from the framework entirely.
The respondent notes that submission 004 has already been made and confines this submission to a matter not covered by it.
The secretariat accepts this submission in part. The conditions are tightened. Upgrading is retained, because a framework that cannot recognise a strong observational signal will misrate the cases where randomisation is not available.
An upgrade for a large effect now requires an explicit statement of the confounding structure that would be needed to produce the observed effect and a reason for regarding it as implausible, and the statement is published with the rating.
Declared interests should appear on the document rather than on a separate page
This submission addresses methodology document, version 4 from the standpoint of a reader who will encounter its output rather than its text.
The draft links to a central conflicts register. The respondent argues that a reader assessing whether to rely on a particular document should not have to leave it to find out who assessed it and what they declared.
The respondent proposes that the interests of every named contributor to a document be printed on that document.
The secretariat notes this submission and records that the draft already provides for it, which the respondent could reasonably have missed because the provision sits in an appendix.
Every document carries the declared interests of its named contributors in its front matter, and the central register exists so that a reader can see a person across all documents rather than one at a time. No amendment arises; the provision has been moved from the appendix into the body of the methodology document so that it is findable.
Preprints and conference abstracts are excluded categorically
This is a submission on methodology document, version 4.
The respondent states that a categorical exclusion removes results that are sometimes the only ones available, and that the reason for excluding them, absence of peer review, is a matter of degree rather than a category.
The respondent proposes that they be included with a downgrade.
The secretariat accepts this submission in part. Such sources are identified and reported and may inform an assessment. They do not contribute to a pooled estimate, because the reporting is generally insufficient for risk-of-bias assessment.
Preprints and conference abstracts are now screened, listed and reported as a distinct evidence class with their status stated, contribute to the narrative assessment, and are excluded from pooled estimates with the reason recorded rather than excluded at screening.
Screening is described in terms that two people would apply differently
The respondent has read methodology document, version 4 in draft and makes a single submission.
The respondent states that the screening criteria in the draft use terms such as clinically relevant without definition, and that agreement between screeners is not measured or reported.
The respondent proposes that criteria be operationalised and that agreement be measured and published.
The respondent supports submission 015 so far as it goes and adds the matter set out here.
The secretariat accepts this submission. A screening decision that cannot be reproduced is a decision the reader cannot check.
Screening criteria are now operationalised as tests that can be applied without further judgement, screening is performed in duplicate, and the agreement observed together with the resolution of disagreements is reported in every review.
Automated assistance in screening should be disclosed and validated
The respondent asks whether any automated tool is used in screening or data extraction, and states that if one is, its performance should be reported in the same way as a human screener's agreement.
The respondent proposes that automated assistance be prohibited.
The secretariat accepts this submission in part. Disclosure and validation are adopted. A prohibition is not, because the alternative to a validated tool is not a human but a smaller search.
The methodology now requires that any automated assistance in screening or extraction be disclosed in the review, that its output be verified by a human for every included record, and that its measured performance against human screening be reported.
Indirectness is defined so broadly that any evidence could be downgraded under it
The respondent read methodology document, version 4 in draft and has confined this submission to a single provision.
The respondent states that the definition covers differences in population, intervention, comparator, outcome and setting without any threshold, so that an assessor who wishes to downgrade can always find a ground.
The respondent proposes that the assessment name the specific difference and state why it would be expected to change the effect.
This point is adjacent to the one made in submission 012 and the respondent puts it in a form the secretariat can act on.
The secretariat accepts this submission. A criterion that can always be satisfied is not a criterion.
An indirectness downgrade now requires the assessor to name the specific difference and to state the mechanism by which it would be expected to change the effect, and the statement is published with the rating so that it can be disputed.
References cited on this page
References are numbered in order of first citation in this document. Each superscript in the text links to its entry below.
- International Organization for Standardization. ISO/IEC 17025:2017 General Requirements for the Competence of Testing and Calibration Laboratories. ISO/IEC Standard 2017;3rd edition. identifier not held by the Institute
Identifiers are reproduced only where the Institute holds them. Where a digital object identifier or PubMed identifier is not shown, the Institute has recorded the journal and year and has not constructed an identifier.