The Lab can check a FlowScript document against 33 reporting checklists, from CONSORT and PRISMA to STROBE, TRIPOD, GRADE and AMSTAR 2. Each checklist runs as a set of executable checks over your typed model: it reconciles the numbers you declared (does the flow diagram add up, does the stated ICER follow from the costs and effects) and reports which fields the standard asks for that your document does not yet carry. Two places in the Lab use those checks. The Reporting standards pane scores the open document item by item, takes you to the line each item is about, and applies the repairs that ship with the checklist. The Standards bench turns a checklist into work: it starts you from a document already shaped for the standard, and proposes one thing to do about every item that is not yet reported, refusing out loud whenever the answer is a judgement or a fact only you have. This page documents both, the rules behind every verdict, and all 33 checklists with their versions, years, scope and starting documents.
At a glance
| Topic | What you need to know |
|---|---|
| Checklists | 33: 31 based on published reporting standards, extensions and appraisal tools, plus 2 clearly labelled flowss syntheses that are not published standards |
| Which ones apply | Chosen from your document itself (its blocks, its design words), most relevant first. Extensions are listed next to their base checklist |
| Verdicts | Reported, Partial, Unknown, Missing and N/A |
| The headline number | met / applicable as two whole numbers. There is no single percentage. Unknown counts as not reported |
| Reporting standards pane | Document status popover (⇧⌘M) ▸ Reporting standards ▸ Score against CONSORT, PRISMA…, or ⌘K Reporting standards… |
| Standards bench | Rail: Standards, or ⌘K Go to ▸ Standards |
| Repairs | A checklist's own fix writes placeholders (TODO markers, n: 0 // TODO), never data. Applying every fix makes a report longer and more honest, not better |
| Starting documents | 18 starters serve 22 checklists directly; 8 extensions start from their base checklist's starter; ROBINS-I, ROBIS and AMSTAR 2 get a generated structure |
| Example data | Every starter is a fictional study. The canvas and every export are stamped Example data — not real until you delete the starter's "EXAMPLE text" notice |
| AI and network | None. Scoring, repairs and scaffolds run in your browser |
On Windows and Linux, read ⌘ as Ctrl.
How scoring works
What a checklist can and cannot judge
A checklist here is data plus checks over your compiled model. That makes two kinds of question decidable: arithmetic (do the counts in a flow diagram reconcile, does a stated ICER, interval or net benefit follow from the numbers beside it) and completeness (is the field the standard asks for present and filled in with something other than a placeholder). It does not grade writing, and it cannot tell whether a sentence in your manuscript says what an item wants. A high count of reported items is a statement about your figure model, not a guarantee about your paper.
The five verdicts
| Verdict (chip) | Meaning | Counted as |
|---|---|---|
| Reported | The document carries what the item asks for | Met |
| Partial | Some of it is there, or a placeholder is waiting to be replaced | Not met |
| Unknown | The engine cannot decide from the figure. The detail names the attribute you could add to make it decidable | Not met |
| Missing | The item is not reported | Not met |
| N/A | The item does not apply to this figure, with a justification on the row | Excluded from the denominator |
Unknown is a real answer and it counts as not met. On a figure model, most manuscript-shaped items land there. That is deliberate: a score should never read as "you may stop checking". If a check itself fails while running, the item is reported as Unknown and flagged, so a broken check lowers the score rather than disappearing from it.
No percentage
The pane shows met / applicable as two integers (for example 31 / 37), a bar with all four counted buckets side by side, and a sentence naming the unknown items as work you still have to confirm by hand. It never shows one confident percentage. If a report's own counts ever fail to add up to its denominator, the pane says so: "This report's counts do not add up to its own denominator — treat the numbers above as unreliable and read the items instead."
Repairs never invent data
Many items ship a fix: a source-to-source repair that writes an obvious placeholder into the right block (registry: "TODO — …", n: 0 // TODO, placeholder: true). Every check ignores a placeholder when deciding Reported. So a fix moves an item from Missing to Partial ("replace the placeholder") or, on checklists such as CHEERS and SQUIRE, to Unknown, but never to Reported. You still have to write the real value.
When scoring runs
Scoring runs only while the Reporting standards pane, the Advisor or the Standards bench is open, and it re-runs as you edit. Only the checklists that apply to your document are scored; the others are listed with a reason. Scoring needs a compiled model of your document. A document with ordinary validation errors still has one, built from the blocks that parsed, and is scored. When the compiler cannot read the document at all (an unterminated string while you are halfway through typing a quoted value is the usual cause), nothing is scored, and the Lab says so rather than reporting that no checklist applies.
Which checklists apply
Each checklist has a test for whether it governs your figure. The Lab orders the ones that apply by relevance, so the checklist you most likely came for is first, and keeps each extension next to its base (CONSORT's extensions follow CONSORT; PRISMA-S, PRISMA-ScR and the PRISMA flow arithmetic follow PRISMA).
| Checklist | Applies when the document… |
|---|---|
| CONSORT | Has two or more arm blocks, or mentions randomisation anywhere |
| CONSORT cluster extension | Applies CONSORT and mentions clusters, intraclass or intracluster correlation, or ICC |
| CONSORT non-inferiority extension | Applies CONSORT and mentions non-inferiority, equivalence or a margin |
| CONSORT pilot extension | Applies CONSORT and mentions a pilot or feasibility, or has a pilot_objective block |
| SPIRIT | Has a protocol, spirit_schedule or spirit_row block, or mentions a protocol |
| PRISMA and PRISMA-S | Has three or more cohort blocks and no arm; or uses review words ("systematic review", "scoping review", "rapid review", "meta-analysis", "records screened", "records identified", "studies included", "reports assessed"); or has a search_source or review block |
| PRISMA flow arithmetic | Applies PRISMA and has at least one cohort |
| PRISMA-ScR | Mentions a scoping review, evidence map, mapping review or charting the data |
| MOOSE | Mentions a synthesis (meta-analysis, pooled, systematic review, or has a meta block) and an observational design (observational, cohort study, case-control, cross-sectional, registry-based, exposure) |
| STROBE | Is not randomised (two or more arm blocks plus randomisation wording), has no diagnostic, roc, calibration, decision_curve, predictor or model_equation block, and either declares an observational design in study.design (cohort, case-control, cross-sectional, observational, registry, surveillance, database), or uses wording that identifies one design (such as person-years, matched controls or prevalence survey), or uses exposure and confounding vocabulary together with two or more cohort blocks or an estimate |
| STROBE cohort, case-control, cross-sectional | Applies STROBE and its design is that design |
| RECORD | Applies STROBE and uses routinely collected data (electronic health records, claims or administrative data, registry-based, data linkage and similar), or has a code_list and a dataset |
| STARD | Has a diagnostic block, or mentions an index test, reference standard, gold standard or diagnostic accuracy, or reports accuracy metrics alongside a roc block |
| TRIPOD | Has a calibration, decision_curve or model_equation block, two or more predictor blocks, a roc block without a diagnostic block, or prediction-model vocabulary (prediction model, prognostic model, risk score, nomogram, c-statistic and similar) |
| TRIPOD+AI | Has a performance block (roc, calibration, decision_curve or eval) and either an ml_protocol, ml_search, ml_compute or model_card block, or a model whose family or title names a machine-learning method |
| Model card disclosure | Has a model_card block, or applies TRIPOD+AI |
| ML evaluation figure | Has an ml_protocol, ml_search or ml_compute block; or an eval and a model; or an ablation and a dataset |
| ARRIVE | Has animal or animal_study blocks, or an experiment whose model: names a laboratory species or strain, or uses animal vocabulary alongside experiment, sample, condition or assay blocks |
| CARE | Has a case_report, timeline or timeline_event block, or a title naming a case report, case study or case series |
| SRQR | Has qual_study, theme, quotation, quote or subtheme blocks, or uses qualitative vocabulary (interviews, focus groups, thematic analysis, grounded theory and similar) |
| COREQ (reflexivity and grounding) | Has qual_study, theme or quotation blocks, or uses qualitative vocabulary |
| TIDieR (replication readiness) | Has a tidier_intervention block, or an arm or qi_project that states its intervention: |
| CHEERS (derived) | Has an econ_model, economic_strategy, icer, psa or econ_report block, or two intervention blocks with a numeric cost: and qalys: or effect:, or one such block plus economic vocabulary (QALY, ICER, cost-effectiveness, cost-utility and similar) |
| SQUIRE (cycles) | Has a qi_project, control_chart or run_chart block, or uses improvement vocabulary (run chart, PDSA, special cause and similar) with a series and a kpi or outcome |
| Cochrane RoB 2 | Names the tool rob2, or has risk-of-bias blocks, names no tool, and is randomised |
| ROBINS-I | Names the tool robins_i, or has risk-of-bias blocks, names no tool, and is not randomised |
| QUADAS-2 | Names the tool quadas2, or has risk-of-bias blocks, names no tool, and has a diagnostic block |
| ROBIS | Names the tool robis |
| AMSTAR 2 | Names the tool amstar2, or has amstar_item blocks |
| GRADE | Has a grade_outcome, sof_row or meta block |
"Names the tool" means a tool: attribute anywhere in the document with that value. "Mentions" means the word appears in a block kind, a block id (underscores and hyphens read as spaces) or any string value, so a block called cluster_a is enough to mention clusters. If nothing applies, the pane says: "No reporting standard claims this figure yet. The checklists are chosen from the document itself — declare a study design, arms, a cohort chain with exclusions, or an economic model, and the standards that govern it appear here."
The Reporting standards pane
Opening and closing it
- Open the document status popover (click the status glyph beside the document title, or press
⇧⌘M), and in its Reporting standards section click Score against CONSORT, PRISMA…. Clicking it again closes the pane. - Press ⌘K and run Reporting standards… ("Score this figure against CONSORT, PRISMA and the rest — item by item, with the repairs that ship"), or choose Reporting standards under Go to.
- Close it with its × ("Close reporting standards") or
Escape.
The pane opens in the column at the right of the canvas, which it shares with the Advisor, the methodologist, Comments, Revision history, Provenance and Publish. Only one of them shows at a time.
Header and the standards row
The header reads Reporting standards with a count of how many checklists apply ("N apply"), or an amber not scored chip when the compiler cannot read the document, and the document's title.
Below it, one chip per applicable checklist, most relevant first, each showing the checklist's name and its met / applicable fraction. Hover a chip for its version and scope. Click a chip to open that checklist.
The summary
For the open checklist:
| Part | What it shows |
|---|---|
| Title line | The checklist's name, version and year, such as "CONSORT 2010 (2010)", and the met / applicable fraction in large type |
| Headline | "N of M applicable items fully reported." or "No item on this checklist applies to this figure." |
| Bar | Four segments in order: Reported, Partial, Unknown, Missing |
| Chips | One per bucket with its count ("12 Reported", "3 Partial", "8 Unknown", "4 Missing"), plus "N excluded" for not-applicable rows ("Excluded from the denominator, with a justification on each row") |
| Unknown note | "N items the engine cannot decide from the figure — counted as not reported, and yours to confirm by hand." |
| Warnings | The unbalanced-counts warning above, if it ever applies |
When the compiler cannot read the document, the Lab clears the scores rather than keeping stale ones: the header shows not scored, the summary and tabs give way to "No checklist was run: the document does not compile." ("The engine scores a compiled model, and there is none — an unterminated string is the usual cause, and it happens every time a quoted value is half-typed. Nothing below is a verdict about your figure. Fix the source and the standards that govern it come straight back."), and every checklist is listed underneath as not run.
The three tabs
| Tab | Count shown | Contents |
|---|---|---|
| To fix | Items you can act on | Every Missing, Partial and Unknown item, ordered by what you can do about it (see below) |
| Checklist | Every item | The whole checklist in the standard's own sections and order |
| Doesn't apply | Standards not offered | Every other checklist, with the reason it is not offered |
The To fix tab
Items are ordered so the quickest to resolve come first: items with a repair and a location, then items with a repair only, then items with a location only, then the rest. Within each group, Missing comes before Partial before Unknown, then checklist order. (This differs from a severity-only list on purpose: this list is for working through inside the editor.)
Each item shows:
- its verdict chip, its number as the standard prints it (such as
13a) and its title; - line N with a crosshair when the verdict points at a place in your source. Click the item's header strip to put your cursor on that line in the Source shelf;
- the verdict's detail, in the checklist's own words, and any evidence it quotes from your document;
- "This item failed while running, so the verdict above is not a judgement — check it by hand." when the check itself failed;
- Show the fix when the checklist ships a repair for the item.
To apply a repair:
- Click Show the fix (it becomes Hide the fix while open; one preview is open at a time). The pane previews the repair by running it on a copy of your document: a description, the lines it would add (prefixed
+), any lines it would remove (prefixed-), and a note. - If the engine accepted the trial run, click Apply this fix. If it did not, the pane says "The engine recompiled this change and refused it, so it is not offered."
- The repair is applied and a toast shows its note with Undo.
⌘Zalso undoes it. If the repair is abandoned at this point, an error toast says so (for example "The fix was abandoned — your document is unchanged").
Every preview ends with: "A fix writes placeholders, never data. The item stays outstanding until you replace them." If the report is older than the checklist, the preview says: "This item's fix is no longer available — the report may be older than the checklist it came from."
When nothing is outstanding: "Nothing outstanding on this checklist: every applicable item is reported and none was left undetermined."
The Checklist tab
The checklist grouped into the standard's own sections (for example Title and abstract, Methods, Results), in first-appearance order, with items in checklist order. Each section header is collapsible and shows its own met / applicable fraction, an "N excluded" chip for not-applicable rows, or "none applies" when every row in the section is excluded. Each item shows its verdict, number, title, location, the requirement in plain words, and the verdict detail. An item with a repair that is not yet reported shows Fix available, which switches to To fix with that repair's preview open.
The Doesn't apply tab
Every checklist that was not offered, listed rather than hidden, each with its name and version, its item count, the reason ("Not offered for this figure: nothing in the document matched what this checklist looks for. Covers …"), and Start from this standard buttons for each starter that ships for it (hover a button for the starter's full title). The list opens with "N standards are not offered for this figure. They are listed with a reason rather than hidden — and any of them can be started from a compliant example." When the compiler cannot read the document, the list explains that none of the checklists was run, and each row says "Not run: the document does not compile, so this checklist's own test for whether it governs your figure was never called. It covers: …"
Starting from a compliant example
When the open checklist has a starter and you either have outstanding items or are not following that standard, the foot of the pane offers Start from a compliant example with one button per starter. A starter replaces your document's content. See Starting from a scaffold for what happens and how to undo it.
The Standards bench
Opening it
- Click Standards on the Lab's tool rail. Click it again to close the workbench sheet.
- Press ⌘K and choose Standards under Go to ("CONSORT, PRISMA and the rest as scaffolds and items you can insert").
- Click an Advisor proposal such as "Start from a CONSORT skeleton" or "Score this figure against STROBE". A proposal about one specific item opens the bench on that item, expanded and scrolled into view, with any filter cleared.
The bench has two shelves, switched at its top: Start from a standard and Improve this document. The strap line reads "compiled, never invented".
Standards bench: Start from a standard
This shelf lists all 33 checklists so you can start a document already shaped for one of them.
Search, filter and cards
| Control | What it does |
|---|---|
| Search field (placeholder "trial, review, cohort, qualitative…") | Filters checklists whose name, id or scope contains what you type, as one phrase, case-insensitively |
| Field of work | Every field, or one of nine fields: Clinical trial; Observational epidemiology; Systematic review and meta-analysis; Diagnostic and prognostic accuracy; Health economics; Quality improvement; Machine-learning evaluation; Qualitative research; Laboratory and animal science |
| Count line | "N of 33 checklists", with the reminder "a scaffold replaces the document" |
If nothing matches: "No checklist matches those words in this field. Try one term at a time, or widen the filter to every field."
Each checklist is a card showing its name and version, a not a published standard badge for the two flowss syntheses, a scored badge when the open document has already been scored against it, and its scope. Expand a card to see:
- "N items on the checklist" and the fields it belongs to;
- the full citation, and a link to the standard's website where one exists;
- See what this document is missing, when the document has been scored against it, which switches to Improve this document on that checklist;
- the scaffold: its title, a badge naming where it came from, a scores no better than an empty document badge where that is true, a note saying exactly what the scaffold does and does not claim, a list of every line marked
TODO("N lines marked TODO — every one is yours to write") when it has placeholders, the full FlowScript, and Start a new document from this scaffold.
At the foot of the shelf, once you have opened a checklist the document has been scored against, a link reads "You have already scored this document against X — N outstanding" and takes you to Improve this document.
Where a scaffold comes from
| Badge | Used for | What it claims |
|---|---|---|
| The layer's own starter | 22 checklists | A complete figure for the standard that compiles and scores well. Every string in it is example text from a fictional study; replace all of it. Scoring the untouched starter tells you about the starter, not your work |
| CONSORT's starter, PRISMA's starter or STROBE's starter (the base checklist's name) | 8 extensions: the three CONSORT extensions, PRISMA-ScR, MOOSE, RECORD, STROBE case-control and STROBE cross-sectional | The extension ships no starter of its own because it extends a base checklist. Score both checklists against it, never one instead of the other. Its example text answers the base checklist's items, not the extension's |
| Structure generated by the bench | ROBINS-I, ROBIS and AMSTAR 2 | These tools appraise somebody else's study or review, so no figure template can answer them. The scaffold declares the instrument's published domains or items, leaves every judgement out, and marks everything only you can write with TODO. Its note says why the layer ships no starter and that it compiles as it stands |
The three generated structures:
| Instrument | What the scaffold contains | Placeholders |
|---|---|---|
| ROBINS-I | An appraisal block (effect of interest, unit of assessment, target trial, confounders, assessors, how disagreements were resolved), a study, an outcome, a confounder, the seven published domains as bias_domain blocks with their signalling questions, a rob_traffic_light view, a caption and alt text | 13 lines |
| ROBIS | An appraisal block (review question, phase 1 relevance, unit of assessment, assessors, disagreements), the review as a study, the four published domains with their signalling questions, the three phase-3 questions A, B and C written out in full as notes, an overall note, a rob_traffic_light view, a caption and alt text | 12 lines |
| AMSTAR 2 | An appraisal block, the review as a study, the sixteen numbered amstar_item blocks each carrying what the item asks, with the critical domains (items 2, 4, 7, 9, 11, 13 and 15) called out, and an appraisal_table view | 20 lines |
AMSTAR 2's scaffold carries the badge scores no better than an empty document. Every AMSTAR 2 item is a judgement about a review you have read, so a document that answers none of them scores exactly what an empty one does. That is the honest result, not a defect; the scaffold still saves you writing out sixteen items. Add answer: yes, partial_yes, no or not_applicable to each item yourself.
Starting from a scaffold
A scaffold is a whole document (several blocks, the references between them and a view line), so it is never appended to your work. Appending it would rename its first block to avoid a clash and break every from: reference to it.
- Expand the checklist's card and read the scaffold.
- Click Start a new document from this scaffold (or Start from a compliant example / Start from this standard in the Reporting standards pane).
- The scaffold replaces the content of the open document at once. A toast reads "Template loaded —" followed by the standard's id or the starter's title (for example "Template loaded — consort"), with Undo. There is no separate confirmation dialog (even though the small print under the button says the studio confirms first); Undo in the toast, or
⌘Z, brings your document back. - Data rows bound to the document from the Data bench are released, because they described the study that was there before. The Data bench says: "The rows bound to this document were released when the template replaced it — they described the study that was here before. Load the file again to switch the column checks back on."
- If the open document already holds exactly that scaffold, nothing happens.
Tip: To keep your current document, create a new one first (⌘K New document), then start it from the scaffold.
Standards bench: Improve this document
This shelf works from a compliance report. For the checklist you choose, it lists every item that is not yet reported and proposes one thing to do about each: either a FlowScript snippet the bench has compiled and checked, or a stated refusal saying why the item is yours to write.
If nothing has been scored (no checklist applies, or the compiler cannot read the document), the shelf says "Nothing has been scored yet." and explains: "This shelf works from a compliance report: it reads the items a checklist did not find met and proposes one thing to do about each. Score the open document against a standard first, and the proposals appear here." A button offers Or start a document from a standard.
Choosing a checklist
Which checklist are you working through? lists every scored checklist with its outstanding count, such as "CONSORT — 6 outstanding", where outstanding counts Missing, Partial and Unknown items. Under it:
- a headline, such as "12 outstanding: the studio can start 5 of them, and 7 need a judgement or a fact only you have." (or "Nothing outstanding on this checklist.");
- for the two syntheses, an amber note that the checklist is not a published reporting standard but a synthesis of community checklists, declared as "synthesis 1".
Filters
| Chip | Shows |
|---|---|
| All N | Every outstanding item |
| Studio can start N | Items with a snippet you can insert |
| Yours to write N | Items the bench refused to write |
How far a snippet has been checked
Every snippet is compiled before it is offered, and its holes are written as the checklist's own placeholders ("TODO — …", 0 // TODO, { not_yet_specified: 0 }), so nothing in it reads as an answer. The bench then scores it on its own with the real checks and withholds any snippet that would answer an item by itself. That is the check the Lab runs today, and the bench says so on screen:
| Badge on the item | What was checked | Note shown |
|---|---|---|
| snippet, unchecked | The snippet was scored on its own, not added to your document and re-scored. It answers nothing standing alone, but in rare cases a snippet that supplies the one thing your document lacked can move an item to Reported even though every value in it is a marker | An amber note above the list: "Every snippet below compiled, and each was checked only against the weaker of the module's two promises — that it answers no item STANDING ALONE. … Read each one, and check the item afterwards." |
| yours to write | No snippet. A refusal says why | See below |
Warning: Because snippets are not yet re-scored against your own document, re-open the item in the Reporting standards pane after you insert one. If it now reads Reported while the snippet still holds TODO markers, replace the markers before you rely on the score.
What an item shows
Collapsed: the verdict chip, the item number, its title, the badge above, and the checklist section. Expanded:
| Part | Contents |
|---|---|
| What the item requires | The requirement, in plain words |
| Why this document did not satisfy it | The checklist's verdict detail |
| Repair note | When the checklist ships its own repair: "The guidelines layer ships its own source-to-source fix for this item. Prefer it, from the compliance rail: it knows where in your document the text belongs, and abandons itself if it would introduce an error." Use Show the fix in the Reporting standards pane |
| Merge note | When the snippet's attributes belong on a block you already have: "These are attributes for the study block you already have — … They are shown inside a block so the snippet compiles on its own; after inserting, merge them into your existing block rather than keeping two." |
| The snippet does not cover all of it | Attributes the advice named that the snippet leaves out, each with its reason (below) |
| Placeholder count | "N values in it are marked TODO." followed by "The bench writes markers instead of plausible values so that nothing here reads as an answer — but it could not score this snippet against your document, so it cannot promise the checklist will still report this item as unmet after you paste it. Check the item afterwards." |
| The snippet | The FlowScript to insert |
| Insert as a … block (for example Insert as a study block) | Appends the snippet. Changes to Inserted — insert again to add another |
When you insert a snippet it is appended to the end of your document, any id that is already taken is renamed (nested ids included), the new block is selected, and a toast names the block kind and its id, such as "Inserted cohort 'approached'".
The reasons an attribute is left out of a snippet:
| Reason | What to do |
|---|---|
| The advice names it but not, in words the bench can read, which block it belongs in | Add it by hand to the block the wording names; a guess would put it where the checker cannot see it |
| The advice puts it on a different block from the one the snippet writes | Add it to that block yourself |
| The value named is not one of the words the vocabulary allows for that attribute | Pick an allowed value yourself; a TODO in that slot would be an error |
| Writing it would assert something about your study (a true/false slot has no placeholder form) | Set it yourself once you know the answer |
Refusals: "Yours to write"
A refused item has its own card headed Yours to write, the reason verbatim, and the line: "There is no button here on purpose, and no feature is missing. A bench that filled this in would be writing a method nobody used into your manuscript, and the checklist would tick it." The reasons are:
| Refusal begins | Why |
|---|---|
| "No snippet: item … asks for a judgement, and the checklist offers a choice between answers rather than a value to write." | Only you can decide it after reading the study |
| "No snippet: item … asks for something only you know — a method, a date, a reason, or a number from your study." | There is no attribute that would answer it; write it in your own words |
| "No snippet: the checklist names … for item … but does not say, in words the bench can read, which block it belongs in." | Add the attributes by hand where the wording says |
| "No snippet: the block this item asks for is reported as met as soon as it EXISTS, …" | An empty block would tick the item; write the block yourself with real values |
| "No snippet: the checklist's advice for item … could not be turned into FlowScript that compiles, …" | Follow the wording by hand |
When a filter leaves nothing to show, the bench says which case you are in: "Nothing outstanding on this checklist. Every item it applies to is reported."; "No proposal on this checklist carries a snippet. Every outstanding item wants a judgement or a fact only you have — switch to “Yours to write” to read them."; or "The bench wrote a snippet for every outstanding item on this checklist. Nothing here is waiting on your judgement alone."
The 33 checklists
Item counts are the rows on each checklist, including any that may come back N/A. Version and year are as published (PRISMA 2020 was published in 2021). The last column is the scaffold title the Standards bench shows when you expand the checklist's card: a title ending "— starting from the … figure" reuses the base checklist's starter, and one ending "— appraisal structure" is a structure the bench generates.
Trials and protocols
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| CONSORT | 2010 | 2010 | 37 | Parallel-group randomised controlled trials | CONSORT 2010 — participant flow for a two-arm trial |
| CONSORT (cluster trials extension) | 2010 extension for cluster randomised trials (2012) | 2012 | 19 | Items the extension adds to or modifies in CONSORT 2010. Score alongside CONSORT | CONSORT (cluster trials extension) — starting from the CONSORT figure |
| CONSORT (non-inferiority and equivalence extension) | 2010 extension for non-inferiority and equivalence trials (2012) | 2012 | 11 | Non-inferiority and equivalence trials. Score alongside CONSORT | CONSORT (non-inferiority and equivalence extension) — starting from the CONSORT figure |
| CONSORT (pilot and feasibility trials extension) | 2010 extension for randomised pilot and feasibility trials (2016) | 2016 | 15 | Randomised pilot and feasibility trials. Score alongside CONSORT | CONSORT (pilot and feasibility trials extension) — starting from the CONSORT figure |
| SPIRIT | 2013 | 2013 | 51 | Clinical trial protocols, including the schedule of enrolment, interventions and assessments | SPIRIT 2013 — protocol skeleton with the schedule of enrolment and assessments |
Systematic reviews
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| PRISMA | 2020 | 2021 | 42 | Systematic reviews, with or without meta-analysis | PRISMA 2020 — flow diagram for a review of new searches |
| PRISMA flow-diagram arithmetic | 2020 | 2021 | 7 | The counting relations of the PRISMA 2020 flow diagram, numbered F1 to F7 by flowss (not PRISMA item numbers) | PRISMA 2020 — flow diagram for a review of new searches |
| PRISMA-S | 2021 | 2021 | 16 | The literature search behind a review (an extension of PRISMA 2020 item 7) | PRISMA 2020 — flow diagram for a review of new searches |
| PRISMA-ScR | 2018 | 2018 | 22 | Scoping reviews | PRISMA-ScR — starting from the PRISMA figure |
| MOOSE | 2000 | 2000 | 35 | Meta-analyses of observational studies in epidemiology. Item numbers are the positions of the published bullets, which the original does not number | MOOSE — starting from the PRISMA figure |
A second PRISMA starter, PRISMA 2020 — flow diagram for an updated review, is offered in the Reporting standards pane for PRISMA and the flow arithmetic.
Observational, diagnostic and prediction studies
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| STROBE | 2007 | 2007 | 34 | Cohort, case-control and cross-sectional studies | STROBE 2007 — participant flow for a cohort study |
| STROBE (cohort) | 2007 | 2007 | 5 | STROBE items specific to cohort studies. Score alongside STROBE | STROBE 2007 — participant flow for a cohort study |
| STROBE (case-control) | 2007 | 2007 | 4 | Items specific to case-control studies. Score alongside STROBE | STROBE (case-control) — starting from the STROBE figure |
| STROBE (cross-sectional) | 2007 | 2007 | 3 | Items specific to cross-sectional studies. Score alongside STROBE | STROBE (cross-sectional) — starting from the STROBE figure |
| RECORD | 2015 | 2015 | 13 | Studies using routinely collected health data. Extends STROBE; score alongside it | RECORD — starting from the STROBE figure |
| STARD | 2015 | 2015 | 34 | Diagnostic accuracy studies: an index test against a reference standard | STARD 2015 — diagnostic-accuracy flow with a 2×2 table |
| TRIPOD | 2015 | 2015 | 37 | Multivariable prediction models: development, validation or both | TRIPOD 2015 — prediction-model development with external validation |
Specialised designs
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| ARRIVE | 2.0 | 2020 | 30 | Animal research: the Essential 10 and the Recommended Set | ARRIVE 2.0 — animal-experiment numbers flow and the Essential 10 |
| CARE | 2013 | 2013 | 13 | Clinical case reports, including the item 7 timeline figure | CARE 2013 — clinical case report with a timeline figure |
| SRQR | 2014 | 2014 | 21 | Qualitative research of any approach | COREQ and SRQR — qualitative focus-group study with a coding tree |
| COREQ (reflexivity and grounding) | 32-item checklist | 2007 | 7 | Interview and focus-group research: who did it, whether the sample and the quoted voices agree, whether the coding tree holds together. A consistency layer over COREQ, items R1 to R7, not a replacement for the 32 items | COREQ and SRQR — qualitative focus-group study with a coding tree |
| TIDieR (replication readiness) | 2014 | 2014 | 6 | Any design that delivered an intervention. Checks every intervention including the comparator is described, with numbers rather than adjectives where a second team must act. Items T1 to T6 | TIDieR 2014 — an intervention and its comparator, both described well enough to repeat |
Economic evaluation and quality improvement
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| CHEERS (derived) | 2022 | 2022 | 28 | Health economic evaluations, with the stated ICER, net benefit, frontier and acceptability curve recomputed from the costs and effects in the document | CHEERS 2022 — cost-utility evaluation whose ICER and frontier re-derive |
| SQUIRE (cycles) | 2.0 | 2016 | 39 | Quality improvement as iterative work: theory of change, attribution, the cycles, and whether the chart could detect the rules it declares | SQUIRE 2.0 — quality-improvement report with its cycles marked on the run chart |
Machine-learned prediction models
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| TRIPOD+AI (machine-learning extension) | 2024 | 2024 | 8 | Clinical prediction models built by machine learning. An extension layer over TRIPOD; run it alongside TRIPOD | TRIPOD+AI — a machine-learned prediction model, with the disclosures a deployable one owes |
| ML evaluation figure (flowss synthesis) | synthesis 1 | 2026 | 12 | Machine-learning evaluation figures. Not a published standard: a synthesis of community reproducibility checklists, items S1 to S12 | Machine-learning evaluation figure (flowss synthesis — not a published standard) |
| Model card / clinical-AI disclosure (flowss synthesis) | synthesis 1 | 2026 | 7 | Intended use, out-of-scope use, ground truth, human oversight, monitoring and known failure modes. Not a published standard, items M1 to M7 | TRIPOD+AI — a machine-learned prediction model, with the disclosures a deployable one owes |
Critical appraisal and certainty
| Checklist | Version | Year | Items | Covers | Scaffold title on the bench |
|---|---|---|---|---|---|
| Cochrane RoB 2 | 22 August 2019 | 2019 | 16 | Risk of bias in the results of randomised trials, judged per result | Cochrane RoB 2 — domain-level assessment of one result |
| ROBINS-I | 2016 | 2016 | 17 | Risk of bias in non-randomised studies of interventions | ROBINS-I — appraisal structure |
| QUADAS-2 | 2011 | 2011 | 16 | Risk of bias and applicability in diagnostic-accuracy studies | QUADAS-2 — risk of bias and applicability across three domains |
| ROBIS | 2016 | 2016 | 10 | Risk of bias in a systematic review itself | ROBIS — appraisal structure |
| AMSTAR 2 | 2017 | 2017 | 17 | Critical appraisal of a systematic review of healthcare interventions | AMSTAR 2 — appraisal structure |
| GRADE | GRADE Handbook (updated October 2013) | 2013 | 19 | Certainty of a body of evidence per outcome, and the Summary of Findings table | GRADE — evidence profile and Summary of Findings table |
Checklists by field of work
| Field | Checklists |
|---|---|
| Clinical trial | CONSORT, SPIRIT, the three CONSORT extensions, Cochrane RoB 2, TIDieR |
| Observational epidemiology | STROBE, STROBE cohort, case-control and cross-sectional, RECORD, ROBINS-I |
| Systematic review and meta-analysis | PRISMA, PRISMA flow arithmetic, PRISMA-S, PRISMA-ScR, MOOSE, AMSTAR 2, ROBIS, Cochrane RoB 2, GRADE |
| Diagnostic and prognostic accuracy | STARD, QUADAS-2, TRIPOD |
| Health economics | CHEERS |
| Quality improvement | SQUIRE, TIDieR |
| Machine-learning evaluation | ML evaluation figure, Model card, TRIPOD+AI, TRIPOD |
| Qualitative research | COREQ (reflexivity), SRQR |
| Laboratory and animal science | ARRIVE |
CARE belongs to no field yet, so it appears only under Every field. Four fields of work in the Lab have no published reporting standard in this set (engineering reliability and risk, project delivery, business and finance, and the general figure); for them the standards views stay empty, which is a fact about the checklists, not a gap in your document.
Numbering that is flowss's own
Where a published checklist does not number its rows, or where a checklist is a flowss synthesis or consistency layer, the item numbers are flowss's and must not be quoted as published item numbers:
| Checklist | Numbering |
|---|---|
| PRISMA flow-diagram arithmetic | F1 to F7 |
| MOOSE | 1 to 35, the positions of the published bullets |
| TRIPOD+AI (machine-learning extension) | A1 to A8 |
| ML evaluation figure | S1 to S12 |
| Model card disclosure | M1 to M7 |
| COREQ (reflexivity and grounding) | R1 to R7 |
| TIDieR (replication readiness) | T1 to T6 |
| GRADE | Named rows (such as "starting certainty", "imprecision", "absolute effect") rather than numbers |
| Cochrane RoB 2 | Domains 1 to 5, then named rows (overall, algorithm, coverage, completeness, domains, result, effect, version, support, duplicate, figure) |
| ROBINS-I | Domains 1 to 7, then named rows (confounders, target, overall, direction, coverage, completeness, domains, support, duplicate, figure) |
| QUADAS-2 | Domains 1 to 4 and their applicability rows, then named rows (algorithm, no-overall, question, tailoring, coverage, completeness, domains, support, figure) |
| ROBIS | Relevance, domains 1 to 4, algorithm, the phase-3 questions A, B and C, and overall |
AMSTAR 2's 17 rows are its 16 published items plus one row for the overall confidence rating. RECORD's 13 rows use RECORD's own decimal numbering (1.1 to 22.1).
The example-data mark
Every starter is a complete study with realistic values, and every one is fictional. Each opens with a comment line such as "Every string below is EXAMPLE text from a fictional trial: replace it." While that line is in your source:
- the canvas shows Example data — not real. "This starter uses values from a fictional study — replace them with yours, then delete its "EXAMPLE text" notice to remove this mark from the figure and its exports." Click Got it to fold it to an Example data chip;
- every view carries a faint "EXAMPLE DATA" watermark and an amber Example data — not real label in its bottom-right corner, on screen and in every export;
- the submission bundle's preflight text opens with the same notice.
To remove the mark:
- Replace every example value with your own.
- Delete the comment line that carries the words "EXAMPLE text from a fictional …" (in most starters it begins "Every string below is EXAMPLE text from a fictional …" or "Every string is EXAMPLE text from …"; a few say "from fictional trials" or "from fictional studies"). The mark is triggered by those words anywhere in the source, in upper-case EXAMPLE, so check that no other line still contains them.
- The mark disappears at the next redraw.
The three generated structures (ROBINS-I, ROBIS, AMSTAR 2) carry TODO markers instead of example values, so they are not stamped.
Tips
- Start from the scaffold of the standard you are reporting to, then replace every value. The scaffold is shaped so that each item has somewhere to go.
- Work down To fix in the Reporting standards pane first: its repairs know where in your document the text belongs. Then use Improve this document for the items no repair covers.
- Click an item's header to jump to the line it is about. A verdict with a location is one edit away.
- Treat Unknown as a to-do, not a pass. Its detail names the attribute that would make it decidable.
- Score an extension alongside its base, never instead of it.
- Replace every
TODOa repair or snippet writes. Until you do, the item stays outstanding by design. - Read the citation and follow the link on a checklist's card when you need the published wording; each item's requirement here is a paraphrase.
Limits and known constraints
- The checks judge your figure model, not your manuscript's prose. A sentence elsewhere in your paper can satisfy an item the Lab reports as Unknown.
- Only checklists whose own test claims your document are scored. There is no control to force-score a checklist that does not apply; its row in Doesn't apply explains why and offers its starter.
- PRISMA-NMA, PRISMA-IPD and the separate PRISMA 2020 for Abstracts checklist are not encoded.
- RoB 2's domain 2 is encoded for the effect of assignment only; the adherence variant gets Unknown for its algorithm. The RoB 2 variants for cluster-randomised and crossover trials, and ROBINS-I's 2024 successor, are not encoded.
- PRISMA's "included" box carries two counts (studies and reports), but a
cohorthas onen, so only the study count is checked. - CLAIM's imaging-specific items are not part of the model-card synthesis.
- The AMSTAR 2 scaffold scores exactly what an empty document scores until you record your own answers. The ROBINS-I and ROBIS scaffolds can satisfy structural items (the declared domains, the figure, the caption) but leave every judgement to you.
- CHEERS item 16 is not applicable when there is no
econ_model, and item 25 whenengagementstates there was no patient or stakeholder involvement. - Starting from a scaffold or starter replaces the open document's content and releases any data rows bound to it.
- Snippets on Improve this document are checked on their own only, not re-scored against your document, so every one carries the snippet, unchecked badge; verify the item after inserting it.
- A snippet that belongs on a block you already have is appended as a second block; merge it into the existing block by hand.
Troubleshooting
The pane says "No checklist was run: the document does not compile." The engine scores a compiled model, and an unterminated string while you are typing a quoted value is the usual cause. Fix the source and the checklists come straight back.
My checklist is not in the standards row. Its test did not claim your document. Open Doesn't apply to see its reason, and add what the checklist looks for (see Which checklists apply), such as a second arm for CONSORT or a tool: robis for ROBIS.
I applied a fix and the item is still not reported. That is by design: a fix writes a placeholder. Replace the TODO with the real value.
Apply this fix is missing. The engine ran the repair on a copy of your document and refused it because it would have introduced an error. Make the edit by hand, following the item's detail.
A toast says "No automatic repair ships for … — this one is a manual edit". The item has no repair. Follow its detail in the pane.
A toast says "The document has to compile before a fix can be applied". Fix the errors shown in the document status popover's Validation section (⇧⌘M), then apply the fix.
Improve this document says "Nothing has been scored yet." Either no checklist applies to your document or the compiler cannot read it. Start from a standard, or add the blocks a checklist looks for, or fix the source.
I inserted a snippet and the item turned Reported although it is full of TODO markers. The bench checks snippets on their own, not inside your document, so in rare cases a marker-filled snippet supplies the one block an item was waiting for. Replace every TODO with the real value before you rely on the score.
The figure still says Example data — not real. Delete the starter's "EXAMPLE text from a fictional …" comment line. Replacing the numbers alone does not remove the mark.
I lost my work when I started from a scaffold. Click Undo in the "Template loaded" toast, or press ⌘Z.
