Skip to content

Guides & reference

Screening, extraction and synthesis

Systematic reviews in Evidence: RIS/BibTeX import, PRISMA screening and kappa, de-duplication, quote-grounded extraction, risk of bias, GRADE, living updates.
Sculptural study of connected forms and structured ideas

This page is the systematic-review handbook for the Evidence studio. It follows a review from the search export a librarian hands you, through PRISMA screening with two or more reviewers, de-duplication, AI extraction and the verification that every claim must pass, to risk of bias, GRADE certainty, pooled effects, the summary of findings and a report whose Methods section and PRISMA flow describe exactly what was done. It also explains how a saved review stays current as new papers are published. For the studio's screens, menus and shortcuts, start with Evidence.

At a glance

StepWhereNeeds
Practise on a finished review⋮ Starter reviews… (for example Systematic review and meta-analysis (PRISMA 2020))Nothing: fictional studies, no AI run
Import search resultsSources flyout, References; or ⋮ Import ▸ Import references…A RIS, BibTeX or PubMed (.nbib) file
Fill in missing abstractsAutomatic, by DOI, during importRecords with a DOI; counts toward literature lookups
Keep the corpus on your accountSave as projectSign-in; Starter or above
Screen recordsRail Screen: status per record, votes, reasonsA saved project; owner or editor to change anything
Measure agreementRail Screen: the κ badgeTwo or more reviewers voting on the same records
Remove duplicatesScreen ▸ De-duplicateA saved project; owner or editor
Read the included studiesSynthesize evidence graph; in a saved project, Re-run with new sourcesUp to 50 included sources with text per run; AI
Check every claimAutomatic quote grounding; optional Strict verificationStrict verification uses AI
Correct a study's numbersInspect a link ▸ Backing claims ▸ Edit numbersOwner or editor (or an unsaved review)
Judge risk of biasScreen ▸ Risk of biasA synthesis; on a saved project, owner or editor to record
Rate certaintyAutomatic GRADE, one rating per relationshipNothing; improves as you appraise
Keep it currentCheck for new papers or Re-run with new sourcesA saved project; owner or editor; AI
Report it⋮ Export ▸ Word, HTML or Markdown reportA synthesis

The workflow, end to end

  1. Run your searches in PubMed, Embase, Scopus or any database, and export each result set as RIS, BibTeX or PubMed (.nbib).
  2. Import each file with References. Every record becomes a source awaiting screening; abstracts missing from the file are fetched by DOI where possible.
  3. Save as project. Screening lives on a saved project, and a project can be saved from imported records alone, before anything is read.
  4. De-duplicate. On the Screen panel, De-duplicate marks copies of the same paper across your exports.
  5. Screen. Each reviewer votes include, maybe or exclude on each record; the panel shows agreement (Cohen's κ) and conflicts. Set each record's status as it moves through title-and-abstract and full-text screening, with a reason for each exclusion.
  6. Read the included studies. In a saved project, clear the Find papers field and press Re-run with new sources at the foot of the Sources flyout. Evidence reads only the records you screened in that have text, up to 50 per run, and builds the graph inside the project. The text must be in the Sources flyout: it is there in the session you imported in, and after you reopen the project later you import the same file again to bring it back (see What a synthesis reads). (For a quick review you have not saved, the same button reads Synthesize evidence graph.)
  7. Verify and correct. Check the backing quotes, and type in a study's real numbers where the abstract rounded or omitted them.
  8. Appraise. Record RoB 2 or ROBINS-I judgements for each study. Every certainty rating updates as you go.
  9. Report. Export the report: the Methods section, PRISMA flow and excluded-studies table are drafted from what you actually recorded.
  10. Keep it living. Check for new papers re-runs your saved search for works published since the last check and stores them for screening. Screen them in, and a later check reads the included ones while they are still inside its search window (see Check for new papers).

Importing references

Click References in the Sources flyout, or ⋮ Import ▸ Import references…, and choose one file. The button's tooltip says what it takes: "Import a search export from PubMed, Embase, Scopus, Zotero or EndNote — RIS, BibTeX or .nbib. Records without an abstract are looked up by DOI where possible." While it works the button reads Importing….

Formats

FormatTypical sourceUsual extension
RISZotero, EndNote, Embase, Scopus.ris
BibTeXZotero, Mendeley, Google Scholar, LaTeX libraries.bib or .bibtex
PubMed / MEDLINEPubMed's "Save" as PubMed format.nbib

The file picker offers .ris, .bib, .bibtex, .nbib and plain-text .txt files.

The format is detected from the file's contents, not its name: Evidence scores the text against all three grammars and refuses a file that matches none rather than guessing. EndNote's own tagged format (.enw) is not accepted; export RIS from EndNote instead. Byte-order marks added by Windows exporters are handled.

What happens to each record

  • Every record with a title or a DOI becomes a source, in file order, with its title, authors (up to 100 kept), year, venue, DOI and link. A record with neither is skipped, because it could be neither screened nor de-duplicated.
  • Every imported record starts at Identified — not yet screened. It is not read by a synthesis until you screen it in.
  • Each record gets a stable identity from its DOI, or from its title and year. Re-importing the same file (after a crash, in another browser, into a reopened project) fills in the same rows rather than doubling them. The identity is also scoped to the export, so the same paper exported from two databases arrives as two records, which de-duplication then finds and counts, as PRISMA requires.
  • A record whose file carries an abstract uses it as its text.
  • A record without an abstract but with a DOI is looked up on OpenAlex, 50 DOIs at a time and up to 200 per import. If OpenAlex has an abstract, it becomes the record's text, and OpenAlex's venue and authors fill in what the file lacked.
  • A record with no abstract anywhere is still kept, with no text. It can be screened on its title, and it yields no claims until you add its PDF or paste its text. Its card says so.
  • Publication years are kept for the evidence timeline and the cumulative analysis.

Importing merges into the list you already have. A row that already has text keeps it (you may have edited it), a row with no text is filled in, and identifying details are added where the existing row has none, never overwritten. The pristine blank card is replaced.

If a saved project is open, the new records are also stored with the project straight away, so they survive the next refresh and reach de-duplication. Only records the project does not already hold are sent, and stored records are left exactly as they are. Viewers can import into their own view, but the records are not stored with the project.

The import note

When the import finishes, the note under Sources lists what came in and what did not, separated by dots. For example, with no project open: "Imported 812 records from RIS (.ris — Zotero, EndNote, Embase, Scopus) · 640 with abstracts · 120 hydrated by DOI · 52 need full text (add the PDF) · save as a project and screen them — only screened-in records are read, 50 per synthesis run."

Part of the noteMeaningNext step
N records from FORMATRecords imported, and the format detectedNone
N with abstractsRecords whose file carried an abstractNone
N hydrated by DOIAbstracts found on OpenAlexNone
N need full text (add the PDF)Looked up, and no abstract exists upstreamAdd each PDF or paste its text
N not looked up — DOI lookup hit its rate limit; retry to finish themThe lookup stopped early; these were never asked aboutImport the same file again later; existing rows are filled in
N not looked up — DOI lookup failed (reason); retry to finish themThe lookup failed for another reasonImport the same file again
N past this import's 200-lookup cap — import in smaller files to hydrate themMore than 200 records needed a lookupSplit the export and import the parts
N skipped — reasonRecords the parser had to drop, with the first reasonSee the skip reasons below
N not saved to this project — you have view accessYou are a viewer on the open projectAsk an editor to import, or work in your own copy
N not saved to this project (reason) — they stay in this session's list; import the file again to retryStoring them with the project failedImport the file again
save as a project and screen them — only screened-in records are read, 50 per synthesis runThe next step, when no project is open (signed out it begins "sign in, save as a project and screen them"; with a project open it is just "only screened-in records are read, 50 per synthesis run")Save as project, then screen

If none of the records came with usable text, an error adds: "None of the N records carried usable text — they are in the corpus to screen, but add the PDFs (or import an export that includes abstracts) before synthesizing."

Import refusals and skip reasons

MessageCause
The file is empty.The file has no text
The file is larger than 8000000 characters. Split the export and import it in parts.The file is over about 8 MB
This does not look like a RIS, BibTeX or PubMed/MEDLINE export. Re-export from your database choosing one of those formats.No grammar matched (the note in brackets says why)
The format was recognised but no record in the file had a title or a DOI.Every record lacked both
Skip reasonFix
No recognisable fields in this record — check for a stray record delimiter.Remove the stray delimiter in the file
No title and no DOI, so this record can be neither screened nor de-duplicated. Re-export with the title field included.Re-export with titles
This BibTeX entry's braces never close — look for a missing } .Fix the entry's braces
This BibTeX entry has no { … } body.Fix or remove the entry
Only the first 25000 records of a file are imported. Split the file and import the rest.Split the file

Each field is kept up to 20,000 characters. Each lookup batch of up to 50 DOIs counts as one literature lookup against your plan's daily allowance (see Evidence).

Records and their screening status

Every source carries a PRISMA 2020 status. The Screen panel's dropdown lists them in the order a record moves through the flow:

Status in the dropdownMeaningRead by a synthesis?Counted in PRISMA as
Identified — not yet screenedFound by a search, screened by nobodyNoIdentified, awaiting screening
Awaiting full textPassed title and abstract, waiting for its full-text decisionNoScreened, awaiting full text
IncludedIn the synthesisYesIncluded
Excluded — screeningExcluded at title and abstractNoExcluded at screening
Excluded — full-textExcluded at full-text assessmentNoFull-text excluded
DuplicateA copy of another recordNoDuplicates removed

Sources you paste, upload as PDFs or add from Find papers carry no status. They are treated as included studies: a quick synthesis of a few papers needs no screening. The Methods section knows the difference, and does not claim that records nobody screened were screened (see What goes into the Methods and PRISMA).

Only included studies are read, pooled and graded. A record that is excluded, or still awaiting a decision, is held out of the graph, the pooled estimates and the certainty ratings, and is counted in its PRISMA box instead.

Saving a corpus as a project

Screening needs a saved project: the votes, statuses and reasons are stored with it, and several reviewers can work on it at once. You can save as soon as you have imported records, without running a synthesis, which is how a 300-record import gets screened before anything is read. Use Save as project on the Screen panel, which reads "Save this corpus as a project to screen every source through PRISMA, record include/exclude votes independently of your co-reviewers, de-duplicate the corpus, and watch inter-rater agreement (Cohen's κ) as you go." Once a synthesis exists, Save as project is also on ⋮ and in Share. Once saved, the Screen panel loads the project's records.

Invite co-reviewers from the Team tab (Plus and above). The Team tab is part of the inspector, which opens once the review has a graph. On a project saved from imported records alone, screen in a few records that have text and press Re-run with new sources to build a first graph; Team is then available.

Warning: On a project saved from imported records alone, the empty canvas still reads Ready to synthesize and offers Synthesize. That button starts a separate, unsaved review and detaches the screen from the project; your screening stays in the project, but the graph it makes does not. Use Re-run with new sources at the foot of the Sources flyout instead.

The Screen panel

Open Screen from the rail. Its badge counts records awaiting a decision. The panel is titled Screening (PRISMA) and has two sections: screening, and Risk of bias (once a synthesis exists).

The header

  • De-duplicate runs de-duplication over the project's records (see De-duplication). It shows … while it works.
  • N incl: records included (and records with no status). Always shown.
  • N excl: records excluded at any stage, including duplicates. Shown when there are any.
  • N awaiting screening: records identified or awaiting full text. Shown when there are any; the same number is the badge on the rail's Screen button.
  • The κ badge, once two or more reviewers have voted (see Dual screening and agreement).
  • N conflict: records your reviewers disagree on that nobody has resolved. You see the conflicts on records you have voted on; the owner sees all of them.

The rows below list every stored record of the project, in the order they were stored. There is no search or filter on this list.

Each record

Each row shows the record's title (the full title on hover), with:

  • The consensus dot. Green for include, grey for exclude, amber for conflict, pale for pending. It appears only on records you have voted on yourself, so that your vote stays independent; the owner also sees it on conflicts, which they are there to settle.
  • Your vote: ✓ (include), ? (maybe) and ✗ (exclude), with the tooltips Vote include, Vote maybe and Vote exclude. Your current vote is highlighted. Click another to change it; clicking the vote you already gave changes nothing. Votes measure agreement between reviewers; they do not change the record's status. A vote that does not reach the server is taken off the screen again.
  • The status dropdown, which sets the record's PRISMA status. Changing it is what moves a record into or out of the synthesis.

When a record's status is an exclusion (Excluded — screening, Excluded — full-text or Duplicate), a Reason for exclusion field appears beneath it. Type a reason, or pick one of the suggested reasons, and press Enter or click away to save it. The suggestions keep the same reason spelled the same way, so counts per reason mean something:

  • Wrong population
  • Wrong intervention or exposure
  • Wrong comparator
  • Wrong outcome
  • Wrong study design
  • Not a primary study
  • Full text not available
  • Language

A reason belongs to the decision it explains. When you move a record to a different exclusion stage, the old reason is not carried over; give the new stage its own. A status that excludes nothing carries no reason.

Every status change can be undone from the view island's Undo ("Undo screening of …"), which restores the previous status and reason. If a change is refused, the dropdown goes back to what it was and the reason is shown.

Who can do what

ActionViewerEditorOwner
See the records, counts and κYesYesYes
VoteNoYesYes
Change a record's status or reasonNoYesYes
Set the status of a record the reviewers disagree onNoNoYes
De-duplicateNoYesYes

Dual screening and agreement

Systematic reviews screen independently with at least two reviewers. In Evidence, each reviewer's votes are their own, and the panel reports how much the reviewers agree beyond chance.

How consensus is decided

Votes on a recordConsensus
At least one include and at least one excludeConflict
A maybe beside an include or an excludeConflict
Two or more reviewers, all includeInclude
Two or more reviewers, all excludeExclude
One reviewer so far, only maybes, or no votesPending

A conflict stays open until the project owner changes the record's status on the dropdown. Doing so records the conflict as resolved. If a reviewer casts a new vote or changes their vote on the record afterwards, the conflict re-opens. An editor who tries to set the status of a record whose votes disagree sees: "The reviewers disagree on this record — one voted include and another exclude, or one is unsure where another decided. The project owner resolves a conflict by setting its final status." This stays true after the conflict is resolved: as long as the votes on a record disagree, only the owner can change its status.

Tip: The owner resolves a conflict by changing the status. If the record already has the status you want to keep, choose another status and then choose the one you want: each change the owner makes records the conflict as resolved.

The κ badge

Cohen's κ is computed for every pair of reviewers over the records both of them rated. The headline figure is the average over pairs, weighted by how many records each pair double-screened, so a reviewer who opened two records cannot outweigh a pair who screened two hundred. The badge shows the value, its label and the overlap, for example κ 0.62 · substantial · 140 double-screened. Its tooltip names how many reviewer pairs it averages. With two reviewers the overlap is the number of records both rated; with three or more it is added up over every pair, so it can be larger than the number of records (three reviewers who all rated the same 100 records show 300). The exported Methods section counts distinct records instead.

κLabel
below 0poor
0 to below 0.2slight
0.2 to below 0.4fair
0.4 to below 0.6moderate
0.6 to below 0.8substantial
0.8 and abovealmost perfect

Sometimes κ cannot be computed, and the badge says so instead of showing a number:

  • κ n/a · no records double-screened: two or more reviewers have voted, but never on the same record.
  • κ n/a · both reviewers gave every shared record the same decision: when both reviewers used one category for everything they both rated (for example, both excluded everything), chance alone predicts complete agreement and κ is undefined. That is not a low score and not a measurement. Screen more records, or ones you are likely to decide differently, to get one.

The same figure, the number of reviewers, the records double-screened and the open and resolved conflicts go into the exported Methods section.

De-duplication

De-duplicate on the Screen panel compares every record in the project and marks copies as Duplicate, then reports "Marked N duplicates." The rules:

  • Two records are the same paper when their DOIs match. DOIs are compared in a canonical form, so https://doi.org/10.1000/x, doi:10.1000/x and 10.1000/x match.
  • Without a DOI match, two records are the same paper when their normalised titles match (titles longer than eight characters), at least one of them has no DOI, and their years, where both give one, are within a year of each other. Two records with different DOIs are never duplicates, however alike their titles.
  • Which copy is kept does not depend on the order of the records. A copy you excluded is kept first (your exclusion stands for the paper), then an included copy, then one awaiting full text, then an unscreened one.
  • Each copy marked duplicate gets the reason "Duplicate of ID (DOI match)" or "Duplicate of ID (title match)", naming the record it duplicates, for the exclusion log.
  • Records you have already excluded are never relabelled, and a record already marked duplicate is left alone.

De-duplication is not on the Undo stack. To restore a record wrongly marked, set its status back on the dropdown.

A living re-run applies the same rules to the records it brings in: a new search result that is a copy of a paper already in the project is stored as a Duplicate of the record that stands for it, and its claims are left out. So a paper you excluded once does not come back as a fresh record under a different identifier.

What a synthesis reads

Two buttons read the corpus. On a review that is not a saved project, Synthesize evidence graph builds a new graph from the included sources that have text, up to 50 per run. In a saved project, Re-run with new sources reads the included sources with text that the project has not read yet (or whose text has changed), up to 50 per run, and merges what it finds into the project's graph. Both read only included studies, and both refuse by name before sending anything:

  • Synthesize evidence graph with nothing screened in: "None of the N sources with text has been screened in — only included studies are read. Save as a project and screen them on the Screen tab first — nothing has been sent."
  • Synthesize evidence graph with more than 50 included sources: "N screened-in sources have text, and one synthesis run takes 50. Screen them down first (or remove sources) — nothing has been sent, and nothing was lost."
  • Re-run with new sources with nothing screened in: "None of the sources with text has been screened in — only included studies are read. Screen them in on the Screen tab first." More than 50 due sources are read over successive re-runs (see Re-run with new sources).
  • Sources without text are skipped by both, and each such source's card warns you beforehand.
Note: A saved project keeps its records' screening, bibliography and claims, but when you open it again in a later session the studio does not load the records' text back into the source list. To read records that have not been read yet, import the same reference file again (it fills the text into the existing rows rather than adding new ones), or add the PDFs or paste the text, then press Re-run with new sources.

A retracted source (one OpenAlex marks retracted) stays in the corpus and in the bibliography, but its claims are left out of the graph, the pooled estimates and the certainty ratings.

Extraction

For each source, the extractor is given the source's title, year and venue and its text, and asked for the atomic empirical claims the text makes. Each claim has:

FieldMeaning
Subject and objectTwo short concept names (two to four words)
RelationOne of twelve, shown as: increases, decreases, causes, prevents, treats, correlates with, associated with, no effect on, moderates, supports, contradicts, part of. Common synonyms the model writes (such as "reduces" or "linked to") are mapped onto these, and a converse phrasing such as "moderated by" is turned round so the arrow points the way the paper says. A claim whose relation matches none of them, or that has no quote, is dropped with an extraction note
QuoteThe exact sentence or span from the text that supports the claim
PopulationWho was studied, when stated
MethodThe study design, when stated (RCT, cohort, meta-analysis, observational…)
EffectThe reported effect, verbatim, for example OR 1.8 (1.2–2.7)
nThe sample size, when stated

The extractor is told to extract only what the text asserts, never to add outside knowledge, to record a null finding as "no effect", and to prefer fewer, well-grounded claims. Text inside a paper is treated as data: instructions that appear in a source are not followed.

Details that affect what you get:

  • Length. Each source contributes at most its first 24,000 characters. A longer source is read only up to there, and an extraction note says so, for example "only the first 24,000 of 61,200 characters were read … so this source's claims speak for that opening section only". For a full text, paste the abstract and results rather than the whole paper.
  • Speed and the time budget. Sources are read three at a time, within a budget of about 55 seconds per run. When the budget runs out first, the run is partial: the sources it never opened are named as unread on every surface. The Inspect overview shows a rose N/M sources read chip, the Pipeline turns Ingest, Extract and Verify red, and an export warns you. Re-run to cover the rest, or synthesize in smaller batches.
  • Extraction notes. Per-source notes (a source whose extraction failed, a truncated source) are listed under Runs ▸ N extraction note(s).
  • Points. On hosted AI a run is charged one point per included source it may open, at most 10, and only after extraction has delivered. A run in which extraction failed on every source is not charged.

After extraction, concept names that normalise to the same form are merged into one concept, keeping the variants as aliases (near-duplicates that differ more, such as an acronym or a spelling variant, are offered for merging under Tidy concepts on the Inspect overview). Claims asserting the same relationship are combined into one link, and the graph, contradictions and gaps are computed. Correlation and association are treated as undirected. Every count of "studies" is a count of distinct sources, so a paper that states something twice is one study; and two reports of the same study (sharing a DOI, or the same trial registration number in their title, link or DOI) count as one study.

Verification

Every extracted claim passes a deterministic check before it can enter the graph. A claim is dropped, with a reason, when:

ReasonWhat it means
quote too short to verifyThe quote is under 16 characters after normalising
no source text availableThe source has no text to check against
quote not found in sourceThe quote is not a verbatim part of the source text
the quote leaves out "…" from its sentence, which negates or qualifies itThe full sentence contains a negation or qualifier the quote cut off
the quote's sentence mentions neither the claim's subject nor its objectThe quote is real but is about something else
from a hypothesis sketch, not a publication — it cannot ground itself; add papers to test itThe claim came from a diagram you sent to Evidence

Matching ignores differences in case, Unicode forms, curly versus straight quotes, dash styles and whitespace, and nothing else. A verified claim also records where its quote sits in the source, so it can be located again.

Numbers are checked too. If an effect, a sample size or arm-level counts attached to a claim do not all appear in the source text, they are removed from the claim and set aside. The claim stays, but those numbers are never pooled, graded or exported as data.

The Inspect overview's N/M grounded chip reports how many claims passed. The Verify stage of the pipeline shows the same figure.

Strict verification

Strict verification is in the prompt's AI settings ("A second pass by the flowss Evidence Agent checks each quote supports its claim."). With it on, after the deterministic check the flowss Evidence Agent reads each grounded claim a second time and drops those whose quote does not actually support it. It applies to the next synthesis or re-run, and adds a model call per source, so runs take longer.

  • N rejected (amber chip) counts claims the strict pass removed.
  • N unverified (rose chip) counts claims the strict pass could not check, for example because it timed out. They are kept, quote-grounded, but nothing checked that the quote supports them. The chip exists so a pass that did not finish cannot look like one that checked everything.
  • The Verify stage reads Quote-grounded + semantic when the pass covered every claim, and names how many were unchecked otherwise. Without strict verification it reads Quote-grounded (support not checked).

The Methods section reports which checks actually ran.

Correcting a study's numbers

Abstracts round, omit intervals, and often report counts instead of ratios. On any link, Inspect ▸ Backing claims lists each claim with its source, design, n, effect and arm-level data, then its quote. Owners and editors (and anyone on an unsaved review) see Edit numbers under each claim, which opens Edit the numbers for this finding:

FieldWhat to enter
Reported effect (verbatim)The effect as the paper writes it, for example OR 1.8 (1.2–2.7)
Participants (n)The sample size
DesignThe study design. Suggestions: RCT, randomised trial, cohort, case-control, cross-sectional, observational, systematic review, meta-analysis
Arm-level dataNone, 2×2 — events / total per arm, Means, SDs and ns, or Events / person-time

Choosing a kind of arm-level data shows its boxes:

KindBoxes
2×2Events (intervention), Total (intervention), Events (control), Total (control)
Means, SDs and nsMean, SD and n for the intervention and for the control
Events / person-timeEvents and Person-time for the intervention and for the control

As you type, the form previews what the numbers pool as ("Pools as … — from …"), or why they define no effect ("No effect from these numbers: …"). When the typed effect and the typed counts point in opposite directions, it warns you that the pooled estimate uses the data, so check the arms are not swapped. A half-filled table is refused with the boxes still needed ("Incomplete — still needed: … Part of a table cannot be pooled."). Click Save numbers to apply, or Cancel.

A saved correction changes the pooled estimate, the forest plot, the GRADE rating and the summary-of-findings row at once. On a saved project it is stored; if storing fails, the change is rolled back and "Could not save those numbers — the finding is unchanged." appears. You cannot edit a claim's quote, source or concepts: those are its identity and its grounding.

Risk of bias

Risk of bias is GRADE's first domain and the one only a reviewer can supply. It sits on the Screen panel under Risk of bias, or opens from a link's Appraise its studies (risk of bias) button. It works on an unsaved synthesis too, so you can watch the ratings move; save the review as a project to keep the judgements (the panel reminds you: "Save this synthesis as a project to keep these judgements.").

The panel lists every study that backs at least one claim in the graph (studies held out by screening are not listed), with a line such as "3 of 7 studies have a risk-of-bias judgement. Until every study is judged, the certainty rating is an upper bound." Each row shows a coloured dot, the study and its state:

StateDotMeaning
Low riskGreenThe instrument's overall judgement is low risk
Some concernsAmberSome concerns (RoB 2), or moderate (ROBINS-I)
SeriousRedHigh risk (RoB 2), or serious (ROBINS-I)
CriticalDark redCritical (ROBINS-I)
StartedGreySome domains are judged, but the instrument cannot yet determine an overall judgement
Not assessedGreyNothing recorded

Click a study to open its form:

  1. Instrument. RoB 2 (randomised trial) or ROBINS-I (non-randomised), preselected from the study's design: a randomised trial gets RoB 2, everything else (including a study with no stated design) ROBINS-I. You can change it. If the study already has judgements, switching asks first ("Switch this study to ROBINS-I? The 3 RoB 2 judgements recorded for it will be discarded — the two instruments share no domains."); on a study with nothing judged yet, the choice is simply kept and saved with the first judgement.
  2. Domains. One dropdown per domain, each starting at — not judged —.
  3. The overall line. "RoB 2 overall: Low risk." once the instrument can decide, or "Not yet determined — N domains still to judge." Any inconsistency between recorded answers and judgements is listed in amber.
InstrumentDomainsJudgements offered
RoB 21. Randomisation, 2. Deviations, 3. Missing data, 4. Measurement, 5. Reported resultLow risk, Some concerns, High risk
ROBINS-I1. Confounding, 2. Selection, 3. Classification, 4. Deviations, 5. Missing data, 6. Measurement, 7. Reported resultLow risk, Moderate, Serious, Critical, No information

The overall judgement is the instrument's own published algorithm applied to your domain judgements; Evidence does not form a second opinion. A ROBINS-I "No information" never counts as low risk. Viewers can read the panel but not change it. Clearing a study's last judgement removes its record.

How risk of bias moves certainty

Each relationship's rating weighs every study by its share of that relationship's pooled estimate (or, where nothing could be pooled, weighs the studies equally and says so):

Share of the estimateEffect
20% or more from studies at critical risk−2
50% or more from studies at high, serious or critical risk−2
20% or more from studies at high, serious or critical risk−1
50% or more from studies with at least some concerns−1
OtherwiseNo downgrade

Until a relationship's studies are appraised, its rating says "risk of bias not assessed — GRADE's study-limitations domain was not applied, so this rating is an upper bound", and the Risk of bias cells of the Summary of findings are amber. When only some of the estimate comes from appraised studies, the reasons say what share is unappraised and that the downgrade is a floor. A study at critical risk that carries less than 20% of the estimate does not downgrade the rating, but the reasons point out that the tool's own guidance is to exclude such a study.

How certainty is rated

Every relationship gets one GRADE certainty rating: High, Moderate, Low or Very low. The same rating colours the graph, fills the evidence table and the Summary of findings, drives the certainty filter and appears on the inspector pill, so they never disagree. Its reasons are listed under the pooled analysis when you inspect a link.

Starting level. A body of evidence starts High only when every study behind it is randomised (or is a systematic review of randomised trials). Otherwise it starts Low, including when any study has no stated design.

Downgrades (each at most one level, except risk of bias):

DomainRule
Risk of biasFrom your appraisals, as above (up to −2)
Inconsistency−1 when the pooled I² is 50% or more (substantial) or 75% or more (considerable); otherwise −1 when the relationship is contradicted by other studies
Imprecision−1 when the pooled 95% interval includes no effect; or when the pooled studies total fewer than 100 participants; or, with nothing pooled, when a single study backs it or the total sample is under 100. A body that starts High also loses a level when imprecision cannot be assessed (no effect estimate or no sample size recorded)
Publication bias−1 when the funnel is asymmetric (Egger's test) and at least 10 studies are pooled

Upgrade. A non-randomised body gains one level for a large pooled effect, but only if it has not been rated down on any domain. Large means a pooled odds, risk or hazard ratio of 2 or more (or 0.5 or less), a standardised mean difference of 0.8 or more, or a correlation of 0.5 or more, from two or more studies, with a 95% interval that excludes no effect. A single study's large estimate never upgrades, and neither does a bare percentage.

Not assessed. Indirectness cannot be judged from claims and is omitted.

The How this was calculated button beside the certainty pill opens the full record of the rating: the result, the inputs, what was excluded and why, the method and its version, parameters, assumptions, diagnostics and limitations. Download reproducible JSON saves it.

Select a link and open Inspect to see its full analysis, top to bottom.

  • The pills. The certainty pill, the number of studies (hover for claims across sources), single study — fragile when one source holds the link up, and How this was calculated.
  • The effect summary. A sentence describing the reported effects, amber when they are heterogeneous.
  • Model. random-effects (the default) or fixed-effect. The choice applies to every link you inspect, and changes the forest plot and every analysis below it. The certainty rating is always computed on the random-effects pool, so it does not change with this switch.
  • Forest plot. One row per study with its estimate, interval and weight, and the pooled diamond with its prediction interval. Effects reported only as text without an interval, and without arm-level data, contribute no weight.
  • Pooled line. For example "Pooled (random-effects): OR 1.42 [1.10, 1.83] · I²=38% · 6 studies", with the Hartung–Knapp interval (HK) for a random-effects pool and the 95% prediction interval (PI). A second How this was calculated opens the record of this pool: which claims it weighted and why others were left out.
  • Publication bias. Egger's test for funnel asymmetry. With fewer than 10 studies it reads "Funnel asymmetry not assessable: Egger's test needs 10 studies to have power and this edge has N … That is "not tested", not "no publication bias"". A detected asymmetry is flagged in amber with a warning sign.
  • Trim-and-fill. Shown when it can run or when asymmetry was flagged: how many studies a publication-bias pattern would have suppressed and the re-pooled estimate. It is a what-if under an assumption, not a correction.
  • E-value. How strongly an unmeasured confounder would have to be associated with both exposure and outcome to explain the estimate away, with its interval limit. For odds and hazard ratios an amber caveat states the rare-outcome assumption it rests on.
  • Absolute effect. For ratio measures, the corresponding risk and number needed to treat or harm at a baseline risk. The line says where the baseline came from: "Baseline = median control-arm risk of the N studies reporting one.", "Baseline set by you for this outcome." or "Baseline assumed — no study on this edge reported a control-arm event rate." Type your population's risk into Assumed baseline risk for OUTCOME (a percentage between 0 and 100) to use it for every link into that outcome while the review is open; clear the box to go back to the studies' own figure. The value is not saved with the project and is not used in exports.
  • Subgroups by design. When the studies span two or more designs, each stratum's pooled estimate and a between-groups test, under the chosen model.
  • Funnel plot. The studies against their precision, with any studies imputed by trim-and-fill and the adjusted centre.
  • Influence. Whether dropping any single study moves the pooled interval across the no-effect line ("Robust: …" or "Not robust: …"). Leave-one-out expands to the estimate without each study, and, where the studies have publication years, cumulative by year ("through 2018 (k=4)"). Studies without a year are left out of the cumulative view, and the panel says how many.
  • Certainty reasons. Every step of the GRADE rating as a list.
  • Backing claims. Each claim's source, design, n, effect and arm-level data, its quote, an amber warning when its typed effect and its counts disagree, and Edit numbers.

The Summary of findings

Table ▸ Summary (or the Summary of findings lens) is the GRADE summary-of-findings table, grouped by outcome. Each block lists the determinants of one outcome with Determinant, Relationship, Effect, Studies, Participants, Risk of bias and Certainty, from highest certainty down. A Risk of bias cell is amber when the domain was not applied to all of the evidence in that row, and the footnotes underneath explain what that means for the rating. The same table, with the same footnotes, is in the exported report.

Contradictions, reconciliation and gaps

Two relationships between the same pair of concepts contradict when their directions oppose:

One findingOpposing finding
Increases, causes or treatsDecreases or prevents
Increases, causes or treatsNo effect
Decreases or preventsNo effect
Correlates or associated withNo effect

Moderates, supports, contradicts and part of never form a contradiction. Contradicted links are amber and animated on the graph, and lose a level for inconsistency unless heterogeneity already cost them one.

For each contradiction, Evidence tests whether a moderator explains it: the study design, the population, or the publication year. When one does, the explanation is shown in green under the contradiction (for example "… — the disagreement tracks study design." or "… — a population/subgroup difference, not a true conflict."). When none does, the explanation says what was and was not testable and the contradiction stays open. The exported report lists reconciled contradictions in their own section.

Gaps are pairs of concepts that are both studied in your corpus but never studied together, and relationships resting on weak evidence. The Inspect overview lists up to eight unstudied pairs with a find button that searches OpenAlex for papers linking them. Ask ▸ Suggested searches to run proposes searches aimed at the gaps and contradictions.

Keeping a review living

A saved project can be brought up to date without starting over. The button at the foot of the Sources flyout depends on the Find papers field:

FieldButtonWhat it does
Holds a queryCheck for new papersRe-searches OpenAlex with the query. If it is the project's saved query, the search covers works published since the last complete check, newest first
EmptyRe-run with new sourcesReads the included sources you added or changed since they were last read

Both are available to owners and editors and show Re-running… while they work. Viewers are refused before anything is spent ("You need edit access to run synthesis.").

Check for new papers

  • The saved query is the one in the Find papers field when the project was saved. Its search window starts at the date of the last complete check (or, before the first one, the date of the search the project was saved from, or else the date the project was created), less a 30-day overlap so that papers indexed late are not missed. Works found again in the overlap are already in the project and are skipped at no cost.
  • Each check takes in the newest 10 new works with abstracts. New works are stored awaiting screening: they are not read, and a check that only finds new papers costs no points. A stored paper is also flagged as a Duplicate when it is a copy of a record the project already holds.
  • To read a new paper, screen it in. A later check reads an included paper when it finds it again, which happens while the paper still falls inside that check's window (the window reaches back 30 days before the previous complete check). A paper published earlier than that is not found again; the studio does not hold the text of stored records, so add its PDF or paste its abstract into the Sources flyout and use Re-run with new sources.
  • Included studies whose abstract OpenAlex has updated are read again.
  • When more new works are waiting than one check takes in, a notice says "N more new papers published since DATE are still waiting — this check took in the newest batch. Press "Check for new papers" again for the next one." The project's search date moves only once every new paper in the window is in the project, so nothing falls between two checks.
  • If nothing is new, a notice (not an error) says so: "No new papers: OpenAlex has no work with an abstract matching "QUERY" published since DATE." If every recent work is already in the project: "Nothing to re-read: all N works published since DATE are already in this project — read and unchanged, or stored for screening — …". If the window holds more than one check can page through: "… Narrow the query so the window fits; the search date has not moved."
  • A query that is not the project's saved query is run as a fresh relevance search with no window, and does not move the saved query's date.

Re-run with new sources

  • Only included sources with text are sent, plus any source the project has not stored yet (which is stored with its status).
  • A source whose text has not changed since it was last read is skipped, so you do not pay to read the same abstract twice. If every included source is unchanged, nothing is sent: "Nothing to re-read: the N screened-in sources with text are unchanged since they were last extracted, so re-running would spend points on the same text and produce the same claims. Add a source, edit one, or run a new search."
  • One re-run reads at most 50 sources. Any beyond are not dropped: "N more sources have new text waiting — one re-run reads 50, so re-run again to read them."
  • If nothing has been screened in: "None of the sources with text has been screened in — only included studies are read. Screen them in on the Screen tab first." With nothing to send at all: "Add new sources (or a saved query) to re-run."

What a re-run reports

The new claims are merged into the project and the graph is recomputed. A green Re-synthesis: line under the button says what changed. The Runs tab records the run with its trigger, time, summary and claims added. Several notices protect you from reading a graph that did not move as a finding:

  • A run that read none of its new sources says so ("Re-run read none of the N new sources — … The evidence graph is unchanged because nothing was read, not because these papers added nothing."), with the fix for the cause: check your AI key and model, or re-run to pick up sources the time budget never opened.
  • A run that read only some names the rest ("N of the M new sources were never opened … so this delta covers the K it did read and says nothing about the rest.").
  • Sources skipped because their text was unchanged are named ("N sources (…) were not re-read: their text is unchanged since it was last extracted, so the findings already in this project stand and no points were spent on them.").

After a re-run, the extraction and grounding counts on screen describe that run, not the whole graph. The Verify stage reads not re-checked this run where a whole-graph figure no longer exists.

Retractions

When Check for new papers finds that OpenAlex now marks a paper already in the project retracted, the paper is flagged even if its abstract is unchanged, and its claims are taken out of the synthesis. When the check otherwise had nothing new to read, its notice adds "N of them are now marked RETRACTED by OpenAlex — their claims were taken out of this project's synthesis." A retraction is never undone automatically.

What goes into the Methods and PRISMA

The exported report drafts its Methods section from what the review actually recorded, and refuses to claim work that was not done.

  • Search. "We searched OpenAlex with the query "…" (last searched DATE)." When the date was not recorded, it says PRISMA item 6 is satisfied only in part. Then the records identified and duplicates removed.
  • Screening. When decisions were recorded: records screened on title and abstract, excluded, assessed in full text, excluded there, and included; records still awaiting screening or full text; and how many were counted as included only because they carry no decision. When no decision was recorded at all, the section says that no screening took place rather than describing one.
  • Agreement. "Records were screened independently by N reviewers; inter-rater agreement over the M records rated by two of them was κ=… (label)", followed by the conflicts still open or resolved. When κ could not be computed, it says why and that no κ is reported.
  • Extraction and verification. How claims were extracted and quote-grounded, how many could not be verified, and whether the strict pass ran and covered everything.
  • Risk of bias and GRADE. Which instruments were used and for how many studies a judgement was determined; that indirectness was not assessed.

The PRISMA counts are arithmetic over the statuses, so every box adds up: records identified (every record), duplicates removed, records screened (excluded at screening, plus awaiting full text, plus assessed at full text), excluded at screening, full-text assessed (full-text excluded plus included), full-text excluded, and included. Records not yet screened and records awaiting full text are reported separately rather than hidden in a box. When any record was excluded, the report adds a PRISMA flow section with an Excluded studies table (Study, Stage, Reason), where the stage is Duplicate, Title/abstract screening or Full-text assessment. The HTML and Word reports also draw the PRISMA 2020 flow diagram, and you can send the flow to another studio from the Pipeline view.

If some records counted as included contributed no claim (for example because a run was partial), the report adds a Coverage of the included records paragraph naming them, so the findings are not read as resting on more studies than they do.

Readiness

Evidence summarises how far a review has come with one of five levels. It is computed from the review on screen and shown in the status popover; for the open project it is also remembered on this device, so your Library board can show it beside the project.

LevelMeaning
EmptyNo claims yet
DraftSome claims are not grounded in their source
SoundEvery claim is grounded, but screening or appraisal is incomplete
ReviewedGrounded, screened and appraised, with a contradiction or a very-low-certainty finding remaining
PublishableEvery check passes
CheckPasses when
Every claim quotes its sourceAll claims are grounded in a verbatim quote
Every record screenedEvery record carries a screening decision and none awaits a full-text decision
Every study appraisedEvery study in the body has a determined risk-of-bias judgement
No unreconciled contradictionNo relationship carries opposing findings
No very-low-certainty findingCertainty has been rated and nothing is very low

A record with no status (pasted text, a PDF, a Find papers result) and an imported record still at Identified — not yet screened both count as carrying no screening decision. So a review built only from pasted papers stays at Sound until it is saved as a project and its records are given a status on the Screen panel.

The status badge counts only the checks that are faults in the evidence (ungrounded claims, contradictions, very-low certainty or certainty not yet rated). Screening and appraisal are listed as the Next step instead.

Tips

  • Import each database's export separately and run De-duplicate after the last one: the cross-database copies are what PRISMA's "duplicates removed" box counts.
  • Save as a project immediately after importing. Records only reach the Screen panel, and your co-reviewers, once the project holds them.
  • Agree your exclusion reasons up front and use the suggested wording, so the exclusion log groups cleanly.
  • Have each reviewer vote before looking at the consensus dots: they appear only after you vote, which keeps your screening independent.
  • When κ reads n/a because everyone excluded everything, screen a sample of records likely to be included as well.
  • Keep the saved query in the Find papers field to use Check for new papers. Clear the field to use Re-run with new sources for PDFs and pasted text. The find buttons on gaps and suggested searches overwrite the field, so check it before you press the button at the foot of the flyout.
  • Open the starter Systematic review and meta-analysis (PRISMA 2020) from ⋮ Starter reviews… to see a finished screen, an exclusion log, RoB 2 judgements and a PRISMA flow before you start your own.
  • Use Edit numbers to type a trial's arm-level counts: the pooled estimate then has an exact standard error instead of resting on a sentence.
  • Watch the amber Risk of bias cells in the Summary of findings: they show exactly which ratings are still upper bounds.

Limits and known constraints

  • One synthesis or re-run reads at most 50 sources. Larger reviews must be screened down, or read over several re-runs.
  • Each source contributes at most its first 24,000 characters; each PDF at most its first 40 pages.
  • A run has about 55 seconds. Large batches may be read in part, and are reported as partial.
  • A reference file can hold up to 8,000,000 characters and 25,000 records; one import looks up at most 200 DOIs.
  • A topic search returns up to 8 works with abstracts; a living check takes in 10 new works at a time.
  • Abstracts, not paywalled full texts, are what OpenAlex supplies. Add PDFs for full texts.
  • The studio does not let you link two reports of one trial by hand; reports are linked automatically only when they share a DOI or a trial registration number.
  • De-duplication is not undoable from Undo; reset statuses by hand.
  • Indirectness is not assessed. Only RoB 2 and ROBINS-I are offered for risk of bias.
  • Votes, screening, de-duplication, stored appraisals and re-runs need a saved project and editor rights; inviting co-reviewers needs Plus or above.
  • The Screen list shows every stored record in one list, with no search or filter.
  • While the votes on a record disagree, only the owner can change its status, even after the conflict has been resolved.
  • A check for new papers only finds again the papers inside its window (30 days before the previous complete check onwards). Older papers it stored for screening must be given their text by hand to be read.
  • The inspector's tabs, including Team, open only once the review has a graph. A project saved from imported records alone shows only the Screen panel until a first re-run builds a graph.

Troubleshooting

You seeWhat it meansWhat to do
This does not look like a RIS, BibTeX or PubMed/MEDLINE export. …The file matches none of the formatsRe-export choosing RIS, BibTeX or PubMed format. From EndNote, choose RIS
N need full text (add the PDF)No abstract exists upstream for these recordsAdd the PDFs or paste the abstracts; you can still screen them on their titles
N not looked up — DOI lookup hit its rate limit; retry to finish themLookups stopped earlyImport the same file again later; rows already filled are left as they are
None of the N sources with text has been screened in — only included studies are read. …Everything with text is still awaiting screeningSave as a project if needed, then screen records in
N screened-in sources have text, and one synthesis run takes 50. …Too many included sourcesExclude or remove sources, or synthesize in batches
You need edit access. / You need edit access to screen.You are a viewerAsk the owner to make you an editor
The reviewers disagree on this record — … The project owner resolves a conflict by setting its final status.The record is in conflict and you are not the ownerAsk the owner to set the final status
Could not record decision.Your vote did not reach the server; it has been removed from screenCheck your connection and vote again
Could not update screening.The status change failed and was revertedTry again
Could not de-duplicate.The de-duplication request failedTry again
… The final status was saved, but the conflict is still recorded as unresolved — set it again to resolve it.Part of a conflict resolution failedSet the status again
Could not save that risk-of-bias judgement.The judgement was not stored and has been revertedTry again
You need edit access to record a risk-of-bias judgement.You are a viewer on this projectAsk the owner to make you an editor
You need edit access to change a finding.A viewer tried to save numbersAsk the owner to make you an editor
That finding is not in this project — reload and try again.The claim changed on the server since you opened the projectReopen the project and edit the numbers again
The review is saved, but N of M risk-of-bias judgements could not be stored …Saving kept the review but lost some judgementsRe-record them under Risk of bias on the Screen panel
Those numbers are not a complete set — …A 2×2, means or rates table is missing a valueFill every box for the kind you chose, or choose None
Could not save those numbers — the finding is unchanged.Storing the correction failed and it was rolled backTry again
You need edit access to run synthesis.A viewer pressed a re-run buttonAsk for editor rights
None of the sources with text has been screened in — only included studies are read. Screen them in on the Screen tab first.Re-run with new sources found nothing included to readScreen records in, then re-run
Add new sources (or a saved query) to re-run.No source has text and the Find papers field is emptyAdd PDFs or text, or put the saved query back in the field
Read none of the N new sources — claim extraction failed on … so nothing was merged and no run was recorded.The AI provider failed on every new sourceCheck your AI key and model in Settings, then re-run
Re-run read none of the N new sources — the time budget stopped the run before it opened …The run ran out of time before reading anything (a notice)Re-run, or add the sources in smaller batches
N more new papers published since DATE are still waiting — this check took in the newest batch. …More new works are in the window than one check takes inPress Check for new papers again
N more sources have new text waiting — one re-run reads 50, so re-run again to read them.More than 50 sources were dueRe-run again
Source lookup failed. / OpenAlex did not answer within 8 s — try again.OpenAlex could not be reached during a checkTry again later
Nothing to re-read: …Every source is unchanged since it was last read (a notice, not an error)Add or edit a source, or run a new search
No new papers: OpenAlex has no work with an abstract matching …Nothing has been published in the windowCheck again later
… Narrow the query so the window fits; the search date has not moved.More recent works match than one check can page throughMake the query more specific
This project holds more than 200000 sources; the run refuses to guess at the rest.The project is too large for a re-run to read safelySplit the review into smaller projects

For problems with the studio itself (saving, exports, the graph), see the troubleshooting table on Evidence.

Something unclear or out of date on this page? Tell us from the Support link in any studio — the flowss team reads every report.

© 2026 Voranox Inc. flowss — Flow Systems Studio. All rights reserved.

This documentation, its text and its examples are protected by copyright. Engine and format names are trademarks of their respective owners — see the terms and copyright and licences.