Figure's analysis side turns the rows on a figure's bench into numbers you can defend: a cleaning bench that finds problems in the data and fixes them in one undoable click, a transform bench for filtering, summarising and reshaping, a statistics bench with 23 analyses behind an assumption gate that refuses the wrong test and says why, a study-design bench for sample size and power, and a ledger that records how every number was computed so a figure can say when one of its results has gone stale. Nothing on these benches is written by AI: every value is computed from your rows by deterministic, cited methods, and the Methods paragraph is composed from what you actually did. This page documents all of it. For the studio itself (bringing data in, drawing panels, the journal check and exports) see Figure.
At a glance
| Topic | What you need to know |
|---|---|
| Where | In Figure: the rail's Data button (tabs Data in, Clean, Transform) and Analyse button (tabs Tests, Design), plus the Data table view |
| Scope | Everything belongs to one figure: its rows, its data ledger, its analyses. Each figure in a manuscript keeps its own |
| Cleaning | Twelve kinds of finding, from blank rows to ambiguous decimal commas, each with a one-click fix where one is safe. Outliers are reported, never removed |
| Transforms | Filter, derive, summarise, aggregate, pivot, wide → long (stack and split), join, and "what a row is" |
| Analyses | 23, from descriptive statistics to Tukey's HSD, dose–response fits, Kaplan–Meier survival, ROC and calibration |
| Assumption gate | Seven checks. Structural failures refuse and name the test that applies; advisory ones suggest a switch but never block |
| On the figure | A result can become a statistical panel, a fitted line or curve with its band on a chart, or a computed significance bracket |
| Provenance | Every recorded result has a ledger entry (L-01, L-02, …). A result whose inputs changed is marked stale |
| Methods | A copy-ready Methods paragraph built from the dataset, the ledger and the analyses, with a family-wise correction sentence and method citations |
| Design | Solve n, power, effect or α for nine designs; power curves; equivalence and non-inferiority margins |
| AI | None on these benches. The analyst's Advisor proposes analyses deterministically; see Figure |
How the pieces fit together
- Rows arrive on the Data in tab. They are this figure's dataset.
- The Clean tab audits them and offers fixes. Every fix is recorded.
- The Transform tab and the Data table reshape and edit them. Every change is recorded.
- The Tests tab runs analyses on the rows as they now stand. Each result is recorded in the ledger with the columns it read.
- Results reach the figure as statistical panels, fitted bands and computed brackets. Bands and brackets cite their ledger entry.
- The Methods paragraph on the Export bench is written from all of the above.
The data ledger is the plain-language list of operations (steps 2 and 3). The provenance ledger is the record of computations (steps 4 and 5). Loading new rows clears the data ledger and the analysis cards, because they described the old rows; entries already drawn on panels stay in the provenance ledger and show as stale until you re-run them.
The Clean bench
Open the rail's Data button and choose the Clean tab. With no rows it reads "Bring data in first — the bench audits it automatically." The Data button and the Clean tab carry a badge with the number of findings.
When there is nothing to report the bench says "No data issues found — blanks, duplicates, mixed types and stray whitespace all clear." Otherwise each finding is a card with a sentence naming the column, the count and examples, and, where the fix is safe, a button that applies it. The button's label describes exactly what it will do. Findings are listed worst first: things that break a figure, then things that distort one, then things that only look untidy.
| Finding | Example message | Fix offered |
|---|---|---|
| Empty rows | 3 rows are completely empty — they count towards every total and contribute nothing. | Delete rows in which every cell is blank. |
| Duplicate rows | 2 rows repeat an earlier row exactly — every sum and count includes them twice. | Delete rows that repeat an earlier row exactly, keeping the first of each. |
| Blank header | A column has no name — a figure has nothing to label its values with. | Delete the column, when it is also empty |
| Mostly blank column (half or more blank) | "dose" is 62% blank (… of …) — too sparse to chart. | None |
| Blank cells | "weight" has 4 blanks — gaps break a line and skew a mean. | For a numeric column, fill the blanks with the column's median |
| Constant column | "site" holds the same value in every row — it cannot distinguish anything on a figure. | Delete the column |
| Numbers stored as text | "od" is text because 2 of 48 values are not numbers (…) | Convert to numbers (values that are not numbers become blank) |
| Ambiguous decimal commas | …their commas could be thousands separators or decimal commas — "1,234" is 1234 one way and 1.234 the other… | Two buttons, one per reading. Nothing is converted on a guess |
| Mixed date formats | "date" mixes date formats — 12 ISO (2026-03-04) and 3 slash (3/4/2026)… | Convert to ISO dates, when one reading fits every row |
| Ambiguous slash dates | "date" writes dates like "03/04/2026", which is 2026-04-03 read day first and 2026-03-04 read month first… | Two buttons: read day first, or read month first. Until you choose, the column is read month first |
| Conflicting or impossible dates | Day-first and month-first dates mixed in one column, or dates the calendar does not have (13/13/2026, 2026-02-30) | None. Correct the odd ones out |
| Stray whitespace | 1 value in "region" has surrounding spaces, splitting 1 category in two — " North" and "North" are counted separately. | Remove the spaces in that column |
| Outliers | 3 values in "rt" sit outside the 1.5×IQR fence (below … or above …): … Nothing was removed — an extreme value is often the real finding. | None, deliberately |
Outliers are checked only on numeric columns with at least 8 values, and not at all when half the column is one value. Below the findings, three bench-wide buttons are always available: Trim all text, Drop empty rows and Drop duplicate rows.
The Data table's column menu offers more cleaning on a single column: Rename…, Coerce to number, Coerce to text, Coerce to date or Coerce to boolean, Fill blanks with the mean, Fill blanks with previous value, Trim whitespace and Delete column. See Figure.
Some operations are refused rather than guessed, and the reason appears above the table:
| Refusal | Why |
|---|---|
| There is already a column called "…" — renaming onto it would merge the two into one, so "…" was left as it is. | Two columns cannot share a name |
| A column needs a name — it was not renamed. | The new name was blank |
| A table needs at least one column — this one was not deleted. | You tried to delete the last column |
| "…" was not converted: it uses commas that could be thousands separators or decimal commas (…), and a guess could change its numbers a thousandfold. Choose a reading on the Clean bench. | Coerce to number from the column menu cannot pick a decimal mark for you |
| "…" was not converted: it mixes day-first dates (…) with month-first ones (…), and no single reading fits every row. Correct the odd ones out first. | Coerce to date cannot pick a date order for you |
The Transform bench
Open Data → Transform. With no rows it reads "Transforms need data on the bench first." Choose a tool from the row of buttons: filter, derive, summarise, aggregate, pivot, wide → long, join and what a row is. Every transform is undoable and lands in the data ledger for your Methods paragraph. Buttons marked "(replaces table)" replace the rows with their result.
filter
Build one or more conditions, then choose keep matches or drop matches and press Apply filter (disabled until there is a condition).
| Part | Options |
|---|---|
| Column | Any column |
| Comparison | =, ≠, >, ≥, <, ≤, contains, doesn't contain, is blank, isn't blank |
| Value | Not needed for the two blank tests. For >, ≥, < and ≤ on a numeric column the value must be a number; on a date column it must be a recognisable date such as 2026-03-04 |
Press + ("Add condition") to add each condition; the cross on a condition removes it. A value the comparison cannot use is refused by name, for example "The condition on "dose" compares numbers, but "high" is not numeric — give it a number to compare against."
derive
Name the New column name (for example "bmi" or "log_dose"), choose a Formula and press Derive column.
| Formula | Inputs |
|---|---|
| Function of a column | A numeric Column and a Function: log10, ln, sqrt, abs, zscore, cumsum, pct-change, rank or round |
| Column ∘ column | Left, Op (+, -, *, /) and Right, both numeric |
| Column ∘ constant | Left, Op and Constant |
| Part of a date | A Date column and the Part: year, month, day or weekday |
| Bin a numeric column | A numeric Column and Bins, from 2 to 50 |
summarise
"One row per group, with the spread the error bars will draw and the n it was taken over." This is how replicate rows become the one-row-per-group table an error-bar chart needs.
- Choose Group by and the Measured column.
- Choose what Error bars show: SEM (standard error of the mean), SD (spread of the observations), 95% CI half-width or 99% CI half-width.
- Read the caption the bench previews, for example "mean ± SEM, n = 6 per group." Any caveat is shown with it, and a request it cannot summarise is refused in red.
- Press Summarise for plotting (replaces table).
The caption is added to the data ledger (so it reaches the Methods paragraph) and is remembered with the new rows: an error-bar chart you then draw from them is captioned with it.
If you have declared an experimental unit (see below), the bench says, for example, "Averaging within each mouse first — n will count mice, not rows." and the caption counts units, for example "mean ± SEM, n = 3 mice (4 technical replicates each, averaged first)." If you have not, and a column's values repeat within each group in a way that looks like a unit, an amber advisory names it beside the n it would change, with a button (for example Declare "mouse_id" as the unit) that opens the what a row is form, pre-filled.
aggregate
Tick one or more Group by columns, choose a Measure and a Function, and press Aggregate (replaces table). The functions are named in full:
| Function | Meaning |
|---|---|
| sum, mean, median, min, max | As named |
| count (rows) | Rows in the group |
| distinct values | Distinct values in the group |
| n (values) | Non-blank values |
| SD | Standard deviation |
| SEM = SD/√n | Standard error of the mean |
| 95% CI half-width, 99% CI half-width | The t-based half-width |
| Q1, Q3, IQR (Q3 − Q1) | Quartiles and their range |
| MAD (unscaled, R constant = 1) | Median absolute deviation, not scaled to estimate an SD |
| geometric mean | As named |
pivot
Choose Rows, Columns, a numeric Value and a Function (the same list), and press Pivot (replaces table).
wide → long
"A plate reader writes one row per sample and a column per replicate. Every chart needs the other shape."
| Control | What it does |
|---|---|
| A button naming the run it found, for example "Replicates are in columns — rep1, rep2, rep3. Stack them into one column." | Appears when Figure detects a run of replicate columns. One click stacks them. The same offer appears on the Data in tab when the rows arrive |
| Columns to stack, Headers become (default "replicate"), Values become (default "value"), Stack into long form | Stack any columns you choose |
| Split one column into several: Column, Separator (default "_"), New names (comma-separated, at least two, for example "well, arm, hour"), Split column | Splits a column whose values join several facts with a separator into one new column per fact |
join
"The plate gives you a sample id and a number; the arm, the dose and the sex live in a second table — another figure in this project, or one you paste here."
- Under Second table, choose another figure in this project that has rows (the picker appears only when one does), or leave it on paste one below and paste the table.
- Choose the column to Match on.
- Choose what happens to Rows with no match: drop them (inner) or keep them, blank (left).
- Press Join the second table.
what a row is: the experimental unit
Many experiments record several rows per subject: four wells per mouse, three fields per dish, repeated measures per patient. Those rows are not independent observations, and counting them as such inflates n. This form, also on the Data in tab under "Several rows per mouse, dish or patient? Say what one row is.", declares the structure. It moves no data; it decides what every n in the figure means.
| Field | Example |
|---|---|
| Which column names the unit? | mouse_id |
| Its numbers restart within… | Choose a column when units are numbered afresh in each group (mice 1 to 3 under "ctrl" and again under "drug" are six mice), or nothing — each value is one unit |
| One unit is a… (required) | mouse, dish, patient |
| One row is a… | well, field, section |
Press Declare it. The form then confirms the declaration, for example: "mouse_id" names each mouse, and each row is one well. Summaries average within a mouse first. The Data in tab's heading becomes, for example, "n counts mice, not rows". Clear removes the declaration. Loading new rows with Analyse data or a file upload clears it too, because it names columns of the old table.
Once declared, the unit is honoured by summarise, by the Group summary analysis, by the error-bar repair on the Charts bench, and by the two-group and k-group comparisons (both t-tests, Mann–Whitney, ANOVA, Kruskal–Wallis, Tukey, Games–Howell, Dunnett), which then compare one mean per unit. A comparison applies the unit only when it is a third column, not the value or the group you are comparing by.
The data ledger and undo
Every cleaning, transform and table edit adds a plain sentence to the figure's data ledger. The last eight lines are shown at the foot of the Export popover under Data ledger — and the figure's name; the whole ledger feeds the Methods paragraph and ships in the repro bundle as data-preparation.txt. While the Data table view is up, ⌘Z and the view island's Undo step back through data changes (50 steps), taking the ledger with them.
The Analyse bench
Open the rail's Analyse button; the first tab is Tests. With no rows it reads "No rows yet." When the manuscript has more than one figure, a line reminds you that analyses and the data ledger belong to the figure on screen.
Running an analysis
- Choose the Analysis. Its form appears below.
- Pick the columns and options. Pickers for numeric inputs list only numeric columns; an optional column shows "all…".
- Read Why this test, the assumption gate's numbered chain, if the analysis has gates.
- Press Run analysis. The result appears as a card below the form.
To let Figure pick the test, fill in a value and a group (or two measures, for a paired design) and press Choose the test. It walks the assumption gate for those columns and puts the chosen test's form on screen, so you see exactly what will run before you run it; or it says, by name, why no test applies, with Switch to and the test that does where there is one.
The analyses
Every card can also be added to the figure as a statistical panel with Add as panel. The last column says what else it can draw onto a chart.
| Analysis | Inputs | What the card reports | Gate | Effect size | Draws on a chart |
|---|---|---|---|---|---|
| Descriptive statistics | Column (optional — blank = all) | n, mean (SD), median [IQR], range, 95% CI for the mean by the t interval, values outside the Tukey fences, missing | (none: describes, does not compare) | ||
| Frequency table | Column | Counts and shares; the most frequent category; long tails pooled as "(other)" | (none) | ||
| Correlation matrix | Method: Pearson (linear) (default) or Spearman (rank) | The matrix over listwise-complete rows and the strongest pair; fragile under 10 rows | (none: every cell is one) | ||
| Linear regression | X (predictor), Y (response) | Equation, R², slope ± SE with its t test, Pearson r | Pearson r with Fisher interval | Fit line and band | |
| t-test (two groups) | Value column, Group column | Each group's mean, SD and n; mean difference with 95% CI; Welch's t, df and p; Cohen's d | Parametric | Hedges' g with a noncentral-t interval (noting it assumes equal variances) | |
| One-sample t-test | Value column, Compare against (μ), default 0 | Mean difference from μ with 95% CI; t, df, p | Cohen's d | ||
| Paired t-test | First measure, Second measure | Mean paired difference with 95% CI; t, df, p; dz | Paired | Cohen's d_z | |
| One-way ANOVA | Value column, Group column | F, df, p; η² | Parametric | Partial η² with Steiger's interval; ω² and Cohen's f noted | |
| Chi-square (independence) | First column, Second column | χ², df, p, Cramér's V, the cross-table; warns when expected counts fall below 5 | Odds ratio (Woolf interval) for a 2×2 table, otherwise Cramér's V | ||
| Mann–Whitney U | Value column, Group column | Each group's median and n; U, z, p (normal approximation with tie correction) | Rank | Rank-biserial correlation | |
| Normality check | Column | Skewness, excess kurtosis, Jarque–Bera and p, with a caveat when the sample is too small to say anything | (none) | ||
| Group summary (mean ± spread) | Measured column, Group by, Error bars show (default SEM) | The per-group summary and its caption, honouring a declared unit | (none: the spread is the magnitude) | ||
| Tukey's HSD (all pairs) | Value column, Group column | Every pair's difference, simultaneous interval and adjusted p; the omnibus ANOVA; Levene's verdict | Parametric | (per-pair differences) | Significance brackets |
| Dunnett (treatments vs control) | Value column, Group column, Control level (exact label), Direction | Each treatment against the control | Parametric | (per-contrast differences) | Significance brackets |
| Curve fit (non-linear least squares) | X (independent), Y (measured), Model (default Michaelis–Menten) | Parameter estimates with SE and 95% CI, R², AICc, RSS | (the parameters are the magnitudes) | Fitted curve and band | |
| Dose–response (4PL/5PL, EC50) | Dose (concentration), Response, Curve per (optional), Model, Shared across curves, Constrain plateaus | EC50 with its 95% CI per curve, Hill slope, parameter table, EC50 ratios against the first curve, shared-versus-separate AICc | (EC50 and Hill slope) | Fitted curves and bands | |
| Games–Howell (unequal variances) | Value column, Group column | Every pair, each on its own Welch df | Parametric | (per-pair differences) | Significance brackets |
| Survival — Kaplan–Meier, log-rank, hazard ratio and RMST | Follow-up time, Event (1) or censored (0), Arm (optional — blank = one cohort), Reference arm (optional — exact label; blank = the last arm in the rows) | Kaplan–Meier per arm, medians, log-rank test, Yusuf–Peto hazard ratio, restricted mean survival time difference, median follow-up by reverse Kaplan–Meier | Hazard ratio (two arms only) | ||
| ROC curve, AUC and the DeLong comparison | Score, Outcome (0/1), Second score (optional — paired comparison), The score came from | AUC with DeLong's interval, the Youden-optimal cut-off, and a paired DeLong comparison of two scores | (the AUC) | ||
| Calibration — slope, intercept, ICI and Brier | Predicted risk (0–1), Outcome (0/1), Risk groups (5, 10 or 20; default 10), The risks came from | Calibration slope and intercept, calibration-in-the-large, ICI with E50, E90 and Emax, Brier score with Murphy's decomposition, Hosmer–Lemeshow reported beside its bin count | (slope, intercept and ICI) | ||
| Student's t-test (equal variances) | Value column, Group column | Each group's mean, SD and n; mean difference with its CI; Student's t on the pooled variance; Cohen's d | Parametric | Hedges' g with a noncentral-t interval | |
| Kruskal–Wallis (rank-based, k groups) | Value column, Group column | Each group's median and n; tie-corrected H against χ², df, p; ε² | Rank | (ε² on the card) | |
| Wilcoxon signed-rank (paired) | First measure, Second measure | Median paired difference; V and p (exact where possible, otherwise the normal approximation); zero differences are discarded and counted | Paired rank | Matched-pairs rank-biserial correlation |
Notes on specific analyses:
| Analysis | Detail |
|---|---|
| Curve fit models | Straight line; Exponential growth or decay; Zero-, first- and second-order integrated rate laws; Michaelis–Menten; Hill equation; Logistic growth; Arrhenius; Four-parameter logistic (dose–response); Five-parameter logistic (asymmetric dose–response). Each is listed with its formula |
| Dose–response | Model is the four-parameter (default) or five-parameter logistic. Shared across curves ticks the parameters a global fit shares: HillSlope (parallel curves), Bottom, Top, LogEC50 (4PL), LogC (5PL), S (5PL); ticking one the model does not have is refused by name. Constrain plateaus: Nothing — fit both plateaus (default), Bottom = 0, or Bottom = 0, Top = 100 (normalised). Rows at a dose of 0 enter the fit as the low-dose plateau. EC50 intervals are formed on log₁₀ dose. An EC50 ratio says whether the curves are parallel, which decides whether the ratio holds at every response level |
| 4PL and 5PL preview | Choosing Dose–response, or either logistic model under Curve fit, shows "4PL/5PL preview: refusal diagnostics are still being validated across runtimes. Independently verify scientific interpretations." The same sentence rides on the card |
| Post-hoc families | Tukey, Dunnett and Games–Howell accept between 2 and 12 groups. Dunnett needs the exact control label (a label that is not in the column is refused with the list of labels) and a Direction: Two-sided (default), One-sided (treatment above control) or One-sided (treatment below control). Their adjusted p-values already control the family-wise error rate, so they are not corrected again in the Methods paragraph |
| Survival | The hazard ratio is the Yusuf–Peto one-step estimator, labelled as such: it is not a Cox model and undershoots a large effect. It is given only when there are exactly two arms; with more, the log-rank test and per-arm medians stand. Without a reference arm, the last arm in the rows is the reference, and the card says so. RMST is restricted to the shorter arm's last observation |
| ROC and calibration | The score came from (or The risks came from): A model fitted on these same rows (apparent) (the default), Internal validation, Bootstrap validation, Cross-validation (out-of-fold predictions), A held-out split of these rows, An external cohort, A later time period (temporal). Apparent, internal and bootstrap validation carry a caveat that the result is optimistic |
| Calibration | The Hosmer–Lemeshow p is reported but is not the card's result and is not counted in the Methods paragraph's family of tests. A predicted risk of exactly 0 or 1 is refused with a suggestion to clip it away from the boundary |
Why this test: the assumption gate
Every comparison declares which assumptions it rests on. The gate checks them on exactly the request you are about to run, and shows its reasoning as a numbered chain under Why this test, each line with its own evidence (for example, that Levene's test found the variances differ, so Welch rather than Student).
| Gate | Kind | What it asks |
|---|---|---|
| What was measured | Structural | Is the outcome a measured number, not a category? |
| How many groups | Structural | Does the test fit the number of levels (two for a t-test, more for ANOVA)? |
| Independent or paired | Structural | Are the groups independent, or is one subject measured under both? |
| What n counts | Structural | Are repeated rows of one mouse, dish or patient being counted as independent? |
| How many in each group | Structural | Are the groups big enough for the test to answer? |
| Shape | Advisory | Is the distribution in each group plausibly normal? |
| Spread | Advisory | Are the group variances equal? |
| Gate set | Used by |
|---|---|
| Parametric (all seven) | t-test (two groups), Student's t-test, One-way ANOVA, Tukey's HSD, Dunnett, Games–Howell |
| Rank (the five structural gates) | Mann–Whitney U, Kruskal–Wallis |
| Paired (what was measured, independence, group size, shape) | Paired t-test |
| Paired rank (what was measured, independence, group size) | Wilcoxon signed-rank |
The verdict decides what happens to the Run analysis button:
| Verdict | What you see |
|---|---|
| Clear | The chain, and Run analysis as normal |
| Route (an advisory gate points elsewhere) | The chain's line in amber and a Switch to button naming the test, beside a live Run analysis. Advice never blocks |
| Refuse (a structural gate fails) | The reason in red in place of Run analysis and Resample this, with a Switch to button where another test applies. Where the refusal is a paired-looking design that could instead be units numbered afresh in each group, a Declare button makes that declaration in one click (it names the column, the unit and the group it is numbered within) |
How the gate decides, in brief:
| Check | Method |
|---|---|
| Normality | D'Agostino–Pearson K² in each group when every group has at least 20 values; D'Agostino's skewness test alone from 8; not tested below 8, and the chain says the test has no power there |
| Equal variance | Levene's test centred on the median (the Brown–Forsythe form). Student's pooled t is chosen only when it does not reject and the groups are the same size |
| Rank tests | Proposed only when they can reach significance at your group sizes; at 3 + 3 no arrangement takes Mann–Whitney below p = 0.1, so it is not proposed there |
| Units | A nested text column, or a nested number that numbers things (by its name, such as "mouse" or "animal_id", or by values 1 to N), is treated as a unit and refuses until declared. Other nested numbers (a dose, a day) are named with a caveat that travels to the card and the Methods sentence |
| Pretesting | When a pretest decided the route, the chain notes that choosing a test by pretesting the same data affects its error rate, and that a pre-registered test is the one to run |
The same refusal also applies when the Advisor or the flowss Figure Agent asks for a test, so a structurally wrong test cannot run by any route. If no more specific reason is available you see "The assumption gate refused this test."
Long-running fits
Curve fits and dose–response fits always run in the background, and so do survival and ROC analyses over 20,000 rows and calibration over 10,000. While one runs, a line reads "Running in the background" with its progress and a Cancel button, Run analysis is disabled, and the canvas stays responsive. An analysis started another way (for example from the Advisor) cancels a fit in progress. A fit that finishes after you switched figure is not applied.
Reading a result card
| Part | What it is |
|---|---|
| Title | For example "Welch t-test — weight by arm" |
| Findings | Up to four lines of results in sentences |
| Comparison table | On Tukey, Dunnett and Games–Howell cards: Pair, Diff., the CI at the family's level, p adj., and a Bracket button per row. Pairs that differ are in bold |
| Parameter table | On curve-fit and dose–response cards: each parameter's estimate, SE and 95% CI, one block per curve |
| Effect line | The effect size and its interval where one exists, with a magnitude word (small, medium, large) where conventional; or "No effect size —" followed by the reason |
| Caveat | Amber italic text about what the result cannot say |
| Trace line | Shown only when something is wrong: "stale ·", "orphan ·" or "untraced ·" followed by the reason (stale means a column it read has changed). A current result shows nothing |
| Hover | Hovering the card shows its provenance: the ledger entry, method and version, the dataset fingerprint and the columns it read |
| Add as panel | Adds the card as a statistical panel |
| Copy text | Copies the title, every finding line, the effect line and the caveat. It reads Copied for a moment |
| How this was calculated | Opens the calculation record |
| Show on panel … | On regression and fit cards: draws the result on a chart |
| The bin icon | Removes the card. If brackets were drawn from it, the first click lists them ("Removing this analysis also removes the bracket drawn from it: …") with Remove both and Keep |
Running the same analysis on unchanged data again reuses the same ledger entry. A card the ledger could not record says "The result stands, but it was not recorded:" and the reason, so an unrecorded number is never mistaken for a recorded one.
Effect sizes and the effect panel
Every inferential result carries a magnitude, with its interval where one can be computed honestly, or a named reason there is none, because a p-value alone does not say how big an effect is.
| Measure | From |
|---|---|
| Hedges' g (with Cohen's d) | Two-group t-tests. The interval is the exact noncentral-t interval |
| Cohen's d (with its large-sample interval) | One-sample and paired t-tests (d_z for paired) |
| Pearson r | Regression, with Fisher's z interval and R² noted |
| Partial η² | One-way ANOVA, with Steiger's interval; ω² and Cohen's f noted |
| Odds ratio | Chi-square on a 2×2 table, with Woolf's interval |
| Cramér's V | Chi-square on larger tables (no interval; the card explains why) |
| Rank-biserial correlation | Mann–Whitney U and (matched-pairs) Wilcoxon signed-rank (no closed-form interval; the card suggests bootstrapping) |
| Hazard ratio | Survival with exactly two arms (Yusuf–Peto) |
Once any card carries an effect size, Effect panel — every result's magnitude opens Effect sizes, one row per result: a forest-style chart with one row per result, a marker at the estimate and a whisker for its interval, drawn in the accent colour when the interval excludes the null value. Measures that do not share a scale get their own strip, each labelled with its null: Standardised mean differences and Correlations (null at 0), Shares of variance and association strength (null at 0) and Ratios (log axis) (null at 1). Results are not pooled: they are different analyses, and a pooled estimate would be a synthesis that never happened. Results without an effect size are listed underneath with the reason. The effect panel is a reading aid on the bench, not an exportable figure. Hide the effect panel closes it.
Resample this
Resample this sits beside Run analysis and works on the request in the form, without adding a card.
| Analysis | What resampling gives |
|---|---|
| t-test (two groups), Student's t-test | The exact permutation test on the card's own statistic, plus a stratified bootstrap interval of the mean difference (each resample draws from each arm separately) |
| Mann–Whitney U | The exact permutation test |
| Descriptive statistics (with one column chosen) | A bootstrap interval for the mean |
| One-sample t-test | A bootstrap interval for the mean difference from μ |
| Paired t-test | A bootstrap interval for the mean paired difference, resampling pairs |
| Everything else | A named refusal explaining why resampling does not apply (for example, a correlation matrix has no single quantity to resample) |
A bootstrap readout shows the card's parametric interval, the percentile interval and the BCa interval side by side, with the resample count and seed ("2000 resamples, seed …") and the words "derived from these rows, not from a clock." The seed comes from the request and the values it reads, so the same request on the same rows always gives the same interval, and changing one cell changes it. A note appears only when the BCa interval is materially asymmetric and a skewness test rejects symmetry. A permutation readout shows "Permutation p =" and the value, marked "(exact)" when every arrangement was enumerated, or "(sampled)" when there were too many.
Resampling is capped at 2,000 resamples over at most 5,000 values; a larger request is refused by name with the cap, never quietly reduced. A structural refusal from the assumption gate applies here too, and hides the button.
Putting results on the figure
Statistical panels
Add as panel draws the card as a table-style panel across the full width of the figure, with its type drawn in points at the size it will print and never below the journal's minimum. Long tables show up to 30 rows and note how many more there are ("… 12 more rows"); columns that do not fit are dropped and counted ("… 3 more columns — too wide for the panel"). A statistical panel has no engine source, so it cannot be restyled or refreshed; its provenance is the analysis, which travels in the Methods text.
Fit lines and bands
A Linear regression, Curve fit or Dose–response card can draw its result onto a chart panel that plots the same two columns.
- On the card, find Show on panel ….
- Press the toggle to choose the band: confidence (the mean response) or prediction (a new observation).
- Press a panel chip (labelled Panel and the letter, for example "Panel A").
The chart is regenerated from its own data link with the fitted line or curves and the band layered on, verified against the ledger entry, and only then drawn. A chip that cannot take the result is greyed with the reason in its tooltip, for example "Panel B was pasted rather than generated, so it has no chart source a fit could be drawn onto." or a sentence saying the fit was computed before your last edit and should be re-run. When no chip can take it, the reason is printed under the chips. A regression whose points all sit on the line (or whose y never varies) records why it has no band, and every chip says so.
The two figure starters Scatter plot with regression line and 95% CI and Dose–response curve (4PL, IC50) open with their fits drawn exactly this way.
Significance brackets
Brackets over pairs of groups should come from a test, not from typing. The computed path:
- Run Tukey's HSD (all pairs), Dunnett (treatments vs control) or Games–Howell (unequal variances) on the value and group a chart panel plots.
- On the card, choose the panel under Brackets on and the notation: stars, exact p or both.
- Press Bracket on any row of the comparison table.
The bracket prints that test's adjusted p and cites the card's ledger entry. Stars follow the APA key: * below .05, ** below .01, *** below .001, and "ns" for a pair the family did not find, judged against the family's own α. Brackets are drawn into the chart so they move with it, and stacked so they never cross. The brackets a card has drawn are listed on it, each with a remove button, and on the panel's card in Layers. Re-running the comparison on changed data moves its brackets to the new result and redraws them.
The Check report's Significance marks section reads every mark back: computed brackets with their test and entry, typed marks from the Annotate bench as "typed, not computed", and in amber any bracket that is stale, whose entry has gone, or whose drawing no longer matches its record. Redraw brackets repairs a drawing that has drifted from its record; a stale bracket is repaired by running the comparison again.
Provenance: the ledger
Every analysis that produces parameters is recorded in the figure's provenance ledger as an entry with an id (L-01, L-02, …), the method and its version, the columns it read (each with a content fingerprint), the parameters it produced and its working. Drawing a band or a bracket records an entry of its own that cites the one it came from.
| Status | Where you see it | Meaning |
|---|---|---|
| current | Nothing (the absence is the signal) | The inputs are unchanged |
| stale | Amber on the card, amber chip on the panel | A column the result read has changed since. Re-run the analysis |
| orphan | Amber | Every input it read has left the bench, or the entry it cites is no longer in the ledger, so it cannot be checked |
| untraced | Grey | Nothing recorded how this was made (a pasted SVG, a statistical panel, or a result the ledger did not record) |
The ledger holds up to 500 entries. Removing an analysis card prunes the entries nothing else refers to; an entry a surviving band or bracket depends on is kept.
How this was calculated
The button on each card opens a dialog titled How this was calculated with the full record of that run: Result, Uncertainty, Inputs, Exclusions, Method and version, Parameters, Transformations, Assumptions, Units, Randomness, Diagnostics and Limitations, headed by the method, the ledger entry and a fingerprint. Download reproducible JSON saves the record. The record describes the rows the run was launched on; while the figure still holds exactly those columns, their values are included.
The Methods paragraph
Press Methods paragraph on the Export popover to copy it. It is available once the figure has an analysis or a line in its data ledger.
| Part | Content |
|---|---|
| The dataset | The dataset ("…") comprised N records across M variables. |
| Data preparation | "Data preparation:" followed by the data ledger, in order |
| Methods | Each analysis's own Methods sentence, once each, including the assumption checks the gate ran |
| Significance | Nothing if no test was run. For one test: "Two-sided p-values below 0.05 were considered statistically significant." For several: how many tests were run, that p-values are reported unadjusted, and how many remain significant after a Holm–Bonferroni correction across them at a family-wise error rate of 0.05 |
| Citations | "Methods cited:" then each method used, with its id, version and reference |
A sample size you solved on the Design tab and recorded appears among the data-preparation steps, prefixed "Sample size:". When the figure's rows are a sample dataset, the paragraph opens with: Example data — not real: these values are an example that ships with the platform, not data from a study. Study this figure sends the paragraph to Study with the rest of the figure.
Note: Which tests form a family is your judgement. The paragraph reports the Holm-adjusted count but never rewrites your p-values. Post-hoc procedures and the calibration card's Hosmer–Lemeshow p are not counted in the family.
The Design bench
Open Analyse → Design. "Plan the study. Nothing here reads your rows — a sample size is a fact about the design, so this bench works on an empty sheet, which is when you need it." It has four panes: Solve, Power curve, Margins and Observed power.
Designs
| Design | Effect means | n counts | Extra input |
|---|---|---|---|
| Two independent means | Cohen's d | subjects per group | |
| Paired means | d_z of the within-pair difference | pairs | |
| One mean against a constant | d against the reference mean | subjects | |
| Two proportions | Risk difference from the baseline | subjects per group | Baseline (p₁) |
| One proportion against a constant | Difference from the null proportion | subjects | Baseline (p₀) |
| Correlation | r | subjects | Baseline (r₀) |
| One-way ANOVA | Cohen's f | subjects per group | Groups (k) |
| Chi-square goodness of fit | Cohen's w | observations in total | Degrees of freedom |
| Survival (log-rank) | The magnitude of the log hazard ratio, abs(ln HR) | EVENTS, not subjects |
The line under the design picker restates what the effect and n mean for it.
Solve
- Choose the Design and what to Solve for: Sample size (n), Power, Detectable effect or Significance level (α).
- Fill in the other three (n, Power as a fraction such as 0.8, Effect, α such as 0.05), the Sides (two-sided or one-sided), the Allocation n₂/n₁, and any extra input the design needs. The form starts at n 64, power 0.8, α 0.05, allocation 1, baseline 0.4, three groups and three degrees of freedom.
- Optionally, under Cluster randomised (optional), enter the ICC (ρ) (between 0 and 1) and Average cluster size (m). The design effect, 1 + (m − 1)ρ, is shown as you type. It applies only to the two-arm designs of subjects; the others say why not.
- Press Solve. The answer appears with any notes.
- Press Record this in the Methods log to add the sentence to the data ledger. It reads Recorded in the Methods log afterwards. Solving several times while exploring adds nothing until you record.
Power curve
Choose Vary: power against n or power against effect, set From and To (10 and 200 to begin with), the fixed value, α and Mark power at (default 0.8), and press Draw the curve. The crossing of the marked power is solved exactly rather than read off the plotted points, for example "Crosses power 0.8 at n 63.766 — solved by the same bisection the sample size uses, not read off the … plotted points.", or "The curve does not cross the marked power inside this range."
Margins
For equivalence and non-inferiority trials. Choose the Margin design: Equivalence (TOST) — two means (default), Non-inferiority — two means, Equivalence (TOST) — two proportions or Non-inferiority — two proportions. Enter Margin (positive) (default 5), Power, α (one-sided) (0.05 gives the 90% two-sided interval bioequivalence asks for), the SD (default 10) or the Control proportion (default 0.7), and for non-inferiority the Direction (higher is better or lower is better). Press Size the trial, and Record this in the Methods log if you want it. For equivalence, the notes explain that both one-sided tests must reject, so β is split, and that sizing it as a non-inferiority trial would understate n by about a quarter at 80% power.
Observed power
Compute observed power explains in place why post-hoc power computed from the effect a study happened to observe is not an answer, rather than producing a misleading number.
Refusals on this bench appear as a red box with the refusal's name and the reason, never as a blank result.
Living panels and new data
A panel drawn from your rows remembers which rows, which recommendation and which cleaning steps produced it. When a new export of the same data arrives, Refresh data re-applies the recorded cleaning to it, checks that every column the chart reads survives, and redraws the panel with your styling. Table edits that cannot be re-applied to new rows (a filter, a cell edit, a reshape) block the refresh by name, and you make the same edits to the new rows on the bench instead; a recorded sort is re-applied. A refresh replaces the figure's rows, so the analyses and cleaning steps that described the old rows are cleared, with a note saying which. The full flow is in Figure.
Tips
- Declare the experimental unit as soon as the rows arrive if you have technical replicates. Every summary, comparison and caption will then count the right thing.
- Use Choose the test rather than picking a test from habit; the chain it shows is also what the Methods paragraph will say was checked.
- Use Summarise for plotting to make error-bar charts. The caption it writes (for example "mean ± SEM, n = 3 mice …") follows the chart into the panel caption and the Methods paragraph.
- Prefer computed brackets from Tukey, Dunnett or Games–Howell over typed significance marks.
- Glance at the effect panel before writing results: it shows at once which intervals exclude no effect.
- After editing the data, look for stale on cards and panels and re-run what it names before exporting.
- Record your sample-size calculation from the Design tab so it reaches the Methods paragraph.
Limits and known constraints
| Limit | Value or behaviour |
|---|---|
| Analyses saved per figure | The 40 most recent are kept when the bench is saved |
| Data ledger | The 200 most recent lines are kept when the bench is saved |
| Provenance ledger | 500 entries per figure |
| Statistical panel tables | 30 rows drawn |
| Post-hoc families | 2 to 12 groups |
| Resampling | 2,000 resamples over at most 5,000 values |
| Background runs | Curve and dose–response fits always; survival and ROC above 20,000 rows; calibration above 10,000 |
| Bins (derive) | 2 to 50 |
| 4PL and 5PL | Preview: verify interpretations independently |
| Survival hazard ratio | Yusuf–Peto one-step estimator for exactly two arms, not a Cox model; no Cox regression is available |
| Effect panel | A reading aid on the bench; it cannot be exported as a figure |
| Derived tables | A result's table can be drawn as a panel but not sent back to the data table; use the Transform bench to reshape the rows themselves |
| Multiplicity | The Methods paragraph reports a Holm correction across all tests on the figure; it does not let you choose families |
Troubleshooting
| You see | Why | What to do |
|---|---|---|
| The gate's reason in red where Run analysis should be | A structural assumption fails (wrong number of groups, a category as the outcome, undeclared repeated measures, a paired design) | Read the chain; press Switch to or Declare, or declare the unit under what a row is |
| "The assumption gate refused this test." | The gate refused without a more specific reason | Check the column types and group counts |
| "Column "group" has 14 distinct values — a post-hoc comparison needs between 2 and 12 categories…" | Too many groups for a post-hoc family | Combine or filter the groups |
| "No group named "…" in column "…" — Dunnett's test compares every treatment with one control…" | The control label does not match | Type the control exactly as it appears in the column |
| "The condition on … compares numbers, but … is not numeric…" or "… is not a recognisable date — use ISO like 2026-03-04." | A filter value the comparison cannot use | Enter a number, or an ISO date |
| "The result stands, but it was not recorded: …" | The ledger refused the entry (for example, it is full) | Remove analyses you no longer need, then run it again |
| "That result has no recorded fit, or that panel has no data link — re-run the analysis and rebuild the chart." | The fit or the chart lost what links them | Re-run the analysis and redraw the chart from the Charts bench |
| A panel chip is greyed under Show on panel … | That panel cannot take this fit (pasted, no data link, different columns, grouped differently) or the fit is stale | Hover the chip for the reason |
| "The fitted panel was refused: …" | The band did not verify against the ledger entry | Re-run the analysis, then try again |
| "The band is drawn, but it was not recorded: …" | The drawing stands but the ledger could not record it | Free ledger space by removing unneeded analyses |
| A Resample this readout names a refusal | The analysis has no resampling rule, the request exceeds the cap, or the gate refuses | Read the reason; most suggest the analysis that can be resampled |
| "There are no rows on this sheet to resample." | The figure has no data | Bring data in first |
| "Clipboard blocked — copy the analysis from the panel instead." | The browser refused clipboard access | Allow clipboard access, or add the card as a panel |
| Methods paragraph is disabled | No analyses and no data-ledger lines yet | Run an analysis or record a step |
| "Cleared with the rows they described: …" after a refresh | Refreshing replaced the rows | Re-run the analyses you need on the new data |
| "Reloaded from the pasted text — earlier table edits, the ledger and past analyses were reset." | Analyse data replaced the rows | Re-run the analyses you need on the new rows. Undo cannot bring back rows replaced this way, so keep your source file |
| A Design refusal in red | The inputs cannot define the design (for example, a power outside 0 to 1) | Read the named reason and adjust the inputs |
