Skip to content

Guides & reference

The Science Studio

Verified science rather than plausible science: what the chemistry, physics and biology checkers refuse to draw, and why.
Sculptural study of connected forms and structured ideas
Status: shipped. This document describes what /science is, the deterministic layer underneath it, and the reasoning behind the design decisions that were not obvious.

The problem this exists to solve

Every diagramming tool on the market — including, until this work, this one — treats a scientific figure as a picture. You drag a hexagon, you type "benzene" under it, you export a PNG. The tool has no idea whether the arrow you drew points the right way, whether the reaction on it balances, whether the equation in the corner is dimensionally possible, or whether the number you copied for Avogadro's constant is off by four orders of magnitude.

That is a strange gap, because most of those questions have exact, mechanical answers. Whether CH4 + O2 → CO2 + H2O balances is not a matter of opinion or of asking a language model nicely; it is the null space of a small integer matrix. Whether x = A sin(t) is meaningful is not a judgement call; the argument of a transcendental function must be dimensionless, and t is a time.

So the Science Studio's premise is: anything that can be checked mechanically, is — on every keystroke, with the working shown, and with the correction offered where one exists.

Two halves

The canvas (components/biology/BiologyComposer.tsx)

An illustrated node/edge canvas across four domains — Biology, Chemistry, Physics, Mathematics — with a curated palette per domain, hand-drawn composite figures, live collaboration, comment pins, ELK auto-layout and export to PNG/SVG/PDF/JSON.

Three kinds of node can sit on it:

Nodedata fieldWhat it is
PictogramiconIdAn Iconify glyph, or a hand-drawn bioShape composite
MoleculesmilesThe real 2-D structure, depicted from SMILES as inline SVG
EquationlatexThe real typeset equation, rendered by KaTeX

The last two matter more than they look. A node carrying smiles: "CC(=O)Oc1ccccc1C(=O)O" is not a picture of aspirin; it is aspirin, in a form a computer can read, check and re-render at any resolution. Same for the LaTeX. This is the difference between a figure a person can read and a figure a machine can verify.

The intelligence layer (lib/science/)

Pure, deterministic, dependency-free modules. No React, no network, no DOM, no model. They run identically in the browser, on the server, in the agent loop and in tests, and the same input always gives the same answer.

ModuleWhat it answers
units.tsWhat kind of thing is this? Dimensions over the seven SI base quantities, a recursive-descent parser for real unit expressions, 150+ units, conversion including the affine temperature scales
quantity.tsHow well is it known? The ± and GUM concise forms, first-order uncertainty propagation, GUM 7.2.6 rounding, ISO 13528 agreement between two measurements
elements.tsThe periodic table — all 118, with weights, configurations, electronegativities, oxidation states and a render layout
constants.tsCODATA 2022: 114 constants with uncertainties, and checkStatedValue to grade a number an author typed
chemistry.tsFormula parsing (nested groups, hydrates, charges) and balance checking
balance.tsExact balancing over the rationals, oxidation states with the rule that fixed each one, redox analysis, ion-electron half-reactions
stoichiometry.tsGrams: molar mass, limiting reagent, theoretical yield, solution prep, dilution, atom economy
solution.tspH (solved properly, not approximated), buffers, Ksp, gases, colligative properties, ΔG with its crossover temperature, Nernst
physics.tsDimensional analysis of a written expression, with ambiguous-symbol readings and the transcendental-argument rule
biostats.tsDistributions, tests, effect sizes, power and sample size, multiplicity correction, statcheck, GRIM
genetics.tsCodon tables, translation, Tm, restriction digests, Hardy–Weinberg, Punnett, HGVS, gene-symbol conventions per organism
identifiers.tsDOI, ORCID, CAS, PDB, UniProt, InChIKey, accessions — validated by their real check digits
notation.tsSI style (BIPM/NIST) and scientific typography — subscripts, isotopes, Greek, and what should be italic
pathways.ts17 canonical pathways as ordered enzyme-annotated data, and a checker that reads a drawn figure against them
taxonomy.tsScientific names under the ICZN, ICN, ICNP and ICTV — parsing, capitalisation, italics, authorities, ranks, and a curated model-organism table with NCBI taxids and lineages
experiment.ts18 study designs, control types and blinding levels; seeded randomisation (simple, block, stratified, minimisation); factorial and fractional-factorial run tables; pseudoreplication detection with the cluster design effect; and a design critique with a remedy and a citation per finding
spectroscopy.tsIR bands, ¹H/¹³C shifts, MS fragments and isotope clusters; degree of unsaturation; peak → candidate identification; cross-check of observed features against a formula; Beer–Lambert; ppm↔Hz
safety.tsGHS pictograms, hazard classes and H/EUH/P statements; laboratory chemicals with CAS, classification and handling notes; the incompatibility pairs and standing hazards that recur in incident reports; waste streams; and PPE that puts the engineering control above the glove
reproducibility.ts12 reporting guidelines (CONSORT, PRISMA, STROBE, ARRIVE, MIQE, STARD, TRIPOD…) with the items a figure can discharge, risk-of-bias tools per design, and guideline recommendation that declines rather than guesses
figure-standards.ts20 journal/format specs, and the arithmetic that turns canvas pixels into printed points
accessibility.tsCIEDE2000, colour-vision simulation, palette and greyscale checking, WCAG contrast, alt text
verify.tsThe orchestrator: which checker has anything to say about this label, and what is the fix

Design decisions worth defending

Silence is a valid answer. Every checker may return nothing, and most often does. A panel that reports on "Mitochondrion" is a panel people switch off, and a switched-off checker catches nothing. The physics analyser therefore has an explicit "is this even physics?" gate: a symbol counts as evidence only when every meaning it can carry is dimensional, so y = 2x + 3 and f(x) = x² + 3x get no verdict at all rather than a wrong one.

Symbols are ambiguous, and pretending otherwise keeps the checker tiny. T is a period, a temperature and a tension. k is a spring constant, the Boltzmann constant, a wavenumber, a rate constant and a thermal conductivity. Each symbol carries a list of candidate meanings; the checker enumerates readings, passes if any is consistent, and names the reading it used — which is what a physicist does when reading someone else's equation.

A finding carries its fix. "Not balanced" is a complaint. "Not balanced — CH₄ + 2O₂ → CO₂ + 2H₂O" is a tool. Wherever a checker can compute the corrected text it goes in fix, and the panel offers it as a one-click repair routed through the undo stack.

The working is shown. Every finding can expand to its derivation. A number a reader cannot check is a number they have to trust, and trust is not what a verification tool should be asking for.

Refuse rather than guess. Balancing reports three outcomes: solved, impossible (naming the element that appears on only one side), and underdetermined — where the species admit more than one independent reaction and returning a single set of coefficients would be a lie. Likewise oxidation states report a fractional result as an average rather than pretending Fe₃O₄ has an integer state, and molarity ↔ molality refuses to convert without a density because nothing relates them without one.

When measurement contradicts the intended design, the module says so. accessibility.ts was built expecting simulated ΔE to separate known-good palettes from known-bad ones. It does not: the worst pair in Okabe–Ito scores 10.7 under deuteranopia, while matplotlib's default red/green scores 13.9 — better. What makes Okabe–Ito safe is that its colours were derived along dichromatic confusion lines, which a pairwise distance cannot recover. Rather than tune the threshold until it produced the expected verdict, the module reports three separate quantities, states that passing is a floor rather than a certificate, and points at provenance for a real guarantee. A checker that is calibrated to nothing is worse than no checker.

An adversarial audit is worth more than another feature. After the kernel was complete and green, six agents were pointed at it with one instruction: find where the code contradicts its own docstring. Thirty-four claims came back; eighteen reproduced and are fixed, and the rest were refuted. Almost every real one was the same shape — a rule stated in prose that the code did not apply. H_CODE_TAGS["constructor"] returned Object.prototype.constructor, so ppeFor(["constructor"]) threw on a module whose header promises it never throws. identifyIR(peaks, null) threw, because a default parameter fills in for undefined and not for null. The formula index resolved C2H6OS to whichever isomer was declared first — dimethyl sulfoxide, unclassified — rather than 2-mercaptoethanol, which is fatal in contact with skin, beside a comment saying ambiguity is resolved by staying quiet. classifyOutcome matched the hint rated inside integ-rated and called a machine-measured outcome subjective, with a reason. GHS Division 1.6 carried "may mass explode in a fire" directly above a note saying no mass-explosion hazard is assigned. An incompatibility labelled "any acid or water" selected only acids, so a bench with calcium carbide and a wash bottle reported clean.

critiqueDesign decided whether the outcome assessor was blinded by scanning the blinded list for the word "assessor" — and BLINDING.single reads "participant (usually; occasionally the assessor instead)", a parenthetical saying the opposite — so a single-blind trial with a subjective primary outcome, the textbook detection-bias case, produced no finding at all. checkSpectrumAgainstFormula reported benzene's own 78 → 51 as a contradiction, because loss 27 listed only HCN: a gap in the reference table became a confident accusation.

None of those were found by the test suites, which were written by the same authors as the code and therefore shared its assumptions. All eighteen are now regression tests. Two of the audit's claims turned out to be the auditor misreading correct behaviour, and those are recorded too — the module is right that double-blind does not blind the outcome assessor unless separately stated, which is precisely why assessor-blinded is a separate level.

The general lesson, which is why this paragraph is in the document: a test suite written alongside its module inherits the module's assumptions, and the defects that survive are exactly the ones both share. An adversarial pass with a different instruction — find where the prose and the code disagree — found eighteen that ten thousand tests did not.

A safety annotation that cries wolf gets switched off. safety.ts matches chemical names against whole labels only — never by substring, never fuzzily. The cost is real: "beaker containing sodium hydroxide" is missed. The benefit is that in a studio which also draws neurons, "sodium channel" never raises an alkali-metal warning, and "sodium cyanate" never returns the hazard sheet for sodium cyanide. Nothing here ever returns "safe" either: the strongest thing it says is "nothing in this table", and every result carries DISCLAIMER — an exported string, so any surface built on the module can put it in front of a user rather than paraphrasing it.

A cluster is read as a cluster. On the M+2 peak alone, three chlorines (100 : 96 : 31 : 3) and one bromine (100 : 97) are the same number — the spectroscopy table's own note on Cl₃ says as much, and says the discriminator is M+4. So the matcher checks it, in both directions: a pattern is refuted by an M+4 it requires and the spectrum lacks, and equally by an M+4 the spectrum has and it cannot produce. Reading a bromobenzene spectrum and being told "consistent with 3 × Cl" leaves a chemist worse off than being told nothing.

The parser is not the gate, and neither is the markup. parseScientificName reads "Cell membrane" as the genus Cell and the epithet membrane — and it is right to, because that is the shape of a binomial. So the nomenclature checker only speaks when the label carries a signal prose does not: it resolves to a curated organism, or the author italicised it. Markers — an abbreviated genus, an author citation, subsp., Candidatus, a hybrid sign — are honoured only on a label already written like a name. That last qualification came from running the studio's own templates through the studio's own checks: the probability template's edge label "× critical t" parsed as a nothotaxon, and the panel advised capitalising "critical". Panel captions do the same thing to abbreviated genera ("A. Overview" is an initial and an epithet to a parser). The cost is that a bare "Quercus robur" is never checked; the benefit is that no cell biologist is ever told to italicise their organelles. Two further false-positive classes came out of the audit and are now tests: an italicised Latin phrase (*in vitro* was told to capitalise "in", because Latin phrases are italicised precisely because that is correct typography), and a bare common name (Mouse resolves to Mus musculus — true, and a useless row, since a common English word has no nomenclature to get right). The fix for the first is structural: the markup has to wrap the whole name, because nobody italicising a binomial italicises half of it. A lexical test was tried first — require the epithet to end in a Latin termination — and abandoned, because it rejects Quercus robur along with Homo sapiens and Zea mays.

Whole-figure context changes what a checker can assert. "Write the genus out before abbreviating it" is advisory in a manuscript — the preceding paragraphs decide it. A figure has to stand alone and verifyFigure holds all of it, so the studio pre-scans every label for a spelled-out genus and reports "E. coli" only when nothing anywhere in the figure writes Escherichia out. That is the same rule promoted from advice to a finding, purely because the evidence needed to settle it is in hand.

Approximations are named, and their error is reported. Weak-acid pH solves the quadratic and tells you what the "x is small" shortcut would have given. A very dilute strong acid is solved from the charge balance, so 1e-8 M HCl comes back just under 7 rather than as the notorious "pH 8 acid". pKw is a function of temperature, so neutral pH at 37 °C is 6.81, not 7.

Where the checks surface

  1. The Science Intelligence panel (ScienceIntelligencePanel.tsx) — nine tabs: Verify, Chem, Physics, Stats, Design, Bio, Publish, Table, Model. Runs on every edit; the Model tab runs on demand (see "Models" below).
  2. The agents — check_science and critique_design are agent tools. The Science Studio specialist, the multi-engine science specialist, the Lab methodologist and the Methods & statistics review specialist all have both, and all are instructed to run the checker before finishing and never to finish on a failing check. critique_design takes its description from the model rather than the document — a bar chart cannot tell you whether the allocation was concealed — so the prompt is explicit that only fields the user actually stated may be passed, and the tool reports the rest as unassessed.
  3. The public API — POST /api/v1/science exposes every operation (verify, balance, redox, oxidation-states, molar-mass, convert, dimension, check-equation, quantity, compare, constant, identifier, notation, pathway, taxonomy, critique-design, randomise, factorial, replication, spectrum, beer-lambert, nmr-frequency, reporting-guideline, risk-of-bias, hazards, figure-standards, palette, alt-text, accessibility) to anything that can make an HTTP request. GET returns the operation catalogue.
  4. The test suite — the studio's own shipped templates are run through verifyFigure, and any unbalanced reaction, wrong constant, dangling edge or unlabelled reaction arrow fails the build. A tool that ships a wrong figure as an example has no business checking anyone else's.

Models — when the drawing computes

Until the model layer existed, everything above read TEXT: a label, an edge label, the caption. A compartment model, a Markov chain or a table of measurements was a picture whose numbers the author had typed. The model layer (lib/science/model/) is the one place a figure carries a number that can be computed with, and it is built as a registry of pure grammars so that a new formalism is a new module, not a new path through the composer.

  • The field. node.data.model and edge.data.model hold { kind, role, q, attrs } (schema.ts): the grammar, the part in it, the quantities as the TEXT a scientist typed ("0.3 /d", "β·S·I/N") and the structured attributes (a CPT, a table). Absence is default — a figure without it serialises byte-identically — and both readers are total and budgeted, so a share link cannot smuggle a megabyte in.
  • The grammars. registry.ts holds the frozen GRAMMARS list; each is a Grammar (contract.ts) with read, a CHEAP verify, parameters, a summary and at least one shipped template. readFigure partitions the canvas by model.kind, so a compartment model, a Markov chain and a Data node coexist on one figure; an edge between two kinds is refused as mixed, and nothing else is.
  • Two halves, two budgets. verifyModel — read plus each grammar's dimensional, row-stochastic, reference and missing-required checks — is the ONLY thing verifyFigure calls, so the kernel's finish gate, the readiness debounce and the API pay for a dimension check and nothing more. Simulation, solving and fitting live in grammars/*-analyse.ts behind the lazy ANALYSERS table (analysers.ts) and run from the Model tab, the agent tool or the API, memoised on the model content key.
  • model is a FindingKind. A rate on a flow that is not per unit time ("γ = 0.1" beside "β = 0.3 /d") is one finding, rule model.compartments.rate-dimension, on the arrow that carries it, with the fix "0.1 /d" in the model's own time unit. It costs rigour 0.34 like every other decidable kind and is BLOCKING in the Kernel, so a decidably wrong model cannot read sound.
  • Every number cites its method. ANALYSIS_METHODS (methods.ts) is the table every ModelResult.method resolves against — id, version, formula, reference — and a pinned result carries a provenance stamp naming the run entry and the input hash it was computed from.
  • The templates prove it. Every grammar ships a template whose captions write numbers ("R₀ = 3.0", "P(Sunny) = 75 %", "sd(x) = 2.14"), and tests/science-templates-verify.test.ts recomputes each of them from the model alone, to the precision written.

On the canvas: the palette's Models category drops nodes with a preset role; a double-click opens the model inspector (one field per declared quantity, the parsed dimension shown live, a red badge and a one-click fix on mismatch); the Model tab lists every model the canvas contains, runs it, shows the working and pins any number back onto the figure; the overlay layer draws the last run on the nodes and arrows (derived, never persisted) and the time scrubber walks a time course without re-integrating it. The ⌘K rows are Run model, Toggle model overlays, Export model, Pin last result, Export ledger, Recompute stale pinned numbers and, for a quantum circuit, Export OpenQASM.

The provenance ledger (lib/science/ledger.ts)

Every run — from the Model tab, an analysis card, an agent tool or the API — is a LedgerEntry: a sequential id (L-01, L-02, …), the method and its version from ANALYSIS_METHODS, the ids of exactly the nodes and edges it read plus any literal inputs (a pasted table's digest, a seed), the results, the working, and two FNV-1a hashes. figureHash is over the whole document and says WHICH figure; inputHash is over the canonical semantic fields of the inputs alone — label, endpoints, the validated model field, never a position, colour or size — with a membership digest @<kind> beside it so an element that JOINS a model (a fourth stock, a second outflow) is noticed even though every recorded id is untouched. "Pin to figure" writes a textOnly node whose label is formatQuantity output and whose data.provenance = { ledger, hash } names the entry; the precision-hygiene checks above then re-verify the label for free.

  • Three verdicts, no recompute. checkLedger compares digests. current: nothing the entry read has changed. stale: an input moved — rule provenance.stale, a rigour WARNING, because the number may well still be right and nothing on the verify path integrates to find out. orphan: the stamp traces to no entry of this ledger (a pin pasted from another figure, a ledger that did not travel) — rule provenance.orphan, a readiness GAP, "could not be traced", never a pass and never a fault. The Model tab's fold counts the pins as current / stale / untraced. A ledger whose checksum does not match is refused whole (provenance.ledger-unreadable), never partly trusted.
  • Two fixes. A stale or orphan finding carries Recompute (a new entry, the same node, the label rewritten) and Detach (the stamp dropped; the number becomes an ordinary claim the label checks still cover), through the same onApplyFix path as every other fix. The fold's "Recompute all" and ⌘K Recompute stale pinned numbers re-run every model a stale or untraced pin came from in dependency order (recomputePlan, Kahn's algorithm over entries that consume other entries; a cycle is refused by name).
  • Where it lives, and where it never goes. One additive key, glyph:biology:ledger:v1 (SCIENCE_KEYS.ledger). It rides in the JSON export, the share link and the SVG <metadata> through ledgerForExport, which is undefined for an empty ledger so a figure that never computed serialises byte-identically. It is NEVER in the live item stream: peers recompute, and until they do an orphan pin is a gap. LEDGER_BUDGETS caps it (500 entries; 256 results, 200 working lines and 64 literals per entry; 4,096 input ids), and when the entry cap bites pruneLedger keeps the entries a pin refers to, transitively, and makes room from the rest.

The grammars (lib/science/model/grammars/)

Twelve are registered. Each is one cheap module on the verify path (read

  • verify) and one *-analyse.ts module behind the lazy ANALYSERS

table; every analyser is pure, seeded where it samples, and refuses BY NAME — { name, message, budget? } — rather than returning NaN. What each reads, what it computes, and what it refuses:

  • compartments — stocks with an initial value (a counting label such as "990 people" is kept, never converted), flows with a rate or an expression over the figure's symbols, parameters. Computes conservation, R₀ (β/γ, or ρ(FV⁻¹) over the infected-tagged stocks), an RK4 or Dormand–Prince time course with residence times, half-lives, peaks, the final size and the steady state with its eigenvalues. Refuses a rate that is not [stock]·T⁻¹ (rate-dimension), an unknown symbol, stocks of mixed dimension, a singular V, a missing infected tag, a missing time unit.
  • reaction — species (c₀, symbol, formula), reaction arrows with mass-action, Michaelis–Menten, Hill or written rate laws, parameter and boundary species. Computes the stoichiometric matrix and its exact conservation laws, K_M / k_cat / V_max from the drawn mechanism, the time course, the steady state and its stability, flux control coefficients, and writes SBML. Refuses a missing rate constant, an unreadable rate, a negative concentration, a steady state that will not converge or is unphysical.
  • circuit — components with a dimensioned value (Ω, F, H, V, A, Hz), wires with named terminals, a ground. Computes nets by union–find, modified nodal analysis at DC or as phasors, every current and power, the Thévenin equivalent, τ and f_c, the resonance of an L–C pair. Refuses a floating net, a loop of ideal voltage sources, no reference node, a singular system, a two-terminal component with three unnamed wires.
  • freebody — a body with m (and g), forces as magnitude and angle, as components, as expressions over the other forces or as unknowns, a pivot. Computes ΣF, ΣM, a = ΣF/m, the free-axis component, and the unknown magnitudes by Gauss–Newton. Refuses underdetermined, indeterminate and inconsistent systems, a missing mass, static friction beyond μN.
  • markov — states, transitions with p (or a rate, making a CTMC). Row sums are checked on the verify path with normalising fixes. Computes the communicating classes and period, π by LU cross-checked by power iteration, mean return and mixing times, absorption probabilities and expected steps by the fundamental matrix. Refuses a row that does not sum to one, p and rate mixed on one chain, a reducible chain with nothing transient, a periodic or non-converged cross-check (partially).
  • petri — places with tokens, transitions, weighted arcs, checked bipartite. Hands the drawn net to lib/petri.ts unchanged for boundedness, reachability, deadlocks, liveness and P-invariants. Refuses by name when the reachability graph is capped, a deadlock search is truncated or a place is unbounded.
  • bayes — variables with states and a CPT, depends arrows. Rows summing to one, the table's shape, a DAG, ≤ 6 states and ≤ 4 parents are the verify path (a bad row carries a bayes:normalise-row fix). Computes exact marginals by variable elimination, P(child ∣ parent) and mutual information along every arrow, posteriors and the most probable explanation under evidence, d-separation, and minimal adjustment sets via lib/causaldag.ts. Refuses a cycle, a missing or mis-shaped CPT, evidence of probability zero, an elimination above 2²⁰ cells.
  • network — a modelled graph (signed edges, Boolean rules, feeding links) or, for the Model tab only, the plain drawing of three or more nodes. Computes degree and density, components, betweenness / closeness / eigenvector / PageRank, cut vertices and cycles, motifs against a degree-preserving null, trophic levels and synchronous Boolean attractors. Refuses a single node, dense centralities above 400 nodes, the cycle cap, the Boolean census above 20 nodes.
  • feynman — vertices, particle lines, external legs with E and p. Computes per vertex the conservation of Q, L_e, L_μ, L_τ and B in exact thirds, colour, the Standard Model coupling the vertex is, Mandelstam s and decay Q-values. Refuses a quantity not conserved, a lone coloured line, a vertex no coupling matches, a decay below the products' mass, a spacelike total momentum.
  • spacetime — events (t, x), worldlines, light rays, a frame S′ with β. Computes γ, every event in S′, intervals classified timelike / spacelike / lightlike, proper time per worldline and per path, simultaneity in S′, and draws the Minkowski plot. Refuses |β| ≥ 1 and any superluminal worldline, and names its path and pair budgets.
  • quantum — qubit lanes, gate tiles with a lane and a column, control lines, SWAP halves, measurements. Computes the state vector for n ≤ 12, probabilities and marginals, Bloch vectors, entanglement entropy across a cut of ≤ 6 qubits, equivalence to Bell / GHZ / QFT for n ≤ 8, and writes OpenQASM 2.0. Refuses a thirteenth lane, a wider cut or equivalence, a state that does not normalise, an unreadable angle; a gate on no lane or a control on its own target is a blocking finding.
  • data — a table node: named columns, rows of finite numbers (at most 5,000 × 64), an optional time column, a mapping from columns to the elements they observe. Computes column summaries; the Fit, Bootstrap and paste cards build on it, and a CSV/TSV dropped on the canvas becomes one. Refuses a ragged row and a column with no finite value.

The analytics (lib/science/analytics/, lib/science/analyse-cards.ts)

Everything here runs ON TOP of the grammars or the pasted table and cites its method the same way. The paste cards describe a sample (n, mean, SD, SEM, quartiles, D'Agostino–Pearson normality), give the t interval for its mean, fit ordinary least squares with a confidence band drawn, run Welch's t between two columns with Cohen's d and Hedges' g, and derive one quantity from two already on the figure with GUM-propagated uncertainty — under data.describe-sample, data.mean-ci, data.ols, data.welch and data.derive, the pasted table's digest recorded as a literal input so the pin goes stale when the data change. Fit (fit-model.ts) takes a catalogue curve — linear, exponential, rate-0/1/2, Michaelis–Menten, Hill, logistic, Arrhenius — or the DRAWN compartment model or reaction network, fits by Levenberg–Marquardt with standard errors, 95 % t intervals, AIC and BIC, and REFUSES as not-identifiable when the Jacobian's rank is below the parameter count (naming the null direction) or two estimates are correlated above 1 − 1e-6; a drawn model's estimates are written back onto their own arrows inside the undo stack. Sensitivity (sensitivity.ts) is central differences and the elasticity ∂ln y/∂ln p, ranked as a tornado, with one- and two-parameter sweeps and their threshold crossings. Monte Carlo (montecarlo.ts) draws only parameters with a stated ± — "β has no stated uncertainty — the studio does not invent data" — by Latin hypercube on mulberry32 under a seed (derived from the model content hash on the canvas, REQUIRED on the API), and returns type-7 quantiles, an outer 5–95 % envelope and P(y > c); the bootstrap resamples a column the same way. GLM (glm.ts) fits gaussian, binomial and poisson families by IRLS with pivoted QR, with VIF and Cook's D, refusing collinear columns and separation by name. Π groups (dimensional.ts) takes the quantities typed or already on the figure and returns Buckingham's groups by exact rational elimination, named where the catalogue knows them (Re, Fr, Nu, …). The four analytics that run over an analyser cite their own table, ANALYTICS_METHODS (analytics.sensitivity, .sweep, .monte-carlo, .bootstrap), recorded in the ledger entry's literals beside the analyser's method, which is the entry's method. Anything estimated above ~200 ms — sweeps, Monte Carlo, a fit through the integrator, a whole run — is a job for lib/workers/science-analytics-worker.ts, science-fit-worker.ts or science-model-worker.ts, with the same functions run inline under SSR and in tests; only summaries cross the boundary, never the samples.

The agent's compute tools (lib/agent/tools.ts, lib/science/toolkit/)

Four tools let a science preset RUN the figure it is judging: analyse_figure (every model, verified cheaply and analysed), simulate_model (the time course, R₀, peaks, final size, steady state), fit_data (a catalogue curve or the drawn model against a pasted table or the Data node) and sensitivity (the elasticities, ranked). They are offered exactly as check_science is — to the presets tagged science, to lab, evidence and data, and to any verify or critique stage — and to nobody else. Each returns every number on a line that BEGINS with the ledger id it was recorded under, and the tool's ledger continues the figure's own numbering, so a computation the figure already recorded is reused under its existing id. The evidence rule (toolkit/evidence.ts) is then enforced, not suggested: a finding may cite only an id a compute tool emitted in this run, and a finding that cites any other id is dropped with a visible warning. pin on any of the four writes the named results onto the figure as text nodes carrying their provenance.

The API (POST /api/v1/science)

The model operations join the checkers: model (readings and the cheap findings), simulate, steady-state and r0 (one integration, one entry), fit, sensitivity, montecarlo (seed required, n ≤ 500), graph, describe, regress, pi-groups, ledger-check (the three verdicts over a figure and its ledger) and quantities (what the labels say, as typed). Every reply that computes carries the new entries and their ids; GET lists the operations with the method ids each may cite and the full method register (ANALYSIS_METHODS + ANALYTICS_METHODS), so a method@version in a reply resolves without a second request. A figure is read whole up to 2,000 elements.

Budgets

Every solver takes a Budget — maxSteps (integration), maxEvals (fits and sweeps), maxIterations (fixed points and solves), maxNodes (structural reads) — and a refusal that hits one names it: refusal.budget = { name, limit, actual }. The canvas runs under DEFAULT_BUDGET (100,000 steps; 20,000 evaluations; 10,000 iterations; 2,000 nodes) and the server under SERVER_BUDGET (20,000 / 20,000 / 10,000 / 2,000). Beside these: Monte Carlo at most 20,000 runs (500 on the server), 64 parameters and 32 outputs, bands on 101 frames; sweeps of 2,000 points, heatmaps of 64 × 64; the network census dense above 400 nodes only by refusal, 200 cycles, 200 null samples, Boolean above 20 nodes refused; 12 qubits, an 8-qubit equivalence and a 6-qubit cut; 6 states and 4 parents per Bayesian variable and 2²⁰ factor cells; a CPT of 4,096 cells and a Data node of 5,000 × 64 in the schema; and the ledger's caps above.

What it deliberately does NOT do

  • It does not judge taste. A figure is not wrong for being sparse, or for using a layout someone else would not have chosen.
  • It does not replace a safety data sheet, a statistician, or peer review. Where a module encodes regulatory or methodological guidance it says so, and says what it is a heuristic for.
  • It does not invent data. When a value is unknown it is omitted, and the omission is visible.
  • It does not call a model to decide anything. Every verdict in this layer is arithmetic. The agents use models; the checkers do not.

Adding a checker

  1. Write the pure module in lib/science/, with a header comment that says why it exists and what it refuses to do.
  2. Write tests/science-<name>.test.ts with real, checkable values from the authoritative source — and a "never throws on garbage" suite.
  3. Add a FindingKind and a checker function in lib/science/verify.ts. Keep it conservative: it must return nothing when it is not confident. A new MODEL grammar is different: write lib/science/model/grammars/<kind>.ts (read + cheap verify + a template) and <kind>-analyse.ts (the expensive half), register the first in registry.ts and the second in analysers.ts, add its methods to methods.ts, and let tests/science-grammar-contract.test.ts and tests/science-templates-verify.test.ts hold it to the contract.
  4. Add the operation to app/api/v1/science/route.ts and its entry to OPERATIONS, so the API documents itself.
  5. If it needs a surface, add a tab or a card to components/biology/ScienceIntelligencePanel.tsx.

Something unclear or out of date on this page? Tell us from the Support link in any studio — the flowss team reads every report.

© 2026 Voranox Inc. flowss — Flow Systems Studio. All rights reserved.

This documentation, its text and its examples are protected by copyright. Engine and format names are trademarks of their respective owners — see the terms and copyright and licences.