← Plus-Minus / API
Tokens

Drive Plus-Minus from your own code

Everything the web page does is available over HTTP: post one measurement sheet — the inputs with their values, units and uncertainties, and the formula that turns them into the result — together with the facts the free engine computed from it, and get back either a metrologist's budget review (every stated uncertainty judged, the contributions a competent evaluation would also carry as ready-to-paste sheet lines, the improvements ranked by share of variance, the correct statement and a verdict) or the uncertainty report section (the methods-and-results prose, the budget table and the thirteen-point reporting checklist a lab report, paper, calibration certificate or internal record needs).

The natural use is a bench script that keeps the uncertainty budget of a standing measurement under review as the sheet changes, or a CI job that refuses to publish a number whose budget the engine cannot compute or whose statement the reply re-rounded. One thing is different from most apps on this platform, and it is the whole of step 4: the arithmetic is not the model's. A free engine — the same units.js and unc.js the browser runs — does the dimensional analysis, the GUM budget, Welch-Satterthwaite, the coverage factor, two Monte Carlo propagations and the rounding, and hands the model a facts object it must quote rather than recompute. An API caller has to produce that object too.

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{ "ok": true,  "data":  { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "details": { ... } } }

Send your token as Authorization: Bearer … on every call. The app slug travels in the body of /guest as {"slug": "plus-minus"}; after that the token itself carries the app, so a run needs only Authorization, Content-Type: application/json and the Idempotency-Key described in step 5.

The request body for /estimate, /run and /run-stream is the input object itself — not wrapped in anything. A body of {"input": {…}} returns 200 and quietly hides every field from the model, so a run that looks fine comes back reviewing nothing. Guard it the way the app's own client does: its mustBeObject() refuses to send anything that is not a plain JSON object, because /estimate happily prices a bare string, a number, null and [] alike.

Error codes

codestatuswhat to do
unauthorized401The token is missing, malformed or expired. Get a new one from the token page.
payment_required402The balance is below min_credits, or a guest token tried a metered run. Call /estimate first, then sign in and top up.
validation_error400The body is not a plain JSON object, or a field is the wrong type — every field of this app's input is a string, facts included. A facts object sent as an object rather than as a JSON string is the common one.
not_found404Unknown job id, or the app slug does not exist. Check the id you polled with.
rate_limited429Too many requests. Back off and retry; do not tight-loop a poll.
internal500A server-side failure. Retry with the same Idempotency-Key so you are not billed twice.

Replaying an Idempotency-Key with a changed body is not a retry and is rejected rather than billed. When you change the input — an edited sheet, a different context_hint, another audience — bump the attempt suffix on the key instead.

The field to get right first: task

Plus-Minus is one app with two lanes, and task is what chooses between them. It is the first field of every request body:

taskwhat comes backextra input fields
"budget"A review of the budget: measurand, contributors judging every stated uncertainty as reasonable, questionable or unsupported, missing contributions each with a ready-to-paste sheet_line, ranked improvements, the statement copied verbatim from the engine, and a verdict of report_as_is, budget_incomplete or recompute_first.none
"report"The uncertainty section: section_title, three to six paragraphs of methods-and-results prose, a budget_table with one row per input that carries an uncertainty, the result_statement verbatim, a thirteen-row checklist (R-001 to R-013) and the caveats the report cannot claim its way out of.audience, review_notes

The system prompt routes on task and never blends the two contracts in one reply. A missing or unrecognised task is not an error: the model answers the closest lane — an audience or review_notes field means report — names the lane it actually answered in the reply's lane field, and says so in notes_on_input. So branch on lane in the reply, never on the task you believe you sent. The app's own normaliser goes further: a reply with no usable lane is classed as report when it carries paragraphs, checklist or budget_table, and as budget otherwise.

The two lanes chain, and the handoff is the point of the app. Run budget over a sheet, act on what it says — paste the missing[].sheet_line lines onto the sheet and recompute, or accept some of them as known gaps — then send the accepted items as the review_notes of a report run over the same sheet. The report lane turns them into the caveats paragraph nobody remembers to write. The web page does this with one button on either side: Add to sheet for the proposed lines, and Write the report from this review for the handoff.

The sheet grammar

sheet is the run's only evidence and it is a small, strict language. The engine parses it line by line; the model never parses it at all, it reads the engine's facts. Get the grammar right and everything downstream works; get it wrong and the engine emits an error-level flag, facts.ok is false, and the reply tells you what to fix on the sheet instead of reviewing a budget that does not exist.

Six kinds of line, and blank lines are ignored:

# a comment: everything from a # to the end of the line is stripped first

g = 4*pi^2*L/T^2                 # a FORMULA line: a name = an expression
L = 0.9950 m ± 0.0005 m  [B rect half]   # an INPUT line: a name = a quantity
corr(L, T) = 0.4                 # a CORRELATION line
level: 95%                       # a DIRECTIVE
reference = 9.8062 m/s^2 ± 0.0001 m/s^2  # a directive that takes a quantity

A line with no = at all, and no corr(...) or directive keyword, is an error. The distinction between an input line and a formula line is made on the right-hand side: if it parses as a quantity it is an input, otherwise it is an expression. Names are [A-Za-z_][A-Za-z0-9_']* — letters, digits, underscore and the prime, starting with a letter or an underscore — and a name may be defined only once, as an input or as a formula but never as both.

Input lines: the quantity forms

Three ways to write a value with its uncertainty, plus one for a value without:

formexamplewhat the engine reads
name = value unit ± u unitL = 0.9950 m ± 0.0005 mThe full form. The uncertainty's unit may differ from the value's as long as the dimensions match — m ± 0.5 mm is converted for you; a mismatch (m ± 0.5 s) is an error, not a warning.
name = value unit ± uT = 2.0012 s ± 0.0008With no unit after the ±, the uncertainty is in the value's unit.
name = value unit ± p%R = 100 ohm ± 0.1%A relative uncertainty: u = |value| · p / 100. The percent sign must follow the number directly.
name = value(uu) unitm = 2.500(12) kgThe concise parenthesis notation: the digits in brackets are the uncertainty in the last places of the value, so this is 2.500 kg ± 0.012 kg. A value written this way is exempt from the digit-hygiene flags, because the notation fixes the digits by construction.
name = value unitn_swings = 20No ± at all: the input is exact, u = 0, and the engine raises an exact-input flag asking you to state an uncertainty unless it is a defined value or a count.

+/- and +- are accepted as ASCII spellings of ±. A value may be written in exponent form (6.626e-34), and the exponent counts when the engine compares the decimal places of a value against those of its uncertainty.

Tags

A bracketed list at the end of an input line says how the uncertainty was evaluated. Tags are separated by spaces or commas, and the brackets must be the last thing on the line before any comment:

tagmeaning
A, typeA, type-AType A: a statistical evaluation of repeated observations.
B, typeB, type-BType B: evaluated by any other means — a specification, a certificate, a resolution, judgement.
n=10The number of repeated observations. It implies Type A, and it sets the degrees of freedom to n - 1 unless dof= says otherwise.
dof=9, nu=9Degrees of freedom directly. Absent and without n=, the input carries infinite degrees of freedom.
k=2The figure on the line is an expanded uncertainty at this coverage factor; the engine divides by k to get the standard uncertainty. This is how a calibration certificate's value goes on the sheet.
rect, rectangular, uniformRectangular distribution.
tri, triangularTriangular distribution.
U, u-shaped, arcsineU-shaped (arcsine) distribution — the classic case being a quantity cycling between two limits.
normal, gauss, gaussianNormal distribution. This is the default when no distribution tag is present.
half, halfwidth, half-widthThe figure after the ± is a half-width, not a standard uncertainty: divide it by the distribution's factor. half on its own implies rect.

The half-width divisors, applied only when half is present:

[rect half]   u = a / sqrt(3)     a 1 mm tape resolution: a = 0.5 mm, u = 0.29 mm
[tri half]    u = a / sqrt(6)
[U half]      u = a / sqrt(2)
[half]        same as [rect half]

Order of operations inside one line: the half-width divisor first, then the division by k= if both are present. A tag the engine does not know is not fatal — it raises a tag-unknown warning flag and the line is otherwise read normally, which is worth knowing because a typo such as [rect halfwidth=1] silently loses the half-width.

Formula lines

A formula line is name = expression. Several are allowed, and they may refer to each other in the order written. The measurand — the quantity the whole budget is about — is the last formula on the sheet unless a measurand: directive names another one. The operators are + - * / ^ (with ** accepted for ^), unary minus, and parentheses; ^ binds tighter than * and /, which bind tighter than + and -. The functions are:

sqrt  abs  exp  ln  log  log10  log2
sin  cos  tan  asin  acos  atan          # trigonometry works in RADIANS
atan2(y, x)   pow(a, b)   min(a, b)   max(a, b)

Every formula is checked dimensionally as it is evaluated: adding metres to seconds, or taking the logarithm of something with a dimension, is an error-level flag rather than a number. The sensitivity coefficients are exact — the engine carries a gradient alongside every value (forward-mode automatic differentiation), so c_i = ∂f/∂x_i is the analytic derivative and not a finite difference. An input defined on the sheet and never used by any formula is reported in facts.unused_inputs and as an unused-input flag.

Correlations

corr(L_1, L_2) = 0.85

One line per pair, with r in [-1, 1]. Both names must be inputs on the sheet. A correlation adds its cross term 2·c_a·c_b·u_a·u_b·r to the variance and is carried into the Monte Carlo sampling through a Cholesky factor. Two honest consequences the engine flags rather than hides: Welch-Satterthwaite ignores correlation terms, so veff is only approximate once you declare any; and the per-input share_pct values are shares of the diagonal variance, so they no longer sum to 100 % of the combined variance. The page's own check for the report lane's budget table is therefore skipped when facts.correlations is non-empty.

Directives

directiveexampleeffect
title:title: Pendulum, bench 3A label for the sheet. It travels in the facts only as part of the sheet text.
measurand:measurand: gNames which formula is the measurand. Without it, the last formula line wins. Naming something that is not a formula is an error.
unit:unit: mmReport the result in this unit instead of the coherent SI one the dimensions imply. It must match the result's dimension, and it may not be an offset unit — a result in kelvin is reported in kelvin, never in degC.
k:k: 2Fix the coverage factor instead of taking it from the t distribution. facts.result.k_source then reads "sheet", and if the t distribution at veff would have wanted a noticeably larger factor the engine raises a k-assumed warning: the stated interval covers less than it claims.
level:level: 95%The coverage probability. Written as a percentage or as a fraction; the default is 95.45 %, the level at which k = 2 is exactly right for a normal distribution with infinite degrees of freedom.
reference =, expected:reference = 9.8062 m/s^2 ± 0.0001 m/s^2A value to compare against. The engine computes the normalised error E_n = (y - y_ref) / sqrt(U² + U_ref²) and raises en-ok or en-fail. A reference written with [k=2] carries an expanded uncertainty already; one written bare is a standard uncertainty and is expanded with the result's own k. This line is not an input to the model and has no sensitivity coefficient.
result:, draft:result: 9.81 m/s^2 ± 0.02 m/s^2 (k=2)Your own draft statement, checked against the computed one: value agreement within U, a missing coverage factor, a ± that is more than a third out, too many significant figures, and a value and an uncertainty rounded to different places. Each is its own flag — draft-value, draft-k-missing, draft-mismatch, draft-digits, draft-place, draft-no-u — and draft-value is one of the two rules that force the verdict to recompute_first.

Built-in constants

A name used in a formula and never defined on the sheet is looked up in the engine's constant table. The exact ones (exact by the 2019 SI definitions) contribute nothing to the budget; the four measured ones carry their CODATA 2018 standard uncertainty and join the budget as Type B inputs, which is why they can appear in facts.inputs and in the report lane's budget table. Defining an input with one of these names shadows the constant and raises a constant-shadow flag:

namevalueunitnote
pi3.14159265358979—pi (exact)
e2.71828182845905—Euler's number (exact)
c299792458m/sspeed of light in vacuum (exact)
h6.62607015e-34J sPlanck constant (exact)
hbar1.054571817e-34J sreduced Planck constant (exact)
k_B1.380649e-23J/KBoltzmann constant (exact)
N_A6.02214076e231/molAvogadro constant (exact)
q_e1.602176634e-19Celementary charge (exact)
g_n9.80665m/s^2standard acceleration of gravity (defined)
R_gas8.314462618J/(mol K)molar gas constant (exact)
sigma_SB5.670374419e-8W/(m^2 K^4)Stefan-Boltzmann constant (exact)
G6.67430e-11 ± 0.00015e-11m^3/(kg s^2)Newtonian constant of gravitation — carries an uncertainty
eps_08.8541878128e-12 ± 1.3e-21F/mvacuum electric permittivity — carries an uncertainty
mu_01.25663706212e-6 ± 1.9e-16N/A^2vacuum magnetic permeability — carries an uncertainty
m_e9.1093837015e-31 ± 2.8e-40kgelectron mass — carries an uncertainty
m_p1.67262192369e-27 ± 5.1e-37kgproton mass — carries an uncertainty
u_amu1.66053906660e-27 ± 5.0e-37kgatomic mass constant — carries an uncertainty

In facts.inputs a constant is marked "constant": true with a note, and an exact one also carries "exact": true. The budget lane judges only the inputs that are actually on the sheet — a contributors entry naming pi is a warning on the page, not a virtue.

Units

Units are parsed into the seven SI base dimensions, so any spelling that resolves to the right dimension works and mixed systems convert silently: m kg s A K mol cd rad sr Hz N Pa J W C V F ohm S Wb T H lm lx Bq Gy Sv kat L, the time units min h d yr, the angles deg arcmin arcsec (and °), the energy and mass units eV cal Wh Ah t u Da, the pressures bar atm mmHg torr psi, the imperial set ft in yd mi nmi lb oz mph kn ha acre gal, the ratios % permille ppm ppb, and the offset temperatures degC °C degF °F. Write compound units with *, / and ^: m/s^2, J/(mol K), kg m^2/s^3. SI prefixes apply wherever the bare symbol is not itself a unit — mm, kPa, µV, MHz.

Two traps worth naming. An offset unit is read as an absolute temperature: T = 20 degC is 293.15 K, and the engine raises an offset-unit flag because a temperature difference of 20 degrees should be written 20 K. And u is the dalton, not micro-anything, while T is the tesla — inputs named after unit symbols are fine (the example sheet's T is a period in seconds), because names and units are read in different positions on the line.

Clipping

The browser sends at most 12,000 characters of sheet, clipping whole lines from the middle and leaving a marker in place of the cut:

# [... 42 lines (3120 characters) cut from the middle ...]

Clip the same way if you send more, and keep the marker: the prompt keys on it, refuses to claim anything about the missing middle, and says so in notes_on_input. The counts also travel in the facts as facts.clipped = {"cut": 3120, "lines_cut": 42}. In practice a sheet that long is several measurements stacked in one file, and the right move is to split it: a budget is about one measurand.

When the engine cannot compute

Parse and dimension failures do not stop the run. They land in facts.errors[], they become flags with "severity": "error", and facts.ok is false — there is then no facts.result and no facts.mc. The prompt's rule for that case is explicit: say what to fix on the sheet and keep the rest short. The web page refuses to spend a run at all in this state; an API caller can still spend one, and the honest thing is to check facts.ok first and save the credits. The same check is what stops a CI job from reporting "the review found nothing serious" about a sheet that never computed.

1. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool — it reads the same storage the app itself uses and prints the token for you.

A guest token can call /me and /estimate. Both lanes are metered, so reviewing a budget and writing a report section each need a personal token from signing in. Nothing about the engine is metered: the arithmetic in step 4 costs nothing and needs no token at all.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://plus-minus.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="aut_YOUR_TOKEN"
#
# To mint a guest token from the command line instead. A guest token is enough for
# /me and /estimate; a budget or report run needs a personal token.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" \
  -d '{"slug": "plus-minus"}'
# {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}

2. A tiny client

One helper that adds the headers, unwraps data and raises on error.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="plus-minus"
TOKEN="$SKILLSAFE_TOKEN"   # from https://plus-minus.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

3. Check the session and the balance

GET /me tells you whether the token is a guest or a person, and what the balance is. The object is small and carries exactly three things: subject_type — guest or user — subject_id, and credits, the wallet balance. There is no username in it, so "signed in" is subject_type === "user" and nothing else; a guest can price a run but cannot start one. Compare credits against min_credits from the next step before you run, so a shortfall surfaces as your own clear message rather than a 402.

call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}

# A guest token answers the same call with subject_type "guest" and cannot run
# either lane. Gate on it before you spend a poll loop finding out:
call me | grep -q '"subject_type":"user"' || {
  echo "sign in at https://plus-minus.skillsafe.ai/tokens.html first" >&2
  exit 1
}

4. Price the run — free

The input object is exactly what the app's own form submits. It is always a JSON object — never a bare string, never wrapped in an input key — and every one of its fields is a string, because the transport takes scalars only. That is why facts is a JSON string rather than a nested object:

fieldtypemeaning
taskstring, required"budget" or "report". The lane. It routes the prompt, and a missing or unknown value degrades to the closest lane rather than failing — the reply names what it answered in lane and in notes_on_input.
sheetstring, requiredThe measurement sheet, in the grammar above. Clipped from the middle at 12,000 characters with the marker kept. This is the text the reply quotes lines from, and the only thing the model knows about the measurement besides context_hint.
context_hintstring, optionalUp to 600 characters about the instrument, the setting and the purpose — "First-year teaching lab. A steel bob on a string, a metre tape, a hand-held stopwatch…". It is what separates a plausible criticism from a guess: without it the model knows the numbers and nothing about how they were obtained. It also sets the language of the prose — write the hint in German and the review comes back in German, with the keys and enum values still in English.
factsstring, always presentThe engine's output, JSON.stringify-ed. Everything numeric in the reply is quoted from here. Shape and the script that computes it are below.
report lane only
audienceenumlab_report (a student or teaching lab), paper (a journal methods section), certificate (a calibration or test certificate, ISO/IEC 17025 7.8) or internal (an internal record). It changes the register, the section title and whether traceability and the standards followed get their own paragraph. Anything else falls back to lab_report.
review_notesstring, optionalUp to 2,000 characters: the items you accepted from a budget review, as text. This is the handoff — the report lane turns them into caveats, which is the paragraph that says what the report cannot claim. Omit it and the caveats are derived from the sheet and the flags alone.
both lanes, rarely needed
retry_notestring, optionalWhat the app sends on its one automatic retry when a reply did not parse: a sentence naming the failure and restating the contract. If you build the same recovery, send the same kind of note rather than resending the identical body — and remember to bump the attempt suffix on the key.

facts: the numbers are the engine's, not the model's

In the browser this object is computed for free, before the run, by units.js and unc.js — the same two files the page loads. The model is told, in the system prompt, never to recompute, re-round or restate any of it with different digits, and to quote facts.result.statement.expanded and .concise verbatim wherever the result is stated. Here is the object for the running example, complete but abbreviated in the long arrays:

{
  "ok": true,
  "measurand": "g",
  "formulas": [ { "name": "g", "expr": "4*pi^2*L/T^2", "unit": "m/s^2", "ok": true } ],
  "inputs": [
    { "name": "L", "value": 0.995, "u": 0.000288675, "unit": "m", "type": "B",
      "distribution": "rectangular", "dof": "infinite", "rel_pct": 0.0290126,
      "sensitivity": 9.85777, "contribution": 0.00284569, "share_pct": 11.6357 },
    { "name": "T", "value": 2.0012, "u": 0.0008, "unit": "s", "type": "A",
      "distribution": "normal", "dof": 9, "rel_pct": 0.039976, "n": 10,
      "sensitivity": -9.8026, "contribution": 0.00784208, "share_pct": 88.3643 },
    { "name": "pi", "value": 3.14159, "u": 0, "unit": "", "type": "B",
      "distribution": "normal", "dof": "infinite", "rel_pct": 0, "constant": true,
      "note": "pi", "exact": true, "sensitivity": 6.24427, "contribution": 0, "share_pct": 0 }
  ],
  "correlations": [],
  "errors": [],
  "unused_inputs": [],
  "flags": [
    { "id": "E-001", "rule": "dominant", "severity": "info",
      "message": "T carries 88 % of the variance of g - improving anything else changes almost nothing",
      "ref": "T" },
    { "id": "E-002", "rule": "low-dof", "severity": "warn",
      "message": "effective degrees of freedom 11.5 (Welch-Satterthwaite) - the coverage factor for 95.45 % is 2.24, not 2",
      "ref": "" },
    { "id": "E-003", "rule": "en-ok", "severity": "info",
      "message": "E_n = 0.12 against the reference 9.80620 m/s^2 - consistent with the reference (|E_n| <= 1)",
      "ref": "" }
  ],
  "clipped": { "cut": 0, "lines_cut": 0 },
  "result": {
    "name": "g", "value": 9.80848, "u": 0.00834243, "U": 0.0186871,
    "k": 2.24, "k_source": "t", "k_from_t": 2.24193, "level_pct": 95.45,
    "veff": 11.5263, "unit": "m/s^2", "unit_si": "m/s^2", "rel_pct": 0.0850533,
    "statement": {
      "expanded": "g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)",
      "standard": "g = 9.8085 m/s^2, u = 0.0083 m/s^2 (standard uncertainty)",
      "concise":  "g = 9.8085(83) m/s^2",
      "value_text": "9.808", "U_text": "0.019", "u_text": "0.0083",
      "k_text": "2.24", "level_text": "95.45 %"
    }
  },
  "mc": {
    "normal": { "n": 60000, "mean": 9.80846, "sd": 0.00834484, "lo": 9.79175, "hi": 9.8251,
                "skew": -0.00740069, "level_pct": 95.45,
                "what": "every input normal - isolates non-linearity" },
    "stated": { "n": 60000, "mean": 9.80848, "sd": 0.00932407, "lo": 9.78944, "hi": 9.82773,
                "skew": 0.0189543, "level_pct": 95.45,
                "what": "the stated distributions - the coverage interval to report" }
  },
  "en": { "value": 0.122144, "reference": 9.8062, "reference_U": 0.000224 }
}

Field by field, the parts a caller actually reads back: inputs[] is the budget — value, standard uncertainty u, type (A or B), distribution, dof (a number or the string "infinite"), the sensitivity coefficient, the contribution |c_i|·u_i and its share_pct of the variance. result carries the combined standard uncertainty u, the coverage factor k with k_source ("t" from the t distribution at veff, or "sheet" when a k: directive fixed it), the expanded uncertainty U, and the three rounded statements. mc.normal re-runs the propagation with every input normal, so its sd against result.u isolates non-linearity; mc.stated uses the distributions you tagged, and its lo/hi is the coverage interval to report when the flags say a symmetric ± misleads. flags[] is the reconciliation list of step 7.

Computing it outside the browser. The engine is two plain scripts served by this host and they run unmodified in Node: give them a window to attach to, evaluate them, and call analyze() then toFacts(). Twenty lines, no dependencies, no token, no charge. Every sample below is the same three steps — fetch /units.js and /unc.js, evaluate them with global.window = {} in scope, then JSON.stringify(window.PMUnc.toFacts(window.PMUnc.analyze(sheet), {cut: 0, lines_cut: 0})) — and the non-JavaScript ones drive that script as a subprocess, which is both the shortest and the only way to be certain your numbers are the ones the app would have produced.

# The sheet, saved as sheet.txt - this is the worked example used everywhere below.
cat > sheet.txt <<'SHEET'
# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m      [B rect half]   # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s      [A n=10]        # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2      # local g from the survey office
SHEET

# The engine, fetched once and kept next to your script. These are the same two
# files the web page loads; re-fetch them when the app is released again.
curl -sS -o units.js "https://plus-minus.skillsafe.ai/units.js"
curl -sS -o unc.js   "https://plus-minus.skillsafe.ai/unc.js"

# facts.js - 12 lines, no dependencies, no token, no charge.
cat > facts.js <<'JS'
const fs = require("fs"), vm = require("vm");
global.window = {};                                   // the two files attach to it
for (const f of ["units.js", "unc.js"]) vm.runInThisContext(fs.readFileSync(f, "utf8"));
const U = window.PMUnc;
const sheet = fs.readFileSync(process.argv[2], "utf8").replace(/\r\n?/g, "\n");
const clip = U.clipMiddle(sheet, U.MAX_SHEET);        // 12,000 chars, whole lines
const facts = U.toFacts(U.analyze(clip.text), { cut: clip.cut, lines_cut: clip.lines_cut });
if (!facts.ok) console.error("engine errors:", JSON.stringify(facts.errors));
process.stdout.write(JSON.stringify(facts));          // ONE line: the `facts` string
JS

FACTS=$(node facts.js sheet.txt)
echo "$FACTS" | head -c 120
# {"ok":true,"measurand":"g","formulas":[{"name":"g","expr":"4*pi^2*L/T^2","unit":"m/s^2",...

The lazy alternative, and what it costs you. The field is a string, so "facts": "{}" is accepted and the run proceeds. What you get back is a review with no numbers in it: nothing to quote, no statement to copy, no share_pct to rank improvements by, no flags to reconcile — and, worse, a model asked to comment on a measurement whose uncertainty nobody computed will write plausible prose anyway. The app itself never does this: every run from the page carries a real facts object, and the page then checks the reply against it. If you cannot run the engine, at least know that you are buying prose, not a budget.

/estimate creates no job and charges nothing. It returns the model binding — model is gpt-5.6-terra, model_alias is gpt-terra, and markup_bps — plus sponsor_enabled and the reservation: hold_credits is what gets held, and min_credits is the balance you must clear to start at all. The hold is a reservation, not the price. It prices the full output cap, so the charged_credits on the settled job is usually far lower. Budget against hold_credits, report against charged_credits.

Price each lane separately — the app throws away the budget lane's estimate the moment you switch lanes, for exactly this reason. A report body carries review_notes and prices differently from a budget body over the same sheet, and facts is usually the largest field in both: a sheet with twenty inputs is twenty budget rows of input tokens whether you read them or not.

# Lane A - the budget review. Build the body with a JSON encoder, never with
# string concatenation: the sheet has newlines and the facts string has quotes.
HINT="First-year teaching lab. A steel bob on a string, a metre tape, a hand-held stopwatch; the period is the mean of ten timings of twenty swings each."

INPUT=$(python3 - "$FACTS" "$HINT" <<'PY'
import json, sys
print(json.dumps({
    "task": "budget",
    "sheet": open("sheet.txt").read(),
    "context_hint": sys.argv[2],
    "facts": sys.argv[1],          # the STRING from node facts.js
}))
PY
)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"gpt-5.6-terra","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":2652,"min_credits":310,"sponsor_enabled":false}}
#
# estimate is FREE. It creates no job and charges nothing. hold_credits is what
# gets RESERVED; charged_credits on the settled job is normally much lower.

# Lane B - the report section over the SAME sheet and the SAME facts, carrying the
# items accepted from the review. This is the handoff.
NOTES="Accepted from the budget review: reaction time of the stopwatch operator (Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the finite-amplitude correction were not evaluated; both are noted as caveats rather than added to the budget."

REPORT_INPUT=$(python3 - "$FACTS" "$HINT" "$NOTES" <<'PY'
import json, sys
print(json.dumps({
    "task": "report",
    "sheet": open("sheet.txt").read(),
    "context_hint": sys.argv[2],
    "audience": "lab_report",      # lab_report | paper | certificate | internal
    "review_notes": sys.argv[3],
    "facts": sys.argv[1],
}))
PY
)

call estimate "$REPORT_INPUT"   # price each lane separately

5. Run it, then poll

POST /run reserves hold_credits, starts the job and returns a job_id at once; the reply arrives when GET /jobs/{id} reports status: "succeeded". Send an Idempotency-Key on every run. The app's key is plus-minus:<task>:<hash of the sheet, hint, audience and notes>:a<attempt>: the lane is in it, so the same sheet through both lanes is two runs, and a retried request with the same key returns the same job instead of billing twice. Bump the attempt suffix only when you mean to run again.

# Same body as the estimate. The key is slug:task:hash:attempt.
KEY="plus-minus:budget:$(shasum -a 256 body.json | cut -c1-16):a1"
JOB=$(curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/run" \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  --data-binary @body.json | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

# Poll until terminal. succeeded | failed | cancelled.
while :; do
  STATE=$(curl -sS "https://api.skillsafe.ai/v1/app-api/jobs/$JOB" \
    -H "Authorization: Bearer $SKILLSAFE_TOKEN")
  STATUS=$(echo "$STATE" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  case "$STATUS" in succeeded|failed|cancelled) break;; esac
  sleep 2
done
echo "$STATE" | python3 -c 'import sys,json;d=json.load(sys.stdin)["data"];print(d["status"],d.get("charged_credits"),"credits, truncated:",d.get("truncated"));print(d["output"]["output"][:400])'

6. Or stream it

POST /run-stream takes the same body and the same key and answers with server-sent events. From a script you see event: delta frames carrying pieces of the reply and one final event: done with the finished job. From a browser page the platform sends only tick heartbeats and the done event — which is why the app's progress card advances on elapsed time between the real signals, and why a streaming preview built in a page would be dead code. Concatenate the deltas; parse only when done arrives.

curl -sN -X POST "https://api.skillsafe.ai/v1/app-api/run-stream" \
  -H "Authorization: Bearer $SKILLSAFE_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Accept: text/event-stream" \
  -H "Idempotency-Key: plus-minus:budget:$(shasum -a 256 body.json | cut -c1-16):a1" \
  --data-binary @body.json
# event: job      data: {"job_id":"job_..."}
# event: delta    data: {"text":"{\"lane\":\"budget\",\"title\":\"g from a simple pendulum - budget review\","}
# event: delta    data: {"text":"\"summary\":\"..."}
# event: done     data: {"status":"succeeded","charged_credits":412,"truncated":false,"output":{"output":"{...}"}}

7. Parse the result

The reply is one JSON object, as a string, inside the job's output.output. Unwrap it twice: once for the envelope, once for the model's text. Then do what the page does before it trusts a single field — check the reply against the facts you sent. Three checks catch nearly everything that goes wrong: the statement must be the engine's verbatim; every engine flag id must appear exactly once in coverage; and each lane's list must be complete — every sheet input judged once (budget), one budget-table row per input with an uncertainty and exactly thirteen checklist rows (report). The page shows the engine's statement regardless of what the model wrote, and lists every disagreement under "The page disagrees with the reply".

# Unwrap twice, then check the statement and the flag coverage with python.
echo "$STATE" | python3 - body.json <<'PY'
import json, sys
state = json.load(sys.stdin)["data"]
facts = json.loads(json.load(open(sys.argv[1]))["facts"])
reply = json.loads(state["output"]["output"])          # the model's single JSON object
want = facts["result"]["statement"]["expanded"]
got = reply["statement"]["expanded"] if reply["lane"] == "budget" else reply["result_statement"]
print("statement verbatim:", got == want)
ids = [f["id"] for f in facts["flags"]]
cov = [c["id"] for c in reply["coverage"]]
print("flags covered once:", sorted(ids) == sorted(cov))
if reply["lane"] == "budget":
    names = sorted(x["name"] for x in facts["inputs"] if not x.get("constant"))
    print("inputs judged:", sorted(c["name"] for c in reply["contributors"]) == names, "| verdict:", reply["verdict"])
else:
    print("13 checklist rows:", len(reply["checklist"]) == 13, "| met:", sum(c["status"] == "met" for c in reply["checklist"]))
PY

The output contract

Every reply carries the common envelope and then one lane body. The keys and enum values are exact; the page re-sequences ids defensively (M-001…, I-001…, R-001…) and falls back on any enum it does not recognise, but a script should treat an unknown value as a defect, not a feature.

{
  "lane": "budget" | "report",
  "title": string,
  "summary": string,
  "coverage": [ { "id": "E-001", "status": "confirmed" | "downgraded" | "dismissed" | "merged", "ref": string, "note": string } ],
  "notes_on_input": string,
  ...lane body
}

The budget body

{
  "measurand": { "name": string, "quantity": string, "method": string },
  "contributors": [ { "name": string, "status": "reasonable" | "questionable" | "unsupported", "note": string } ],
  "missing": [ { "id": "M-001", "what": string, "type": "A" | "B",
                 "distribution": "normal" | "rectangular" | "triangular" | "u_shaped",
                 "how_to_evaluate": string, "sheet_line": string, "priority": "high" | "medium" | "low" } ],
  "improvements": [ { "id": "I-001", "action": string, "targets": [string], "expected_effect": string, "effort": "low" | "medium" | "high" } ],
  "statement": { "expanded": string, "concise": string, "wording": string },
  "verdict": "report_as_is" | "budget_incomplete" | "recompute_first",
  "verdict_reason": string
}

statement.expanded and statement.concise are copies of facts.result.statement; wording is the one sentence a report would carry and may quote only digits that occur in facts.result.statement or facts.mc.stated. Every contributors[].name is a sheet input, each exactly once. Every missing[].sheet_line parses in the sheet grammar under a new name — the page refuses to add one that does not, and says so on its card. The verdict rule the page re-derives: recompute_first if any flag has severity error or an en-fail / draft-value flag is present; else budget_incomplete if any missing item is high; else report_as_is.

The report body

{
  "audience": "lab_report" | "paper" | "certificate" | "internal",
  "section_title": string,
  "paragraphs": [string],
  "budget_table": [ { "source": string, "value_text": string, "u_text": string, "unit": string, "type": "A" | "B",
                      "distribution": string, "dof_text": string, "sensitivity_text": string, "contribution_text": string, "share_pct": number } ],
  "result_statement": string,
  "checklist": [ { "id": "R-001", "item": string, "status": "met" | "not_met" | "n_a", "note": string } ],
  "caveats": [string]
}

result_statement equals facts.result.statement.expanded. The budget table has one row per input with u > 0 (a CODATA constant with an uncertainty included), its share_pct values within 0.6 of the engine's and summing to about 100 when there are no correlations. The checklist is exactly thirteen rows, R-001 to R-013, in the order the system prompt lists: measurand defined; model function stated; every input has a value and a unit; every input has a standard uncertainty; Type stated; distribution and degrees of freedom stated; correlations stated or independence asserted; combined standard uncertainty given; effective degrees of freedom given; coverage factor and probability given; expanded uncertainty in the result's unit or relative; result and uncertainty rounded to the same place with at most two significant figures in U; method of evaluation named.

The flags

The engine's flag rules, so a script can filter on them: dim-add, dim-func, dim-pow, div-zero, domain, name-unknown, unit-unknown, unit-mismatch, unit-mismatch-result, formula-syntax, no-formula, corr-range, corr-unknown (all severity error — no result); low-dof, k-assumed, nonlinear, mc-bias, asymmetric, rel-large, result-rel-large, few-observations, u-digits, value-digits, tag-unknown, en-fail, draft-k-missing, draft-mismatch, draft-digits, draft-place, draft-value, draft-no-u (severity warn); dominant, negligible, correlated, distribution-effect, exact-input, type-b-normal, constant-uncertain, constant-shadow, offset-unit, unused-input, en-ok (severity info).

8. Use it as a gate

A measurement sheet kept beside a report can be checked on every change: the engine alone (no credits) settles whether the sheet still computes and whether the statement in the report matches it; the budget lane (metered) says whether the budget is complete. The zero-cost half is the one worth wiring first — a result: line on the sheet holding the report's own statement makes the engine's draft-* flags a free regression test on the number you publish.

#!/bin/sh
# Free gate: the sheet must compute and the report's statement must match the engine.
# Put the report's statement on the sheet as `result: ...` and fail on any draft-* flag.
set -e
curl -sS https://plus-minus.skillsafe.ai/units.js -o units.js
curl -sS https://plus-minus.skillsafe.ai/unc.js -o unc.js
node -e '
const fs=require("fs"),vm=require("vm"); global.window={};
vm.runInThisContext(fs.readFileSync("units.js","utf8")); vm.runInThisContext(fs.readFileSync("unc.js","utf8"));
const a=window.PMUnc.analyze(fs.readFileSync(process.argv[1],"utf8"));
if(!a.ok){console.error("sheet does not compute:",a.errors.map(e=>e.message).join("; "));process.exit(1);}
const bad=a.flags.filter(f=>/^draft-|^en-fail$/.test(f.rule));
console.log(a.statement.expanded);
if(bad.length){console.error(bad.map(f=>f.id+" "+f.rule+": "+f.message).join("\n"));process.exit(2);}
' measurement.sheet

Truncation and partial results

When the balance sits between min_credits and hold_credits, the run is not refused: it executes with a reduced output cap and comes back with truncated: true on the finished job and on the streaming done event. What you hold then is a prefix of the reply. The envelope and coverage come first, so in the budget lane a truncated reply may carry every flag's status and the judged contributors while missing, improvements, statement and verdict — the fields a gate reads — are absent or cut mid-string. In the report lane the paragraphs are the longest strings and come early, so what gets lost is the budget table, the checklist and the caveats.

Check the flag before you treat a reply as complete. A truncated budget review with two judged inputs parses cleanly and looks like a short review. The page's answer is to close the JSON structure that arrived, render every section that parsed, and say "the stream ended early" above it; a script should do the same or retry. The right retry is a top-up or a smaller sheet with the attempt suffix on the Idempotency-Key incremented — never a repair that appends braces and calls the result a review.

One more honest limit: the sheet is clipped from the middle at 12,000 characters, whole lines at a time, with a comment line marking the cut, and both the engine and the model work on what remains. A budget over a clipped sheet is a budget of the inputs that survived; the statement the engine computes is still exact for those inputs, but it is not the measurement's. Keep one measurement per sheet — at a few dozen inputs a sheet is nowhere near the budget.