Drive Plus-Minus from your own code
Everything the web page does is available over HTTP: post one measurement sheet — the inputs with their values, units and uncertainties, and the formula that turns them into the result — together with the facts the free engine computed from it, and get back either a metrologist's budget review (every stated uncertainty judged, the contributions a competent evaluation would also carry as ready-to-paste sheet lines, the improvements ranked by share of variance, the correct statement and a verdict) or the uncertainty report section (the methods-and-results prose, the budget table and the thirteen-point reporting checklist a lab report, paper, calibration certificate or internal record needs).
The natural use is a bench script that keeps the uncertainty budget of a standing measurement under
review as the sheet changes, or a CI job that refuses to publish a number whose budget the engine
cannot compute or whose statement the reply re-rounded. One thing is different from most apps on
this platform, and it is the whole of step 4: the arithmetic is not the model's.
A free engine — the same units.js and unc.js the browser runs — does the
dimensional analysis, the GUM budget, Welch-Satterthwaite, the coverage factor, two Monte Carlo
propagations and the rounding, and hands the model a facts object it must quote rather
than recompute. An API caller has to produce that object too.
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{ "ok": true, "data": { ... } }
{ "ok": false, "error": { "code": "...", "message": "...", "details": { ... } } }
Send your token as Authorization: Bearer … on every call. The app slug travels in the
body of /guest as {"slug": "plus-minus"}; after that the token
itself carries the app, so a run needs only Authorization,
Content-Type: application/json and the Idempotency-Key described in
step 5.
The request body for /estimate, /run and /run-stream is the
input object itself — not wrapped in anything. A body of {"input": {…}} returns 200
and quietly hides every field from the model, so a run that looks fine comes back reviewing
nothing. Guard it the way the app's own client does: its mustBeObject() refuses to
send anything that is not a plain JSON object, because /estimate happily prices a bare
string, a number, null and [] alike.
Error codes
| code | status | what to do |
|---|---|---|
unauthorized | 401 | The token is missing, malformed or expired. Get a new one from the token page. |
payment_required | 402 | The balance is below min_credits, or a guest token tried a metered run. Call /estimate first, then sign in and top up. |
validation_error | 400 | The body is not a plain JSON object, or a field is the wrong type — every field of this app's input is a string, facts included. A facts object sent as an object rather than as a JSON string is the common one. |
not_found | 404 | Unknown job id, or the app slug does not exist. Check the id you polled with. |
rate_limited | 429 | Too many requests. Back off and retry; do not tight-loop a poll. |
internal | 500 | A server-side failure. Retry with the same Idempotency-Key so you are not billed twice. |
Replaying an Idempotency-Key with a changed body is not a retry and is
rejected rather than billed. When you change the input — an edited sheet, a different
context_hint, another audience — bump the attempt suffix on the key instead.
The field to get right first: task
Plus-Minus is one app with two lanes, and task is what chooses between them. It is the
first field of every request body:
task | what comes back | extra input fields |
|---|---|---|
"budget" | A review of the budget: measurand, contributors judging every stated uncertainty as reasonable, questionable or unsupported, missing contributions each with a ready-to-paste sheet_line, ranked improvements, the statement copied verbatim from the engine, and a verdict of report_as_is, budget_incomplete or recompute_first. | none |
"report" | The uncertainty section: section_title, three to six paragraphs of methods-and-results prose, a budget_table with one row per input that carries an uncertainty, the result_statement verbatim, a thirteen-row checklist (R-001 to R-013) and the caveats the report cannot claim its way out of. | audience, review_notes |
The system prompt routes on task and never blends the two contracts in one reply. A
missing or unrecognised task is not an error: the model answers the closest lane — an
audience or review_notes field means report — names the lane
it actually answered in the reply's lane field, and says so in
notes_on_input. So branch on lane in the reply, never on the
task you believe you sent. The app's own normaliser goes further: a reply with no
usable lane is classed as report when it carries
paragraphs, checklist or budget_table, and as
budget otherwise.
The two lanes chain, and the handoff is the point of the app. Run
budget over a sheet, act on what it says — paste the missing[].sheet_line
lines onto the sheet and recompute, or accept some of them as known gaps — then send the accepted
items as the review_notes of a report run over the same sheet. The report
lane turns them into the caveats paragraph nobody remembers to write. The web page
does this with one button on either side: Add to sheet for the proposed lines, and
Write the report from this review for the handoff.
The sheet grammar
sheet is the run's only evidence and it is a small, strict language. The engine parses
it line by line; the model never parses it at all, it reads the engine's facts. Get
the grammar right and everything downstream works; get it wrong and the engine emits an
error-level flag, facts.ok is false, and the reply tells you what to fix
on the sheet instead of reviewing a budget that does not exist.
Six kinds of line, and blank lines are ignored:
# a comment: everything from a # to the end of the line is stripped first
g = 4*pi^2*L/T^2 # a FORMULA line: a name = an expression
L = 0.9950 m ± 0.0005 m [B rect half] # an INPUT line: a name = a quantity
corr(L, T) = 0.4 # a CORRELATION line
level: 95% # a DIRECTIVE
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # a directive that takes a quantity
A line with no = at all, and no corr(...) or directive keyword, is an
error. The distinction between an input line and a formula line is made on the right-hand side: if
it parses as a quantity it is an input, otherwise it is an expression. Names are
[A-Za-z_][A-Za-z0-9_']* — letters, digits, underscore and the prime, starting with a
letter or an underscore — and a name may be defined only once, as an input or as a formula but
never as both.
Input lines: the quantity forms
Three ways to write a value with its uncertainty, plus one for a value without:
| form | example | what the engine reads |
|---|---|---|
name = value unit ± u unit | L = 0.9950 m ± 0.0005 m | The full form. The uncertainty's unit may differ from the value's as long as the dimensions match — m ± 0.5 mm is converted for you; a mismatch (m ± 0.5 s) is an error, not a warning. |
name = value unit ± u | T = 2.0012 s ± 0.0008 | With no unit after the ±, the uncertainty is in the value's unit. |
name = value unit ± p% | R = 100 ohm ± 0.1% | A relative uncertainty: u = |value| · p / 100. The percent sign must follow the number directly. |
name = value(uu) unit | m = 2.500(12) kg | The concise parenthesis notation: the digits in brackets are the uncertainty in the last places of the value, so this is 2.500 kg ± 0.012 kg. A value written this way is exempt from the digit-hygiene flags, because the notation fixes the digits by construction. |
name = value unit | n_swings = 20 | No ± at all: the input is exact, u = 0, and the engine raises an exact-input flag asking you to state an uncertainty unless it is a defined value or a count. |
+/- and +- are accepted as ASCII spellings of ±. A value may
be written in exponent form (6.626e-34), and the exponent counts when the engine
compares the decimal places of a value against those of its uncertainty.
Tags
A bracketed list at the end of an input line says how the uncertainty was evaluated. Tags are separated by spaces or commas, and the brackets must be the last thing on the line before any comment:
| tag | meaning |
|---|---|
A, typeA, type-A | Type A: a statistical evaluation of repeated observations. |
B, typeB, type-B | Type B: evaluated by any other means — a specification, a certificate, a resolution, judgement. |
n=10 | The number of repeated observations. It implies Type A, and it sets the degrees of freedom to n - 1 unless dof= says otherwise. |
dof=9, nu=9 | Degrees of freedom directly. Absent and without n=, the input carries infinite degrees of freedom. |
k=2 | The figure on the line is an expanded uncertainty at this coverage factor; the engine divides by k to get the standard uncertainty. This is how a calibration certificate's value goes on the sheet. |
rect, rectangular, uniform | Rectangular distribution. |
tri, triangular | Triangular distribution. |
U, u-shaped, arcsine | U-shaped (arcsine) distribution — the classic case being a quantity cycling between two limits. |
normal, gauss, gaussian | Normal distribution. This is the default when no distribution tag is present. |
half, halfwidth, half-width | The figure after the ± is a half-width, not a standard uncertainty: divide it by the distribution's factor. half on its own implies rect. |
The half-width divisors, applied only when half is present:
[rect half] u = a / sqrt(3) a 1 mm tape resolution: a = 0.5 mm, u = 0.29 mm
[tri half] u = a / sqrt(6)
[U half] u = a / sqrt(2)
[half] same as [rect half]
Order of operations inside one line: the half-width divisor first, then the division by
k= if both are present. A tag the engine does not know is not fatal — it raises a
tag-unknown warning flag and the line is otherwise read normally, which is worth
knowing because a typo such as [rect halfwidth=1] silently loses the half-width.
Formula lines
A formula line is name = expression. Several are allowed, and they may refer to each
other in the order written. The measurand — the quantity the whole budget is about
— is the last formula on the sheet unless a measurand: directive names another
one. The operators are + - * / ^ (with ** accepted for ^),
unary minus, and parentheses; ^ binds tighter than * and
/, which bind tighter than + and -. The functions are:
sqrt abs exp ln log log10 log2
sin cos tan asin acos atan # trigonometry works in RADIANS
atan2(y, x) pow(a, b) min(a, b) max(a, b)
Every formula is checked dimensionally as it is evaluated: adding metres to seconds, or taking the
logarithm of something with a dimension, is an error-level flag rather than a number. The
sensitivity coefficients are exact — the engine carries a gradient alongside every value
(forward-mode automatic differentiation), so c_i = ∂f/∂x_i is the analytic derivative
and not a finite difference. An input defined on the sheet and never used by any formula is
reported in facts.unused_inputs and as an unused-input flag.
Correlations
corr(L_1, L_2) = 0.85
One line per pair, with r in [-1, 1]. Both names must be inputs on the
sheet. A correlation adds its cross term 2·c_a·c_b·u_a·u_b·r to the variance and is
carried into the Monte Carlo sampling through a Cholesky factor. Two honest consequences the engine
flags rather than hides: Welch-Satterthwaite ignores correlation terms, so
veff is only approximate once you declare any; and the per-input
share_pct values are shares of the diagonal variance, so they no longer sum
to 100 % of the combined variance. The page's own check for the report lane's budget table is
therefore skipped when facts.correlations is non-empty.
Directives
| directive | example | effect |
|---|---|---|
title: | title: Pendulum, bench 3 | A label for the sheet. It travels in the facts only as part of the sheet text. |
measurand: | measurand: g | Names which formula is the measurand. Without it, the last formula line wins. Naming something that is not a formula is an error. |
unit: | unit: mm | Report the result in this unit instead of the coherent SI one the dimensions imply. It must match the result's dimension, and it may not be an offset unit — a result in kelvin is reported in kelvin, never in degC. |
k: | k: 2 | Fix the coverage factor instead of taking it from the t distribution. facts.result.k_source then reads "sheet", and if the t distribution at veff would have wanted a noticeably larger factor the engine raises a k-assumed warning: the stated interval covers less than it claims. |
level: | level: 95% | The coverage probability. Written as a percentage or as a fraction; the default is 95.45 %, the level at which k = 2 is exactly right for a normal distribution with infinite degrees of freedom. |
reference =, expected: | reference = 9.8062 m/s^2 ± 0.0001 m/s^2 | A value to compare against. The engine computes the normalised error E_n = (y - y_ref) / sqrt(U² + U_ref²) and raises en-ok or en-fail. A reference written with [k=2] carries an expanded uncertainty already; one written bare is a standard uncertainty and is expanded with the result's own k. This line is not an input to the model and has no sensitivity coefficient. |
result:, draft: | result: 9.81 m/s^2 ± 0.02 m/s^2 (k=2) | Your own draft statement, checked against the computed one: value agreement within U, a missing coverage factor, a ± that is more than a third out, too many significant figures, and a value and an uncertainty rounded to different places. Each is its own flag — draft-value, draft-k-missing, draft-mismatch, draft-digits, draft-place, draft-no-u — and draft-value is one of the two rules that force the verdict to recompute_first. |
Built-in constants
A name used in a formula and never defined on the sheet is looked up in the engine's constant
table. The exact ones (exact by the 2019 SI definitions) contribute nothing to the budget; the four
measured ones carry their CODATA 2018 standard uncertainty and join the budget as Type B
inputs, which is why they can appear in facts.inputs and in the report lane's
budget table. Defining an input with one of these names shadows the constant and raises a
constant-shadow flag:
| name | value | unit | note |
|---|---|---|---|
pi | 3.14159265358979 | — | pi (exact) |
e | 2.71828182845905 | — | Euler's number (exact) |
c | 299792458 | m/s | speed of light in vacuum (exact) |
h | 6.62607015e-34 | J s | Planck constant (exact) |
hbar | 1.054571817e-34 | J s | reduced Planck constant (exact) |
k_B | 1.380649e-23 | J/K | Boltzmann constant (exact) |
N_A | 6.02214076e23 | 1/mol | Avogadro constant (exact) |
q_e | 1.602176634e-19 | C | elementary charge (exact) |
g_n | 9.80665 | m/s^2 | standard acceleration of gravity (defined) |
R_gas | 8.314462618 | J/(mol K) | molar gas constant (exact) |
sigma_SB | 5.670374419e-8 | W/(m^2 K^4) | Stefan-Boltzmann constant (exact) |
G | 6.67430e-11 ± 0.00015e-11 | m^3/(kg s^2) | Newtonian constant of gravitation — carries an uncertainty |
eps_0 | 8.8541878128e-12 ± 1.3e-21 | F/m | vacuum electric permittivity — carries an uncertainty |
mu_0 | 1.25663706212e-6 ± 1.9e-16 | N/A^2 | vacuum magnetic permeability — carries an uncertainty |
m_e | 9.1093837015e-31 ± 2.8e-40 | kg | electron mass — carries an uncertainty |
m_p | 1.67262192369e-27 ± 5.1e-37 | kg | proton mass — carries an uncertainty |
u_amu | 1.66053906660e-27 ± 5.0e-37 | kg | atomic mass constant — carries an uncertainty |
In facts.inputs a constant is marked "constant": true with a
note, and an exact one also carries "exact": true. The budget lane judges
only the inputs that are actually on the sheet — a contributors entry naming
pi is a warning on the page, not a virtue.
Units
Units are parsed into the seven SI base dimensions, so any spelling that resolves to the right
dimension works and mixed systems convert silently:
m kg s A K mol cd rad sr Hz N Pa J W C V F ohm S Wb T H lm lx Bq Gy Sv kat L, the time
units min h d yr, the angles deg arcmin arcsec (and °), the
energy and mass units eV cal Wh Ah t u Da, the pressures
bar atm mmHg torr psi, the imperial set
ft in yd mi nmi lb oz mph kn ha acre gal, the ratios
% permille ppm ppb, and the offset temperatures degC °C degF °F. Write
compound units with *, / and ^: m/s^2,
J/(mol K), kg m^2/s^3. SI prefixes apply wherever the bare symbol is not
itself a unit — mm, kPa, µV, MHz.
Two traps worth naming. An offset unit is read as an absolute temperature:
T = 20 degC is 293.15 K, and the engine raises an offset-unit flag
because a temperature difference of 20 degrees should be written 20 K.
And u is the dalton, not micro-anything, while T is the tesla — inputs
named after unit symbols are fine (the example sheet's T is a period in seconds),
because names and units are read in different positions on the line.
Clipping
The browser sends at most 12,000 characters of sheet, clipping whole lines from the middle and leaving a marker in place of the cut:
# [... 42 lines (3120 characters) cut from the middle ...]
Clip the same way if you send more, and keep the marker: the prompt keys on it, refuses to claim
anything about the missing middle, and says so in notes_on_input. The counts also
travel in the facts as facts.clipped = {"cut": 3120, "lines_cut": 42}. In practice a
sheet that long is several measurements stacked in one file, and the right move is to split it: a
budget is about one measurand.
When the engine cannot compute
Parse and dimension failures do not stop the run. They land in facts.errors[], they
become flags with "severity": "error", and facts.ok is false
— there is then no facts.result and no facts.mc. The prompt's rule for
that case is explicit: say what to fix on the sheet and keep the rest short. The web page refuses
to spend a run at all in this state; an API caller can still spend one, and the honest thing is to
check facts.ok first and save the credits. The same check is what stops a CI job from
reporting "the review found nothing serious" about a sheet that never computed.
1. Get a token
The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. Nothing on that page needs a developer tool — it reads the same storage the app itself uses and prints the token for you.
A guest token can call /me and /estimate. Both lanes are
metered, so reviewing a budget and writing a report section each need a personal
token from signing in. Nothing about the engine is metered: the arithmetic in step 4 costs nothing
and needs no token at all.
# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
# https://plus-minus.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="aut_YOUR_TOKEN"
#
# To mint a guest token from the command line instead. A guest token is enough for
# /me and /estimate; a budget or report run needs a personal token.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" \
-d '{"slug": "plus-minus"}'
# {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}
# Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot run a metered lane.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest",
data=json.dumps({"slug": "plus-minus"}).encode(),
method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
print(TOKEN)
// Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "plus-minus" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest",
bytes.NewReader([]byte(`{"slug": "plus-minus"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct {
Data struct {
Token string `json:"token"`
} `json:"data"`
}
_ = json.NewDecoder(guestRes.Body).Decode(&guest)
fmt.Println(guest.Data.Token)
// Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
var http = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\": \"plus-minus\"}"))
.build();
HttpResponse<String> guest = http.send(guestReq, HttpResponse.BodyHandlers.ofString());
System.out.println(guest.body()); // {"ok":true,"data":{"token":"aut_...","subject_type":"guest"}}
# Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot run a metered lane.
require "json"
require "net/http"
require "uri"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
req = Net::HTTP::Post.new(uri)
req["Content-Type"] = "application/json"
req.body = JSON.generate({ "slug" => "plus-minus" })
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_HTTPHEADER, ["Content-Type: application/json"]);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode(["slug" => "plus-minus"]));
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
$TOKEN = $payload["data"]["token"];
// Open https://plus-minus.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot run a metered lane.
using System.Net.Http.Json;
using System.Text.Json;
var http = new HttpClient();
var guestRes = await http.PostAsJsonAsync(
"https://api.skillsafe.ai/v1/app-api/guest",
new { slug = "plus-minus" });
var guest = await guestRes.Content.ReadFromJsonAsync<JsonElement>();
var token = guest.GetProperty("data").GetProperty("token").GetString();
2. A tiny client
One helper that adds the headers, unwraps data and raises on error.
# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="plus-minus"
TOKEN="$SKILLSAFE_TOKEN" # from https://plus-minus.skillsafe.ai/tokens.html
call() { # call <path> [json-body]
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
fi
}
import json, os, urllib.error, urllib.request
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "plus-minus"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN") # from https://plus-minus.skillsafe.ai/tokens.html
def call(path, body=None, headers=None):
"""Returns the unwrapped `data`, or raises with the API error code."""
if body is not None and not isinstance(body, dict):
raise TypeError("the request body must be a JSON object, not a bare string")
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(f"{BASE}/{path}", data=data,
method="POST" if body is not None else "GET")
req.add_header("Authorization", f"Bearer {TOKEN}")
if body is not None:
req.add_header("Content-Type", "application/json")
for k, v in (headers or {}).items():
req.add_header(k, v)
try:
with urllib.request.urlopen(req) as r:
payload = json.load(r)
except urllib.error.HTTPError as e:
payload = json.load(e)
if not payload.get("ok"):
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code')}: {err.get('message')}")
return payload["data"]
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "plus-minus";
const TOKEN = "aut_YOUR_TOKEN"; // from https://plus-minus.skillsafe.ai/tokens.html
async function call(path, body, extraHeaders) {
if (body !== undefined && (body === null || typeof body !== "object" || Array.isArray(body))) {
throw new TypeError("the request body must be a JSON object, not a bare string");
}
const res = await fetch(`${BASE}/${path}`, {
method: body ? "POST" : "GET",
headers: {
Authorization: `Bearer ${TOKEN}`,
...(body ? { "Content-Type": "application/json" } : {}),
...(extraHeaders || {}),
},
body: body ? JSON.stringify(body) : undefined,
});
const payload = await res.json();
if (!payload.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
return payload.data;
}
package main
import (
"bufio"
"bytes"
"crypto/sha256"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"os/exec"
"strings"
"time"
)
const (
base = "https://api.skillsafe.ai/v1/app-api"
slug = "plus-minus"
)
var token = os.Getenv("SKILLSAFE_TOKEN") // from https://plus-minus.skillsafe.ai/tokens.html
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(path string, body any, hdr map[string]string) (json.RawMessage, error) {
method := http.MethodGet
var rdr io.Reader
if body != nil {
method = http.MethodPost
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
for k, v := range hdr {
req.Header.Set(k, v)
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
import java.net.URI;
import java.net.http.*;
import java.util.Map;
public class PlusMinus {
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final String SLUG = "plus-minus";
static final String TOKEN = System.getenv().getOrDefault("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String call(String path, String jsonBody, Map<String, String> extra) throws Exception {
if (jsonBody != null && !jsonBody.trim().startsWith("{")) {
throw new IllegalArgumentException("the request body must be a JSON object");
}
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + TOKEN);
if (jsonBody != null) {
b.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(jsonBody));
} else {
b.GET();
}
if (extra != null) extra.forEach(b::header);
HttpResponse<String> res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
// The envelope is always {"ok":true,"data":...} or {"ok":false,"error":...}.
return res.body();
}
}
require "json"
require "net/http"
require "uri"
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "plus-minus"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "aut_YOUR_TOKEN") # from https://plus-minus.skillsafe.ai/tokens.html
def call(path, body = nil, extra = {})
raise TypeError, "the request body must be a JSON object" if body && !body.is_a?(Hash)
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
if body
req["Content-Type"] = "application/json"
req.body = JSON.generate(body)
end
extra.each { |k, v| req[k] = v }
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
raise "#{payload['error']['code']}: #{payload['error']['message']}" unless payload["ok"]
payload["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "plus-minus";
define("TOKEN", getenv("SKILLSAFE_TOKEN") ?: "aut_YOUR_TOKEN"); // from /tokens.html
function call(string $path, ?array $body = null, array $extra = []) {
$ch = curl_init(BASE . "/" . $path);
$headers = array_merge(["Authorization: Bearer " . TOKEN], $extra);
if ($body !== null) {
$headers[] = "Content-Type: application/json";
curl_setopt($ch, CURLOPT_POST, true);
// The body must encode as an object: an empty PHP array would encode
// as [] and be rejected as a validation_error.
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body, JSON_UNESCAPED_SLASHES));
}
curl_setopt($ch, CURLOPT_HTTPHEADER, $headers);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
if (empty($payload["ok"])) {
throw new RuntimeException($payload["error"]["code"] . ": " . $payload["error"]["message"]);
}
return $payload["data"];
}
using System.Net.Http.Json;
using System.Text.Json;
static class PlusMinus
{
const string Base = "https://api.skillsafe.ai/v1/app-api";
const string Slug = "plus-minus";
static readonly string Token =
Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "aut_YOUR_TOKEN";
static readonly HttpClient Http = new();
public static async Task<JsonElement> Call(string path, object? body = null,
(string, string)? extraHeader = null)
{
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post,
$"{Base}/{path}");
req.Headers.Add("Authorization", $"Bearer {Token}");
if (extraHeader is { } h) req.Headers.Add(h.Item1, h.Item2);
if (body is not null) req.Content = JsonContent.Create(body);
var res = await Http.SendAsync(req);
var payload = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!payload.GetProperty("ok").GetBoolean())
{
var e = payload.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return payload.GetProperty("data");
}
}
3. Check the session and the balance
GET /me tells you whether the token is a guest or a person, and what the balance is.
The object is small and carries exactly three things: subject_type —
guest or user — subject_id, and credits, the
wallet balance. There is no username in it, so "signed in" is
subject_type === "user" and nothing else; a guest can price a run but cannot
start one. Compare credits against min_credits from the next step before
you run, so a shortfall surfaces as your own clear message rather than a 402.
call me
# {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}
# A guest token answers the same call with subject_type "guest" and cannot run
# either lane. Gate on it before you spend a poll loop finding out:
call me | grep -q '"subject_type":"user"' || {
echo "sign in at https://plus-minus.skillsafe.ai/tokens.html first" >&2
exit 1
}
me = call("me")
signed_in = me["subject_type"] == "user"
print(me["subject_type"], me.get("credits"), "signed in" if signed_in else "guest")
if not signed_in:
raise SystemExit("a guest token cannot run the budget or the report lane")
const me = await call("me");
const signedIn = me.subject_type === "user";
console.log(me.subject_type, me.credits, signedIn ? "signed in" : "guest");
if (!signedIn) throw new Error("a guest token cannot run the budget or the report lane");
raw, err := call("me", nil, nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
SubjectID string `json:"subject_id"`
Credits int `json:"credits"`
}
_ = json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits)
if me.SubjectType != "user" {
panic("a guest token cannot run the budget or the report lane")
}
System.out.println(PlusMinus.call("me", null, null));
// {"ok":true,"data":{"subject_type":"user","subject_id":"usr_...","credits":51234}}
// "signed in" is subject_type.equals("user") - there is no username field, and a
// guest token is refused by both lanes.
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
abort "sign in first - a guest token cannot run a lane" unless me["subject_type"] == "user"
<?php
$me = call("me");
echo $me["subject_type"], " ", $me["credits"], PHP_EOL;
$signedIn = $me["subject_type"] === "user";
if (!$signedIn) {
throw new RuntimeException("a guest token cannot run the budget or the report lane");
}
var me = await PlusMinus.Call("me");
var signedIn = me.GetProperty("subject_type").GetString() == "user";
Console.WriteLine($"{me.GetProperty("subject_type")} {me.GetProperty("credits")} {signedIn}");
if (!signedIn) throw new Exception("a guest token cannot run the budget or the report lane");
4. Price the run — free
The input object is exactly what the app's own form submits. It is always a JSON
object — never a bare string, never wrapped in an input key — and
every one of its fields is a string, because the transport takes scalars only.
That is why facts is a JSON string rather than a nested object:
| field | type | meaning |
|---|---|---|
task | string, required | "budget" or "report". The lane. It routes the prompt, and a missing or unknown value degrades to the closest lane rather than failing — the reply names what it answered in lane and in notes_on_input. |
sheet | string, required | The measurement sheet, in the grammar above. Clipped from the middle at 12,000 characters with the marker kept. This is the text the reply quotes lines from, and the only thing the model knows about the measurement besides context_hint. |
context_hint | string, optional | Up to 600 characters about the instrument, the setting and the purpose — "First-year teaching lab. A steel bob on a string, a metre tape, a hand-held stopwatch…". It is what separates a plausible criticism from a guess: without it the model knows the numbers and nothing about how they were obtained. It also sets the language of the prose — write the hint in German and the review comes back in German, with the keys and enum values still in English. |
facts | string, always present | The engine's output, JSON.stringify-ed. Everything numeric in the reply is quoted from here. Shape and the script that computes it are below. |
| report lane only | ||
audience | enum | lab_report (a student or teaching lab), paper (a journal methods section), certificate (a calibration or test certificate, ISO/IEC 17025 7.8) or internal (an internal record). It changes the register, the section title and whether traceability and the standards followed get their own paragraph. Anything else falls back to lab_report. |
review_notes | string, optional | Up to 2,000 characters: the items you accepted from a budget review, as text. This is the handoff — the report lane turns them into caveats, which is the paragraph that says what the report cannot claim. Omit it and the caveats are derived from the sheet and the flags alone. |
| both lanes, rarely needed | ||
retry_note | string, optional | What the app sends on its one automatic retry when a reply did not parse: a sentence naming the failure and restating the contract. If you build the same recovery, send the same kind of note rather than resending the identical body — and remember to bump the attempt suffix on the key. |
facts: the numbers are the engine's, not the model's
In the browser this object is computed for free, before the run, by units.js and
unc.js — the same two files the page loads. The model is told, in the system prompt,
never to recompute, re-round or restate any of it with different digits, and to quote
facts.result.statement.expanded and .concise verbatim wherever the result
is stated. Here is the object for the running example, complete but abbreviated in the long arrays:
{
"ok": true,
"measurand": "g",
"formulas": [ { "name": "g", "expr": "4*pi^2*L/T^2", "unit": "m/s^2", "ok": true } ],
"inputs": [
{ "name": "L", "value": 0.995, "u": 0.000288675, "unit": "m", "type": "B",
"distribution": "rectangular", "dof": "infinite", "rel_pct": 0.0290126,
"sensitivity": 9.85777, "contribution": 0.00284569, "share_pct": 11.6357 },
{ "name": "T", "value": 2.0012, "u": 0.0008, "unit": "s", "type": "A",
"distribution": "normal", "dof": 9, "rel_pct": 0.039976, "n": 10,
"sensitivity": -9.8026, "contribution": 0.00784208, "share_pct": 88.3643 },
{ "name": "pi", "value": 3.14159, "u": 0, "unit": "", "type": "B",
"distribution": "normal", "dof": "infinite", "rel_pct": 0, "constant": true,
"note": "pi", "exact": true, "sensitivity": 6.24427, "contribution": 0, "share_pct": 0 }
],
"correlations": [],
"errors": [],
"unused_inputs": [],
"flags": [
{ "id": "E-001", "rule": "dominant", "severity": "info",
"message": "T carries 88 % of the variance of g - improving anything else changes almost nothing",
"ref": "T" },
{ "id": "E-002", "rule": "low-dof", "severity": "warn",
"message": "effective degrees of freedom 11.5 (Welch-Satterthwaite) - the coverage factor for 95.45 % is 2.24, not 2",
"ref": "" },
{ "id": "E-003", "rule": "en-ok", "severity": "info",
"message": "E_n = 0.12 against the reference 9.80620 m/s^2 - consistent with the reference (|E_n| <= 1)",
"ref": "" }
],
"clipped": { "cut": 0, "lines_cut": 0 },
"result": {
"name": "g", "value": 9.80848, "u": 0.00834243, "U": 0.0186871,
"k": 2.24, "k_source": "t", "k_from_t": 2.24193, "level_pct": 95.45,
"veff": 11.5263, "unit": "m/s^2", "unit_si": "m/s^2", "rel_pct": 0.0850533,
"statement": {
"expanded": "g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)",
"standard": "g = 9.8085 m/s^2, u = 0.0083 m/s^2 (standard uncertainty)",
"concise": "g = 9.8085(83) m/s^2",
"value_text": "9.808", "U_text": "0.019", "u_text": "0.0083",
"k_text": "2.24", "level_text": "95.45 %"
}
},
"mc": {
"normal": { "n": 60000, "mean": 9.80846, "sd": 0.00834484, "lo": 9.79175, "hi": 9.8251,
"skew": -0.00740069, "level_pct": 95.45,
"what": "every input normal - isolates non-linearity" },
"stated": { "n": 60000, "mean": 9.80848, "sd": 0.00932407, "lo": 9.78944, "hi": 9.82773,
"skew": 0.0189543, "level_pct": 95.45,
"what": "the stated distributions - the coverage interval to report" }
},
"en": { "value": 0.122144, "reference": 9.8062, "reference_U": 0.000224 }
}
Field by field, the parts a caller actually reads back: inputs[] is the budget — value,
standard uncertainty u, type (A or B),
distribution, dof (a number or the string "infinite"),
the sensitivity coefficient, the contribution |c_i|·u_i and its
share_pct of the variance. result carries the combined standard
uncertainty u, the coverage factor k with k_source
("t" from the t distribution at veff, or "sheet" when a
k: directive fixed it), the expanded uncertainty U, and the three rounded
statements. mc.normal re-runs the propagation with every input normal, so its
sd against result.u isolates non-linearity; mc.stated uses
the distributions you tagged, and its lo/hi is the coverage interval to
report when the flags say a symmetric ± misleads. flags[] is the
reconciliation list of step 7.
Computing it outside the browser. The engine is two plain scripts served by this
host and they run unmodified in Node: give them a window to attach to, evaluate them,
and call analyze() then toFacts(). Twenty lines, no dependencies, no
token, no charge. Every sample below is the same three steps — fetch /units.js and
/unc.js, evaluate them with global.window = {} in scope, then
JSON.stringify(window.PMUnc.toFacts(window.PMUnc.analyze(sheet), {cut: 0, lines_cut: 0}))
— and the non-JavaScript ones drive that script as a subprocess, which is both the shortest and the
only way to be certain your numbers are the ones the app would have produced.
# The sheet, saved as sheet.txt - this is the worked example used everywhere below.
cat > sheet.txt <<'SHEET'
# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office
SHEET
# The engine, fetched once and kept next to your script. These are the same two
# files the web page loads; re-fetch them when the app is released again.
curl -sS -o units.js "https://plus-minus.skillsafe.ai/units.js"
curl -sS -o unc.js "https://plus-minus.skillsafe.ai/unc.js"
# facts.js - 12 lines, no dependencies, no token, no charge.
cat > facts.js <<'JS'
const fs = require("fs"), vm = require("vm");
global.window = {}; // the two files attach to it
for (const f of ["units.js", "unc.js"]) vm.runInThisContext(fs.readFileSync(f, "utf8"));
const U = window.PMUnc;
const sheet = fs.readFileSync(process.argv[2], "utf8").replace(/\r\n?/g, "\n");
const clip = U.clipMiddle(sheet, U.MAX_SHEET); // 12,000 chars, whole lines
const facts = U.toFacts(U.analyze(clip.text), { cut: clip.cut, lines_cut: clip.lines_cut });
if (!facts.ok) console.error("engine errors:", JSON.stringify(facts.errors));
process.stdout.write(JSON.stringify(facts)); // ONE line: the `facts` string
JS
FACTS=$(node facts.js sheet.txt)
echo "$FACTS" | head -c 120
# {"ok":true,"measurand":"g","formulas":[{"name":"g","expr":"4*pi^2*L/T^2","unit":"m/s^2",...
# The engine is JavaScript, so the honest way to reproduce the app's numbers is to
# run the app's own engine. facts.js is the 12-line script shown on the cURL tab;
# fetch units.js and unc.js next to it once.
import json, subprocess, pathlib, urllib.request
HOST = "https://plus-minus.skillsafe.ai"
for name in ("units.js", "unc.js"):
if not pathlib.Path(name).exists():
urllib.request.urlretrieve(f"{HOST}/{name}", name)
SHEET = """# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office
"""
pathlib.Path("sheet.txt").write_text(SHEET, encoding="utf-8")
FACTS = subprocess.run(["node", "facts.js", "sheet.txt"],
capture_output=True, check=True).stdout.decode("utf-8")
facts = json.loads(FACTS) # parse it only to inspect it
if not facts["ok"]:
raise SystemExit(f"the engine could not compute a result: {facts['errors']}")
print(facts["result"]["statement"]["expanded"])
# g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
# FACTS - the STRING, not the parsed object - is what goes in the request body.
// In Node the engine needs no subprocess at all: give the two files a `window`
// to attach to and evaluate them. This is facts.mjs.
import fs from "node:fs";
import vm from "node:vm";
const HOST = "https://plus-minus.skillsafe.ai";
async function loadEngine() {
globalThis.window = {}; // units.js and unc.js attach here
for (const name of ["/units.js", "/unc.js"]) {
const src = await fetch(HOST + name).then((r) => r.text());
vm.runInThisContext(src); // defines window.PMUnits, window.PMUnc
}
return globalThis.window.PMUnc;
}
const SHEET = [
"# g from a simple pendulum",
"g = 4*pi^2*L/T^2",
"L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm",
"T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings",
"reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office",
].join("\n");
const U = await loadEngine();
const clip = U.clipMiddle(SHEET, U.MAX_SHEET); // 12,000 chars, whole lines
const analysis = U.analyze(clip.text);
const FACTS = JSON.stringify(
U.toFacts(analysis, { cut: clip.cut, lines_cut: clip.lines_cut })
);
if (!analysis.ok) throw new Error("the engine could not compute a result");
console.log(analysis.statement.expanded);
// g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
// FACTS is a STRING and travels as one; fs.writeFileSync("facts.json", FACTS) if
// you would rather keep the engine step in its own job.
// Go drives the same facts.js as a subprocess. Keep units.js, unc.js and facts.js
// beside the binary; they are static files from https://plus-minus.skillsafe.ai.
const sheet = `# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office
`
func engineFacts(sheet string) (string, error) {
if err := os.WriteFile("sheet.txt", []byte(sheet), 0o644); err != nil {
return "", err
}
out, err := exec.Command("node", "facts.js", "sheet.txt").Output()
if err != nil {
return "", fmt.Errorf("the engine failed: %w", err)
}
return string(out), nil
}
facts, err := engineFacts(sheet)
if err != nil {
panic(err)
}
// Parse it only to check it; the request carries the string unchanged.
var probe struct {
OK bool `json:"ok"`
Result struct {
Statement struct {
Expanded string `json:"expanded"`
} `json:"statement"`
} `json:"result"`
}
_ = json.Unmarshal([]byte(facts), &probe)
if !probe.OK {
panic("the engine could not compute a result - fix the sheet before spending a run")
}
fmt.Println(probe.Result.Statement.Expanded)
// g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
// Java drives the same facts.js as a subprocess. units.js, unc.js and facts.js are
// static files from https://plus-minus.skillsafe.ai, fetched once.
import java.nio.charset.StandardCharsets;
import java.nio.file.*;
static final String SHEET = String.join("\n",
"# g from a simple pendulum",
"g = 4*pi^2*L/T^2",
"L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm",
"T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings",
"reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office");
static String engineFacts(String sheet) throws Exception {
Files.write(Path.of("sheet.txt"), sheet.getBytes(StandardCharsets.UTF_8));
Process p = new ProcessBuilder("node", "facts.js", "sheet.txt")
.redirectErrorStream(false).start();
String facts = new String(p.getInputStream().readAllBytes(), StandardCharsets.UTF_8);
if (p.waitFor() != 0) throw new RuntimeException("the engine failed");
return facts;
}
String FACTS = engineFacts(SHEET);
if (!FACTS.contains("\"ok\":true")) {
throw new RuntimeException("the engine could not compute a result - fix the sheet first");
}
// FACTS is one JSON string and goes into the body as a string VALUE, escaped once.
# Ruby drives the same facts.js as a subprocess. units.js, unc.js and facts.js are
# static files from https://plus-minus.skillsafe.ai, fetched once.
require "json"
require "open3"
SHEET = <<~SHEET
# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office
SHEET
def engine_facts(sheet)
File.write("sheet.txt", sheet)
out, err, status = Open3.capture3("node", "facts.js", "sheet.txt")
raise "the engine failed: #{err}" unless status.success?
out
end
FACTS = engine_facts(SHEET)
facts = JSON.parse(FACTS)
abort "the engine could not compute a result: #{facts['errors']}" unless facts["ok"]
puts facts["result"]["statement"]["expanded"]
# g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
<?php
// PHP drives the same facts.js as a subprocess. units.js, unc.js and facts.js are
// static files from https://plus-minus.skillsafe.ai, fetched once.
$SHEET = "# g from a simple pendulum\n"
. "g = 4*pi^2*L/T^2\n"
. "L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm\n"
. "T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings\n"
. "reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office\n";
function engine_facts(string $sheet): string {
file_put_contents("sheet.txt", $sheet);
$out = [];
$code = 0;
exec("node facts.js sheet.txt", $out, $code);
if ($code !== 0) {
throw new RuntimeException("the engine failed");
}
return implode("", $out);
}
$FACTS = engine_facts($SHEET);
$facts = json_decode($FACTS, true);
if (empty($facts["ok"])) {
throw new RuntimeException("the engine could not compute a result - fix the sheet first");
}
echo $facts["result"]["statement"]["expanded"], PHP_EOL;
// g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
// C# drives the same facts.js as a subprocess. units.js, unc.js and facts.js are
// static files from https://plus-minus.skillsafe.ai, fetched once.
using System.Diagnostics;
using System.Text.Json;
const string Sheet = """
# g from a simple pendulum
g = 4*pi^2*L/T^2
L = 0.9950 m ± 0.0005 m [B rect half] # tape resolution 1 mm
T = 2.0012 s ± 0.0008 s [A n=10] # mean of 10 timings of 20 swings
reference = 9.8062 m/s^2 ± 0.0001 m/s^2 # local g from the survey office
""";
static string EngineFacts(string sheet)
{
File.WriteAllText("sheet.txt", sheet);
var psi = new ProcessStartInfo("node", "facts.js sheet.txt")
{
RedirectStandardOutput = true,
};
using var proc = Process.Start(psi)!;
var facts = proc.StandardOutput.ReadToEnd();
proc.WaitForExit();
if (proc.ExitCode != 0) throw new Exception("the engine failed");
return facts;
}
var FACTS = EngineFacts(Sheet);
var probe = JsonDocument.Parse(FACTS).RootElement;
if (!probe.GetProperty("ok").GetBoolean())
throw new Exception("the engine could not compute a result - fix the sheet first");
Console.WriteLine(probe.GetProperty("result").GetProperty("statement")
.GetProperty("expanded").GetString());
// g = 9.808 m/s^2 ± 0.019 m/s^2 (k = 2.24, 95.45 %)
The lazy alternative, and what it costs you. The field is a string, so
"facts": "{}" is accepted and the run proceeds. What you get back is a review with no
numbers in it: nothing to quote, no statement to copy, no share_pct to
rank improvements by, no flags to reconcile — and, worse, a model asked to comment on a
measurement whose uncertainty nobody computed will write plausible prose anyway. The app itself
never does this: every run from the page carries a real facts object, and the page then checks the
reply against it. If you cannot run the engine, at least know that you are buying prose, not a
budget.
/estimate creates no job and charges nothing. It returns the model
binding — model is gpt-5.6-terra, model_alias is
gpt-terra, and markup_bps — plus sponsor_enabled and the
reservation: hold_credits is what gets held, and min_credits is the
balance you must clear to start at all. The hold is a reservation, not the price. It
prices the full output cap, so the charged_credits on the settled job is usually far
lower. Budget against hold_credits, report against charged_credits.
Price each lane separately — the app throws away the budget lane's estimate the moment you switch
lanes, for exactly this reason. A report body carries review_notes and prices
differently from a budget body over the same sheet, and facts is usually the largest
field in both: a sheet with twenty inputs is twenty budget rows of input tokens whether you read
them or not.
# Lane A - the budget review. Build the body with a JSON encoder, never with
# string concatenation: the sheet has newlines and the facts string has quotes.
HINT="First-year teaching lab. A steel bob on a string, a metre tape, a hand-held stopwatch; the period is the mean of ten timings of twenty swings each."
INPUT=$(python3 - "$FACTS" "$HINT" <<'PY'
import json, sys
print(json.dumps({
"task": "budget",
"sheet": open("sheet.txt").read(),
"context_hint": sys.argv[2],
"facts": sys.argv[1], # the STRING from node facts.js
}))
PY
)
call estimate "$INPUT"
# {"ok":true,"data":{"model":"gpt-5.6-terra","model_alias":"gpt-terra",
# "markup_bps":1000,"hold_credits":2652,"min_credits":310,"sponsor_enabled":false}}
#
# estimate is FREE. It creates no job and charges nothing. hold_credits is what
# gets RESERVED; charged_credits on the settled job is normally much lower.
# Lane B - the report section over the SAME sheet and the SAME facts, carrying the
# items accepted from the review. This is the handoff.
NOTES="Accepted from the budget review: reaction time of the stopwatch operator (Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the finite-amplitude correction were not evaluated; both are noted as caveats rather than added to the budget."
REPORT_INPUT=$(python3 - "$FACTS" "$HINT" "$NOTES" <<'PY'
import json, sys
print(json.dumps({
"task": "report",
"sheet": open("sheet.txt").read(),
"context_hint": sys.argv[2],
"audience": "lab_report", # lab_report | paper | certificate | internal
"review_notes": sys.argv[3],
"facts": sys.argv[1],
}))
PY
)
call estimate "$REPORT_INPUT" # price each lane separately
HINT = ("First-year teaching lab. A steel bob on a string, a metre tape, a hand-held "
"stopwatch; the period is the mean of ten timings of twenty swings each.")
NOTES = ("Accepted from the budget review: reaction time of the stopwatch operator "
"(Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the "
"finite-amplitude correction were not evaluated; both are noted as caveats "
"rather than added to the budget.")
# Lane A - the budget review. Every value is a string; FACTS is the string from
# node facts.js, NOT the parsed object.
INPUT = {
"task": "budget",
"sheet": SHEET,
"context_hint": HINT,
"facts": FACTS,
}
est = call("estimate", INPUT)
print(est["hold_credits"], est["min_credits"], est["model"])
# 2652 310 gpt-5.6-terra - estimate is free and creates no job
# Lane B - the report section over the same sheet and the same facts.
REPORT_INPUT = {
"task": "report",
"sheet": SHEET,
"context_hint": HINT,
"audience": "lab_report", # lab_report | paper | certificate | internal
"review_notes": NOTES,
"facts": FACTS,
}
print(call("estimate", REPORT_INPUT)["hold_credits"]) # price each lane separately
const HINT =
"First-year teaching lab. A steel bob on a string, a metre tape, a hand-held " +
"stopwatch; the period is the mean of ten timings of twenty swings each.";
const NOTES =
"Accepted from the budget review: reaction time of the stopwatch operator " +
"(Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the " +
"finite-amplitude correction were not evaluated; both are noted as caveats " +
"rather than added to the budget.";
// Lane A - the budget review. FACTS is the STRING from the engine step.
const INPUT = {
task: "budget",
sheet: SHEET,
context_hint: HINT,
facts: FACTS,
};
const est = await call("estimate", INPUT);
console.log(est.hold_credits, est.min_credits, est.model);
// 2652 310 gpt-5.6-terra - estimate is free and creates no job
// Lane B - the report section over the same sheet and the same facts.
const REPORT_INPUT = {
task: "report",
sheet: SHEET,
context_hint: HINT,
audience: "lab_report", // lab_report | paper | certificate | internal
review_notes: NOTES,
facts: FACTS,
};
console.log((await call("estimate", REPORT_INPUT)).hold_credits); // price each lane
const hint = "First-year teaching lab. A steel bob on a string, a metre tape, a hand-held " +
"stopwatch; the period is the mean of ten timings of twenty swings each."
const notes = "Accepted from the budget review: reaction time of the stopwatch operator " +
"(Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the " +
"finite-amplitude correction were not evaluated; both are noted as caveats " +
"rather than added to the budget."
// Lane A - the budget review. Every value is a string; facts is the engine's
// JSON as a STRING, marshalled into the body like any other string.
input := map[string]any{
"task": "budget",
"sheet": sheet,
"context_hint": hint,
"facts": facts,
}
raw, err := call("estimate", input, nil)
if err != nil {
panic(err)
}
var est struct {
Model string `json:"model"`
HoldCredits int `json:"hold_credits"`
MinCredits int `json:"min_credits"`
}
_ = json.Unmarshal(raw, &est)
fmt.Println(est.Model, est.HoldCredits, est.MinCredits)
// gpt-5.6-terra 2652 310 - estimate is free and creates no job
// Lane B - the report section over the same sheet and the same facts.
reportInput := map[string]any{
"task": "report",
"sheet": sheet,
"context_hint": hint,
"audience": "lab_report", // lab_report | paper | certificate | internal
"review_notes": notes,
"facts": facts,
}
_, _ = call("estimate", reportInput, nil) // price each lane separately
// Build the body with a JSON library in real code; escaped by hand here so the
// shape is visible. FACTS is one long JSON string and is escaped ONCE, as a value.
static String jsonString(String s) { // minimal escaper for the sample
return "\"" + s.replace("\\", "\\\\").replace("\"", "\\\"").replace("\n", "\\n") + "\"";
}
static final String HINT = "First-year teaching lab. A steel bob on a string, a metre tape, "
+ "a hand-held stopwatch; the period is the mean of ten timings of twenty swings each.";
static final String NOTES = "Accepted from the budget review: reaction time of the stopwatch "
+ "operator (Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the "
+ "finite-amplitude correction were not evaluated; both are noted as caveats rather than "
+ "added to the budget.";
String INPUT = "{"
+ "\"task\":" + jsonString("budget") + ","
+ "\"sheet\":" + jsonString(SHEET) + ","
+ "\"context_hint\":" + jsonString(HINT) + ","
+ "\"facts\":" + jsonString(FACTS)
+ "}";
System.out.println(PlusMinus.call("estimate", INPUT, null));
// {"ok":true,"data":{"model":"gpt-5.6-terra","hold_credits":2652,"min_credits":310,...}}
String REPORT_INPUT = "{"
+ "\"task\":" + jsonString("report") + ","
+ "\"sheet\":" + jsonString(SHEET) + ","
+ "\"context_hint\":" + jsonString(HINT) + ","
+ "\"audience\":" + jsonString("lab_report") + "," // lab_report|paper|certificate|internal
+ "\"review_notes\":" + jsonString(NOTES) + ","
+ "\"facts\":" + jsonString(FACTS)
+ "}";
System.out.println(PlusMinus.call("estimate", REPORT_INPUT, null)); // price each lane
HINT = "First-year teaching lab. A steel bob on a string, a metre tape, a hand-held " \
"stopwatch; the period is the mean of ten timings of twenty swings each."
NOTES = "Accepted from the budget review: reaction time of the stopwatch operator " \
"(Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the " \
"finite-amplitude correction were not evaluated; both are noted as caveats " \
"rather than added to the budget."
# Lane A - the budget review. FACTS is the STRING from the engine step.
INPUT = {
"task" => "budget",
"sheet" => SHEET,
"context_hint" => HINT,
"facts" => FACTS
}
est = call("estimate", INPUT)
puts "#{est['hold_credits']} #{est['min_credits']} #{est['model']}"
# 2652 310 gpt-5.6-terra - estimate is free and creates no job
REPORT_INPUT = {
"task" => "report",
"sheet" => SHEET,
"context_hint" => HINT,
"audience" => "lab_report", # lab_report | paper | certificate | internal
"review_notes" => NOTES,
"facts" => FACTS
}
puts call("estimate", REPORT_INPUT)["hold_credits"] # price each lane separately
<?php
$HINT = "First-year teaching lab. A steel bob on a string, a metre tape, a hand-held "
. "stopwatch; the period is the mean of ten timings of twenty swings each.";
$NOTES = "Accepted from the budget review: reaction time of the stopwatch operator "
. "(Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the "
. "finite-amplitude correction were not evaluated; both are noted as caveats "
. "rather than added to the budget.";
// Lane A - the budget review. $FACTS is the STRING from the engine step.
$INPUT = [
"task" => "budget",
"sheet" => $SHEET,
"context_hint" => $HINT,
"facts" => $FACTS,
];
$est = call("estimate", $INPUT);
echo $est["hold_credits"], " ", $est["min_credits"], " ", $est["model"], PHP_EOL;
// 2652 310 gpt-5.6-terra - estimate is free and creates no job
$REPORT_INPUT = [
"task" => "report",
"sheet" => $SHEET,
"context_hint" => $HINT,
"audience" => "lab_report", // lab_report | paper | certificate | internal
"review_notes" => $NOTES,
"facts" => $FACTS,
];
echo call("estimate", $REPORT_INPUT)["hold_credits"], PHP_EOL; // price each lane
const string Hint = "First-year teaching lab. A steel bob on a string, a metre tape, " +
"a hand-held stopwatch; the period is the mean of ten timings of twenty swings each.";
const string Notes = "Accepted from the budget review: reaction time of the stopwatch " +
"operator (Type B, rectangular, about 0.2 s per timing, divided by 20 swings) and the " +
"finite-amplitude correction were not evaluated; both are noted as caveats rather " +
"than added to the budget.";
// Lane A - the budget review. FACTS is the STRING from the engine step.
var input = new Dictionary<string, object?>
{
["task"] = "budget",
["sheet"] = Sheet,
["context_hint"] = Hint,
["facts"] = FACTS,
};
var est = await PlusMinus.Call("estimate", input);
Console.WriteLine($"{est.GetProperty("hold_credits")} {est.GetProperty("min_credits")}");
// 2652 310 - estimate is free and creates no job
var reportInput = new Dictionary<string, object?>
{
["task"] = "report",
["sheet"] = Sheet,
["context_hint"] = Hint,
["audience"] = "lab_report", // lab_report | paper | certificate | internal
["review_notes"] = Notes,
["facts"] = FACTS,
};
await PlusMinus.Call("estimate", reportInput); // price each lane separately
5. Run it, then poll
POST /run reserves hold_credits, starts the job and returns a
job_id at once; the reply arrives when GET /jobs/{id} reports
status: "succeeded". Send an Idempotency-Key on every run. The app's key is
plus-minus:<task>:<hash of the sheet, hint, audience and notes>:a<attempt>:
the lane is in it, so the same sheet through both lanes is two runs, and a retried request with the
same key returns the same job instead of billing twice. Bump the attempt suffix only when you mean
to run again.
# Same body as the estimate. The key is slug:task:hash:attempt.
KEY="plus-minus:budget:$(shasum -a 256 body.json | cut -c1-16):a1"
JOB=$(curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/run" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
--data-binary @body.json | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
# Poll until terminal. succeeded | failed | cancelled.
while :; do
STATE=$(curl -sS "https://api.skillsafe.ai/v1/app-api/jobs/$JOB" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN")
STATUS=$(echo "$STATE" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
case "$STATUS" in succeeded|failed|cancelled) break;; esac
sleep 2
done
echo "$STATE" | python3 -c 'import sys,json;d=json.load(sys.stdin)["data"];print(d["status"],d.get("charged_credits"),"credits, truncated:",d.get("truncated"));print(d["output"]["output"][:400])'
import hashlib, time
def run_and_wait(body, attempt=1, poll=2.0):
digest = hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()[:16]
key = f"{SLUG}:{body['task']}:{digest}:a{attempt}"
job = call("run", body, headers={"Idempotency-Key": key})
while True:
state = call(f"jobs/{job['job_id']}")
if state["status"] in ("succeeded", "failed", "cancelled"):
return state
time.sleep(poll)
state = run_and_wait(budget_input)
print(state["status"], state.get("charged_credits"), "credits; truncated:", state.get("truncated"))
raw = state["output"]["output"] # the model's reply: one JSON object as a string
import { createHash } from "node:crypto";
async function runAndWait(body, attempt = 1, pollMs = 2000) {
const digest = createHash("sha256").update(JSON.stringify(body)).digest("hex").slice(0, 16);
const key = `${SLUG}:${body.task}:${digest}:a${attempt}`;
const job = await call("run", body, { "Idempotency-Key": key });
for (;;) {
const state = await call(`jobs/${job.job_id}`);
if (["succeeded", "failed", "cancelled"].includes(state.status)) return state;
await new Promise(r => setTimeout(r, pollMs));
}
}
const state = await runAndWait(budgetInput);
console.log(state.status, state.charged_credits, "credits; truncated:", state.truncated);
const raw = state.output.output; // the model's reply as a string
func runAndWait(body map[string]string, attempt int) (map[string]any, error) {
enc, _ := json.Marshal(body)
sum := sha256.Sum256(enc)
key := fmt.Sprintf("%s:%s:%x:a%d", slug, body["task"], sum[:8], attempt)
job, err := call("run", body, map[string]string{"Idempotency-Key": key})
if err != nil {
return nil, err
}
for {
state, err := call("jobs/"+job["job_id"].(string), nil, nil)
if err != nil {
return nil, err
}
switch state["status"] {
case "succeeded", "failed", "cancelled":
return state, nil
}
time.Sleep(2 * time.Second)
}
}
// state, _ := runAndWait(budgetInput, 1)
// raw := state["output"].(map[string]any)["output"].(string)
static JsonNode runAndWait(Map<String, String> body, int attempt) throws Exception {
byte[] enc = MAPPER.writeValueAsBytes(body);
String digest = HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(enc)).substring(0, 16);
String key = SLUG + ":" + body.get("task") + ":" + digest + ":a" + attempt;
JsonNode job = call("run", body, Map.of("Idempotency-Key", key));
while (true) {
JsonNode state = call("jobs/" + job.get("job_id").asText(), null, Map.of());
String status = state.get("status").asText();
if (status.equals("succeeded") || status.equals("failed") || status.equals("cancelled")) return state;
Thread.sleep(2000);
}
}
// JsonNode state = runAndWait(budgetInput, 1);
// String raw = state.get("output").get("output").asText();
require "digest"
def run_and_wait(body, attempt: 1, poll: 2)
digest = Digest::SHA256.hexdigest(JSON.generate(body))[0, 16]
key = "#{SLUG}:#{body['task']}:#{digest}:a#{attempt}"
job = call("run", body, headers: { "Idempotency-Key" => key })
loop do
state = call("jobs/#{job['job_id']}")
return state if %w[succeeded failed cancelled].include?(state["status"])
sleep poll
end
end
state = run_and_wait(budget_input)
puts "#{state['status']} #{state['charged_credits']} credits; truncated: #{state['truncated']}"
raw = state["output"]["output"]
function runAndWait(array $body, int $attempt = 1, int $poll = 2): array {
$digest = substr(hash("sha256", json_encode($body)), 0, 16);
$key = SLUG . ":" . $body["task"] . ":" . $digest . ":a" . $attempt;
$job = call("run", $body, ["Idempotency-Key: $key"]);
while (true) {
$state = call("jobs/" . $job["job_id"]);
if (in_array($state["status"], ["succeeded", "failed", "cancelled"], true)) return $state;
sleep($poll);
}
}
$state = runAndWait($budgetInput);
echo $state["status"], " ", $state["charged_credits"] ?? "", " credits; truncated: ", var_export($state["truncated"] ?? false, true), "\n";
$raw = $state["output"]["output"];
static async Task<JsonElement> RunAndWait(Dictionary<string, string> body, int attempt = 1)
{
var enc = JsonSerializer.SerializeToUtf8Bytes(body);
var digest = Convert.ToHexString(SHA256.HashData(enc)).ToLowerInvariant()[..16];
var key = $"{PlusMinus.Slug}:{body["task"]}:{digest}:a{attempt}";
var job = await PlusMinus.Call("run", body, new() { ["Idempotency-Key"] = key });
while (true)
{
var state = await PlusMinus.Call($"jobs/{job.GetProperty("job_id").GetString()}");
var status = state.GetProperty("status").GetString();
if (status is "succeeded" or "failed" or "cancelled") return state;
await Task.Delay(2000);
}
}
// var state = await RunAndWait(budgetInput);
// var raw = state.GetProperty("output").GetProperty("output").GetString();
6. Or stream it
POST /run-stream takes the same body and the same key and answers with server-sent
events. From a script you see event: delta frames carrying pieces of the reply and one
final event: done with the finished job. From a browser page the platform sends only
tick heartbeats and the done event — which is why the app's progress card
advances on elapsed time between the real signals, and why a streaming preview built in a page would
be dead code. Concatenate the deltas; parse only when done arrives.
curl -sN -X POST "https://api.skillsafe.ai/v1/app-api/run-stream" \
-H "Authorization: Bearer $SKILLSAFE_TOKEN" \
-H "Content-Type: application/json" \
-H "Accept: text/event-stream" \
-H "Idempotency-Key: plus-minus:budget:$(shasum -a 256 body.json | cut -c1-16):a1" \
--data-binary @body.json
# event: job data: {"job_id":"job_..."}
# event: delta data: {"text":"{\"lane\":\"budget\",\"title\":\"g from a simple pendulum - budget review\","}
# event: delta data: {"text":"\"summary\":\"..."}
# event: done data: {"status":"succeeded","charged_credits":412,"truncated":false,"output":{"output":"{...}"}}
def run_stream(body, attempt=1):
digest = hashlib.sha256(json.dumps(body, sort_keys=True).encode()).hexdigest()[:16]
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(body).encode(), method="POST")
for k, v in {"Authorization": f"Bearer {TOKEN}", "Content-Type": "application/json",
"Accept": "text/event-stream", "Idempotency-Key": f"{SLUG}:{body['task']}:{digest}:a{attempt}"}.items():
req.add_header(k, v)
event, buf, done = None, [], None
with urllib.request.urlopen(req) as r:
for line in r:
line = line.decode().rstrip("\n")
if line.startswith("event:"):
event = line[6:].strip()
elif line.startswith("data:"):
data = json.loads(line[5:])
if event == "delta":
buf.append(data.get("text", ""))
elif event == "done":
done = data
raw = (done or {}).get("output", {}).get("output") or "".join(buf)
return done, raw
async function runStream(body, attempt = 1) {
const digest = createHash("sha256").update(JSON.stringify(body)).digest("hex").slice(0, 16);
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json",
Accept: "text/event-stream", "Idempotency-Key": `${SLUG}:${body.task}:${digest}:a${attempt}` },
body: JSON.stringify(body)
});
const reader = res.body.getReader(), dec = new TextDecoder();
let pending = "", event = null, buf = "", done = null;
for (;;) {
const { value, done: end } = await reader.read();
if (end) break;
pending += dec.decode(value, { stream: true });
let i;
while ((i = pending.indexOf("\n")) !== -1) {
const line = pending.slice(0, i); pending = pending.slice(i + 1);
if (line.startsWith("event:")) event = line.slice(6).trim();
else if (line.startsWith("data:")) {
const data = JSON.parse(line.slice(5));
if (event === "delta") buf += data.text || ""; // a STRING, not an object
else if (event === "done") done = data;
}
}
}
return { done, raw: (done && done.output && done.output.output) || buf };
}
func runStream(body map[string]string, attempt int) (map[string]any, string, error) {
enc, _ := json.Marshal(body)
sum := sha256.Sum256(enc)
req, _ := http.NewRequest("POST", base+"/run-stream", bytes.NewReader(enc))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Accept", "text/event-stream")
req.Header.Set("Idempotency-Key", fmt.Sprintf("%s:%s:%x:a%d", slug, body["task"], sum[:8], attempt))
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, "", err
}
defer res.Body.Close()
var event, buf string
var done map[string]any
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<24)
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event:"):
event = strings.TrimSpace(line[6:])
case strings.HasPrefix(line, "data:"):
var data map[string]any
json.Unmarshal([]byte(line[5:]), &data)
if event == "delta" {
buf += fmt.Sprint(data["text"])
} else if event == "done" {
done = data
}
}
}
raw := buf
if out, ok := done["output"].(map[string]any); ok {
raw = out["output"].(string)
}
return done, raw, nil
}
static String[] runStream(Map<String, String> body, int attempt) throws Exception {
byte[] enc = MAPPER.writeValueAsBytes(body);
String digest = HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(enc)).substring(0, 16);
HttpRequest req = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + TOKEN).header("Content-Type", "application/json")
.header("Accept", "text/event-stream")
.header("Idempotency-Key", SLUG + ":" + body.get("task") + ":" + digest + ":a" + attempt)
.POST(HttpRequest.BodyPublishers.ofByteArray(enc)).build();
HttpResponse<java.io.InputStream> res = CLIENT.send(req, HttpResponse.BodyHandlers.ofInputStream());
String event = null, doneJson = null; StringBuilder buf = new StringBuilder();
try (var in = new java.io.BufferedReader(new java.io.InputStreamReader(res.body()))) {
String line;
while ((line = in.readLine()) != null) {
if (line.startsWith("event:")) event = line.substring(6).trim();
else if (line.startsWith("data:")) {
JsonNode data = MAPPER.readTree(line.substring(5));
if ("delta".equals(event)) buf.append(data.path("text").asText(""));
else if ("done".equals(event)) doneJson = data.toString();
}
}
}
JsonNode done = doneJson == null ? null : MAPPER.readTree(doneJson);
String raw = done != null && done.has("output") ? done.get("output").get("output").asText() : buf.toString();
return new String[] { doneJson, raw };
}
def run_stream(body, attempt: 1)
digest = Digest::SHA256.hexdigest(JSON.generate(body))[0, 16]
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri, "Authorization" => "Bearer #{TOKEN}", "Content-Type" => "application/json",
"Accept" => "text/event-stream",
"Idempotency-Key" => "#{SLUG}:#{body['task']}:#{digest}:a#{attempt}")
req.body = JSON.generate(body)
event = nil; buf = +""; done = nil
Net::HTTP.start(uri.host, uri.port, use_ssl: true) do |http|
http.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
if line.start_with?("event:") then event = line[6..].strip
elsif line.start_with?("data:")
data = JSON.parse(line[5..])
if event == "delta" then buf << (data["text"] || "")
elsif event == "done" then done = data end
end
end
end
end
end
[done, done&.dig("output", "output") || buf]
end
function runStream(array $body, int $attempt = 1): array {
$digest = substr(hash("sha256", json_encode($body)), 0, 16);
$ch = curl_init(BASE . "/run-stream");
$event = null; $buf = ""; $done = null; $pending = "";
curl_setopt_array($ch, [
CURLOPT_POST => true, CURLOPT_POSTFIELDS => json_encode($body),
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . TOKEN, "Content-Type: application/json", "Accept: text/event-stream",
"Idempotency-Key: " . SLUG . ":" . $body["task"] . ":$digest:a$attempt"],
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$event, &$buf, &$done, &$pending) {
$pending .= $chunk;
while (($i = strpos($pending, "\n")) !== false) {
$line = substr($pending, 0, $i); $pending = substr($pending, $i + 1);
if (str_starts_with($line, "event:")) $event = trim(substr($line, 6));
elseif (str_starts_with($line, "data:")) {
$data = json_decode(substr($line, 5), true);
if ($event === "delta") $buf .= $data["text"] ?? "";
elseif ($event === "done") $done = $data;
}
}
return strlen($chunk);
},
]);
curl_exec($ch); curl_close($ch);
return [$done, $done["output"]["output"] ?? $buf];
}
static async Task<(JsonElement? done, string raw)> RunStream(Dictionary<string, string> body, int attempt = 1)
{
var enc = JsonSerializer.SerializeToUtf8Bytes(body);
var digest = Convert.ToHexString(SHA256.HashData(enc)).ToLowerInvariant()[..16];
using var req = new HttpRequestMessage(HttpMethod.Post, $"{PlusMinus.Base}/run-stream");
req.Headers.Authorization = new("Bearer", PlusMinus.Token);
req.Headers.Accept.ParseAdd("text/event-stream");
req.Headers.Add("Idempotency-Key", $"{PlusMinus.Slug}:{body["task"]}:{digest}:a{attempt}");
req.Content = new ByteArrayContent(enc) { Headers = { ContentType = new("application/json") } };
using var res = await PlusMinus.Http.SendAsync(req, HttpCompletionOption.ResponseHeadersRead);
using var reader = new StreamReader(await res.Content.ReadAsStreamAsync());
string? ev = null; var buf = new StringBuilder(); JsonElement? done = null;
while (await reader.ReadLineAsync() is { } line)
{
if (line.StartsWith("event:")) ev = line[6..].Trim();
else if (line.StartsWith("data:"))
{
var data = JsonDocument.Parse(line[5..]).RootElement;
if (ev == "delta") buf.Append(data.TryGetProperty("text", out var t) ? t.GetString() : "");
else if (ev == "done") done = data.Clone();
}
}
var raw = done?.GetProperty("output").GetProperty("output").GetString() ?? buf.ToString();
return (done, raw);
}
7. Parse the result
The reply is one JSON object, as a string, inside the job's output.output. Unwrap it
twice: once for the envelope, once for the model's text. Then do what the page does before it
trusts a single field — check the reply against the facts you sent. Three checks
catch nearly everything that goes wrong: the statement must be the engine's verbatim;
every engine flag id must appear exactly once in coverage; and each lane's list must
be complete — every sheet input judged once (budget), one budget-table row per input with an
uncertainty and exactly thirteen checklist rows (report). The page shows the engine's statement
regardless of what the model wrote, and lists every disagreement under "The page disagrees with
the reply".
# Unwrap twice, then check the statement and the flag coverage with python.
echo "$STATE" | python3 - body.json <<'PY'
import json, sys
state = json.load(sys.stdin)["data"]
facts = json.loads(json.load(open(sys.argv[1]))["facts"])
reply = json.loads(state["output"]["output"]) # the model's single JSON object
want = facts["result"]["statement"]["expanded"]
got = reply["statement"]["expanded"] if reply["lane"] == "budget" else reply["result_statement"]
print("statement verbatim:", got == want)
ids = [f["id"] for f in facts["flags"]]
cov = [c["id"] for c in reply["coverage"]]
print("flags covered once:", sorted(ids) == sorted(cov))
if reply["lane"] == "budget":
names = sorted(x["name"] for x in facts["inputs"] if not x.get("constant"))
print("inputs judged:", sorted(c["name"] for c in reply["contributors"]) == names, "| verdict:", reply["verdict"])
else:
print("13 checklist rows:", len(reply["checklist"]) == 13, "| met:", sum(c["status"] == "met" for c in reply["checklist"]))
PY
def check_reply(raw, body):
facts = json.loads(body["facts"])
reply = json.loads(raw.strip().strip("`").removeprefix("json").strip()) # tolerate a fence
problems = []
want = facts["result"]["statement"]["expanded"]
got = reply.get("statement", {}).get("expanded") if reply["lane"] == "budget" else reply.get("result_statement")
if got != want:
problems.append(f"statement differs: {got!r} vs engine {want!r}")
ids = sorted(f["id"] for f in facts["flags"]); cov = sorted(c["id"] for c in reply.get("coverage", []))
if ids != cov:
problems.append(f"flags not covered exactly once: engine {ids}, reply {cov}")
if reply["lane"] == "budget":
names = sorted(x["name"] for x in facts["inputs"] if not x.get("constant"))
if sorted(c["name"] for c in reply.get("contributors", [])) != names:
problems.append("not every sheet input was judged exactly once")
else:
if len(reply.get("checklist", [])) != 13:
problems.append(f"checklist has {len(reply.get('checklist', []))} rows, not 13")
want_rows = sorted(x["name"] for x in facts["inputs"] if x["u"] > 0)
if sorted(r["source"] for r in reply.get("budget_table", [])) != want_rows:
problems.append("budget table rows do not match the inputs with an uncertainty")
return reply, problems
reply, problems = check_reply(raw, budget_input)
print(reply["title"]); print(reply.get("verdict")); print(problems or "clean against the engine")
function checkReply(raw, body) {
const facts = JSON.parse(body.facts);
const t = raw.trim().replace(/^```[a-z]*\s*/i, "").replace(/```\s*$/, "");
const reply = JSON.parse(t.slice(t.indexOf("{"), t.lastIndexOf("}") + 1));
const problems = [];
const want = facts.result.statement.expanded;
const got = reply.lane === "budget" ? reply.statement?.expanded : reply.result_statement;
if (got !== want) problems.push(`statement differs: ${JSON.stringify(got)} vs engine ${JSON.stringify(want)}`);
const ids = facts.flags.map(f => f.id).sort().join(), cov = (reply.coverage || []).map(c => c.id).sort().join();
if (ids !== cov) problems.push(`flags not covered exactly once: engine ${ids} reply ${cov}`);
if (reply.lane === "budget") {
const names = facts.inputs.filter(x => !x.constant).map(x => x.name).sort().join();
if ((reply.contributors || []).map(c => c.name).sort().join() !== names) problems.push("not every sheet input was judged exactly once");
} else {
if ((reply.checklist || []).length !== 13) problems.push(`checklist has ${(reply.checklist || []).length} rows, not 13`);
const rows = facts.inputs.filter(x => x.u > 0).map(x => x.name).sort().join();
if ((reply.budget_table || []).map(r => r.source).sort().join() !== rows) problems.push("budget table rows do not match the inputs with an uncertainty");
}
return { reply, problems };
}
const { reply, problems } = checkReply(raw, budgetInput);
console.log(reply.title, reply.verdict, problems.length ? problems : "clean against the engine");
type Flag struct{ ID string `json:"id"` }
type Facts struct {
Result struct{ Statement struct{ Expanded string `json:"expanded"` } `json:"statement"` } `json:"result"`
Flags []Flag `json:"flags"`
Inputs []struct{ Name string `json:"name"`; U float64 `json:"u"`; Constant bool `json:"constant"` } `json:"inputs"`
}
func checkReply(raw string, body map[string]string) (map[string]any, []string) {
var facts Facts
json.Unmarshal([]byte(body["facts"]), &facts)
raw = strings.TrimSpace(raw)
raw = raw[strings.Index(raw, "{") : strings.LastIndex(raw, "}")+1]
var reply map[string]any
json.Unmarshal([]byte(raw), &reply)
var problems []string
got := ""
if reply["lane"] == "budget" {
if st, ok := reply["statement"].(map[string]any); ok { got, _ = st["expanded"].(string) }
} else {
got, _ = reply["result_statement"].(string)
}
if got != facts.Result.Statement.Expanded {
problems = append(problems, "statement differs from the engine's")
}
ids := map[string]int{}
for _, f := range facts.Flags { ids[f.ID]++ }
for _, c := range reply["coverage"].([]any) { ids[c.(map[string]any)["id"].(string)]-- }
for id, n := range ids { if n != 0 { problems = append(problems, "flag "+id+" not covered exactly once") } }
if reply["lane"] == "report" {
if rows, ok := reply["checklist"].([]any); !ok || len(rows) != 13 { problems = append(problems, "checklist is not 13 rows") }
}
return reply, problems
}
static List<String> checkReply(String raw, Map<String, String> body) throws Exception {
JsonNode facts = MAPPER.readTree(body.get("facts"));
String t = raw.trim();
JsonNode reply = MAPPER.readTree(t.substring(t.indexOf('{'), t.lastIndexOf('}') + 1));
List<String> problems = new ArrayList<>();
String want = facts.at("/result/statement/expanded").asText();
String got = reply.get("lane").asText().equals("budget") ? reply.at("/statement/expanded").asText() : reply.path("result_statement").asText();
if (!want.equals(got)) problems.add("statement differs: " + got + " vs engine " + want);
Set<String> ids = new TreeSet<>(), cov = new TreeSet<>();
facts.get("flags").forEach(f -> ids.add(f.get("id").asText()));
reply.path("coverage").forEach(c -> cov.add(c.get("id").asText()));
if (!ids.equals(cov) || cov.size() != reply.path("coverage").size()) problems.add("flags not covered exactly once");
if (reply.get("lane").asText().equals("report") && reply.path("checklist").size() != 13) problems.add("checklist has " + reply.path("checklist").size() + " rows, not 13");
if (reply.get("lane").asText().equals("budget")) {
Set<String> names = new TreeSet<>(), judged = new TreeSet<>();
facts.get("inputs").forEach(x -> { if (!x.path("constant").asBoolean(false)) names.add(x.get("name").asText()); });
reply.path("contributors").forEach(c -> judged.add(c.get("name").asText()));
if (!names.equals(judged)) problems.add("not every sheet input was judged exactly once");
}
return problems;
}
def check_reply(raw, body)
facts = JSON.parse(body["facts"])
t = raw.strip
reply = JSON.parse(t[t.index("{")..t.rindex("}")])
problems = []
want = facts.dig("result", "statement", "expanded")
got = reply["lane"] == "budget" ? reply.dig("statement", "expanded") : reply["result_statement"]
problems << "statement differs: #{got.inspect} vs engine #{want.inspect}" unless got == want
ids = facts["flags"].map { |f| f["id"] }.sort
cov = (reply["coverage"] || []).map { |c| c["id"] }.sort
problems << "flags not covered exactly once" unless ids == cov
if reply["lane"] == "budget"
names = facts["inputs"].reject { |x| x["constant"] }.map { |x| x["name"] }.sort
problems << "not every sheet input was judged exactly once" unless (reply["contributors"] || []).map { |c| c["name"] }.sort == names
else
problems << "checklist has #{(reply['checklist'] || []).size} rows, not 13" unless (reply["checklist"] || []).size == 13
end
[reply, problems]
end
reply, problems = check_reply(raw, budget_input)
puts reply["title"], reply["verdict"].to_s, (problems.empty? ? "clean against the engine" : problems)
function checkReply(string $raw, array $body): array {
$facts = json_decode($body["facts"], true);
$t = trim($raw);
$reply = json_decode(substr($t, strpos($t, "{"), strrpos($t, "}") - strpos($t, "{") + 1), true);
$problems = [];
$want = $facts["result"]["statement"]["expanded"];
$got = $reply["lane"] === "budget" ? ($reply["statement"]["expanded"] ?? null) : ($reply["result_statement"] ?? null);
if ($got !== $want) $problems[] = "statement differs: " . var_export($got, true) . " vs engine " . var_export($want, true);
$ids = array_map(fn($f) => $f["id"], $facts["flags"]); sort($ids);
$cov = array_map(fn($c) => $c["id"], $reply["coverage"] ?? []); sort($cov);
if ($ids !== $cov) $problems[] = "flags not covered exactly once";
if ($reply["lane"] === "budget") {
$names = array_map(fn($x) => $x["name"], array_filter($facts["inputs"], fn($x) => empty($x["constant"]))); sort($names);
$judged = array_map(fn($c) => $c["name"], $reply["contributors"] ?? []); sort($judged);
if (array_values($names) !== $judged) $problems[] = "not every sheet input was judged exactly once";
} elseif (count($reply["checklist"] ?? []) !== 13) {
$problems[] = "checklist has " . count($reply["checklist"] ?? []) . " rows, not 13";
}
return [$reply, $problems];
}
[$reply, $problems] = checkReply($raw, $budgetInput);
echo $reply["title"], "\n", $reply["verdict"] ?? "", "\n", $problems ? implode("\n", $problems) : "clean against the engine", "\n";
static (JsonElement reply, List<string> problems) CheckReply(string raw, Dictionary<string, string> body)
{
var facts = JsonDocument.Parse(body["facts"]).RootElement;
var t = raw.Trim();
var reply = JsonDocument.Parse(t[t.IndexOf('{')..(t.LastIndexOf('}') + 1)]).RootElement;
var problems = new List<string>();
var lane = reply.GetProperty("lane").GetString();
var want = facts.GetProperty("result").GetProperty("statement").GetProperty("expanded").GetString();
var got = lane == "budget"
? (reply.TryGetProperty("statement", out var st) && st.TryGetProperty("expanded", out var e) ? e.GetString() : null)
: (reply.TryGetProperty("result_statement", out var rs) ? rs.GetString() : null);
if (got != want) problems.Add($"statement differs: {got} vs engine {want}");
var ids = facts.GetProperty("flags").EnumerateArray().Select(f => f.GetProperty("id").GetString()).OrderBy(x => x).ToList();
var cov = reply.TryGetProperty("coverage", out var c) ? c.EnumerateArray().Select(x => x.GetProperty("id").GetString()).OrderBy(x => x).ToList() : new();
if (!ids.SequenceEqual(cov)) problems.Add("flags not covered exactly once");
if (lane == "budget")
{
var names = facts.GetProperty("inputs").EnumerateArray().Where(x => !(x.TryGetProperty("constant", out var k) && k.GetBoolean())).Select(x => x.GetProperty("name").GetString()).OrderBy(x => x);
var judged = reply.TryGetProperty("contributors", out var cs) ? cs.EnumerateArray().Select(x => x.GetProperty("name").GetString()).OrderBy(x => x) : Enumerable.Empty<string>();
if (!names.SequenceEqual(judged)) problems.Add("not every sheet input was judged exactly once");
}
else if (!(reply.TryGetProperty("checklist", out var cl) && cl.GetArrayLength() == 13)) problems.Add("checklist is not 13 rows");
return (reply, problems);
}
The output contract
Every reply carries the common envelope and then one lane body. The keys and enum values are exact; the page re-sequences ids defensively (M-001…, I-001…, R-001…) and falls back on any enum it does not recognise, but a script should treat an unknown value as a defect, not a feature.
{
"lane": "budget" | "report",
"title": string,
"summary": string,
"coverage": [ { "id": "E-001", "status": "confirmed" | "downgraded" | "dismissed" | "merged", "ref": string, "note": string } ],
"notes_on_input": string,
...lane body
}
The budget body
{
"measurand": { "name": string, "quantity": string, "method": string },
"contributors": [ { "name": string, "status": "reasonable" | "questionable" | "unsupported", "note": string } ],
"missing": [ { "id": "M-001", "what": string, "type": "A" | "B",
"distribution": "normal" | "rectangular" | "triangular" | "u_shaped",
"how_to_evaluate": string, "sheet_line": string, "priority": "high" | "medium" | "low" } ],
"improvements": [ { "id": "I-001", "action": string, "targets": [string], "expected_effect": string, "effort": "low" | "medium" | "high" } ],
"statement": { "expanded": string, "concise": string, "wording": string },
"verdict": "report_as_is" | "budget_incomplete" | "recompute_first",
"verdict_reason": string
}
statement.expanded and statement.concise are copies of
facts.result.statement; wording is the one sentence a report would carry
and may quote only digits that occur in facts.result.statement or
facts.mc.stated. Every contributors[].name is a sheet input, each exactly
once. Every missing[].sheet_line parses in the sheet grammar under a new name — the
page refuses to add one that does not, and says so on its card. The verdict rule the page re-derives:
recompute_first if any flag has severity error or an
en-fail / draft-value flag is present; else budget_incomplete
if any missing item is high; else report_as_is.
The report body
{
"audience": "lab_report" | "paper" | "certificate" | "internal",
"section_title": string,
"paragraphs": [string],
"budget_table": [ { "source": string, "value_text": string, "u_text": string, "unit": string, "type": "A" | "B",
"distribution": string, "dof_text": string, "sensitivity_text": string, "contribution_text": string, "share_pct": number } ],
"result_statement": string,
"checklist": [ { "id": "R-001", "item": string, "status": "met" | "not_met" | "n_a", "note": string } ],
"caveats": [string]
}
result_statement equals facts.result.statement.expanded. The budget table
has one row per input with u > 0 (a CODATA constant with an uncertainty included),
its share_pct values within 0.6 of the engine's and summing to about 100 when there are
no correlations. The checklist is exactly thirteen rows, R-001 to R-013, in the order the system
prompt lists: measurand defined; model function stated; every input has a value and a unit; every
input has a standard uncertainty; Type stated; distribution and degrees of freedom stated;
correlations stated or independence asserted; combined standard uncertainty given; effective degrees
of freedom given; coverage factor and probability given; expanded uncertainty in the result's unit or
relative; result and uncertainty rounded to the same place with at most two significant figures in U;
method of evaluation named.
The flags
The engine's flag rules, so a script can filter on them: dim-add, dim-func,
dim-pow, div-zero, domain, name-unknown,
unit-unknown, unit-mismatch, unit-mismatch-result,
formula-syntax, no-formula, corr-range,
corr-unknown (all severity error — no result); low-dof,
k-assumed, nonlinear, mc-bias, asymmetric,
rel-large, result-rel-large, few-observations,
u-digits, value-digits, tag-unknown, en-fail,
draft-k-missing, draft-mismatch, draft-digits,
draft-place, draft-value, draft-no-u (severity
warn); dominant, negligible, correlated,
distribution-effect, exact-input, type-b-normal,
constant-uncertain, constant-shadow, offset-unit,
unused-input, en-ok (severity info).
8. Use it as a gate
A measurement sheet kept beside a report can be checked on every change: the engine alone (no
credits) settles whether the sheet still computes and whether the statement in the report matches
it; the budget lane (metered) says whether the budget is complete. The zero-cost half is the one
worth wiring first — a result: line on the sheet holding the report's own statement
makes the engine's draft-* flags a free regression test on the number you publish.
#!/bin/sh
# Free gate: the sheet must compute and the report's statement must match the engine.
# Put the report's statement on the sheet as `result: ...` and fail on any draft-* flag.
set -e
curl -sS https://plus-minus.skillsafe.ai/units.js -o units.js
curl -sS https://plus-minus.skillsafe.ai/unc.js -o unc.js
node -e '
const fs=require("fs"),vm=require("vm"); global.window={};
vm.runInThisContext(fs.readFileSync("units.js","utf8")); vm.runInThisContext(fs.readFileSync("unc.js","utf8"));
const a=window.PMUnc.analyze(fs.readFileSync(process.argv[1],"utf8"));
if(!a.ok){console.error("sheet does not compute:",a.errors.map(e=>e.message).join("; "));process.exit(1);}
const bad=a.flags.filter(f=>/^draft-|^en-fail$/.test(f.rule));
console.log(a.statement.expanded);
if(bad.length){console.error(bad.map(f=>f.id+" "+f.rule+": "+f.message).join("\n"));process.exit(2);}
' measurement.sheet
# Metered gate: fail the build when the review says the budget is incomplete.
import subprocess
facts = subprocess.check_output(["node", "facts.js", "measurement.sheet"]).decode() # the step-4 script
body = {"task": "budget", "sheet": open("measurement.sheet").read(), "context_hint": "CI check", "facts": facts}
state = run_and_wait(body)
if state["status"] != "succeeded" or state.get("truncated"):
raise SystemExit(f"run did not complete: {state['status']} truncated={state.get('truncated')}")
reply, problems = check_reply(state["output"]["output"], body)
if problems:
raise SystemExit("reply disagrees with the engine: " + "; ".join(problems))
print(reply["statement"]["expanded"])
if reply["verdict"] != "report_as_is":
for m in reply["missing"]:
print(f" {m['id']} [{m['priority']}] {m['what']}\n {m['sheet_line']}")
raise SystemExit(f"verdict: {reply['verdict']} - {reply['verdict_reason']}")
// Metered gate for a CI step (Node 20+).
import { execFileSync } from "node:child_process";
import { readFileSync } from "node:fs";
const sheet = readFileSync("measurement.sheet", "utf8");
const facts = execFileSync("node", ["facts.js", "measurement.sheet"]).toString(); // the step-4 script
const body = { task: "budget", sheet, context_hint: "CI check", facts };
const state = await runAndWait(body);
if (state.status !== "succeeded" || state.truncated) { console.error("run did not complete", state.status, state.truncated); process.exit(1); }
const { reply, problems } = checkReply(state.output.output, body);
if (problems.length) { console.error("reply disagrees with the engine:", problems); process.exit(1); }
console.log(reply.statement.expanded);
if (reply.verdict !== "report_as_is") {
for (const m of reply.missing) console.log(` ${m.id} [${m.priority}] ${m.what}\n ${m.sheet_line}`);
console.error(`verdict: ${reply.verdict} - ${reply.verdict_reason}`); process.exit(2);
}
func main() {
sheet, _ := os.ReadFile("measurement.sheet")
facts, err := exec.Command("node", "facts.js", "measurement.sheet").Output() // the step-4 script
if err != nil {
log.Fatal(err)
}
body := map[string]string{"task": "budget", "sheet": string(sheet), "context_hint": "CI check", "facts": string(facts)}
state, err := runAndWait(body, 1)
if err != nil || state["status"] != "succeeded" || state["truncated"] == true {
log.Fatalf("run did not complete: %v %v", state["status"], err)
}
reply, problems := checkReply(state["output"].(map[string]any)["output"].(string), body)
if len(problems) > 0 {
log.Fatalf("reply disagrees with the engine: %v", problems)
}
fmt.Println(reply["statement"].(map[string]any)["expanded"])
if reply["verdict"] != "report_as_is" {
for _, m := range reply["missing"].([]any) {
mm := m.(map[string]any)
fmt.Printf(" %s [%s] %s\n %s\n", mm["id"], mm["priority"], mm["what"], mm["sheet_line"])
}
log.Fatalf("verdict: %v - %v", reply["verdict"], reply["verdict_reason"])
}
}
public static void main(String[] args) throws Exception {
String sheet = Files.readString(Path.of("measurement.sheet"));
Process p = new ProcessBuilder("node", "facts.js", "measurement.sheet").start(); // the step-4 script
String facts = new String(p.getInputStream().readAllBytes());
Map<String, String> body = new LinkedHashMap<>(Map.of("task", "budget", "sheet", sheet, "context_hint", "CI check", "facts", facts));
JsonNode state = runAndWait(body, 1);
if (!state.get("status").asText().equals("succeeded") || state.path("truncated").asBoolean(false)) {
System.err.println("run did not complete: " + state.get("status")); System.exit(1);
}
String raw = state.get("output").get("output").asText();
List<String> problems = checkReply(raw, body);
if (!problems.isEmpty()) { System.err.println("reply disagrees with the engine: " + problems); System.exit(1); }
JsonNode reply = MAPPER.readTree(raw.substring(raw.indexOf('{'), raw.lastIndexOf('}') + 1));
System.out.println(reply.at("/statement/expanded").asText());
if (!reply.get("verdict").asText().equals("report_as_is")) {
reply.get("missing").forEach(m -> System.out.printf(" %s [%s] %s%n %s%n", m.get("id").asText(), m.get("priority").asText(), m.get("what").asText(), m.get("sheet_line").asText()));
System.err.println("verdict: " + reply.get("verdict").asText() + " - " + reply.get("verdict_reason").asText()); System.exit(2);
}
}
sheet = File.read("measurement.sheet")
facts = `node facts.js measurement.sheet` # the step-4 script
body = { "task" => "budget", "sheet" => sheet, "context_hint" => "CI check", "facts" => facts }
state = run_and_wait(body)
abort "run did not complete: #{state['status']} truncated=#{state['truncated']}" unless state["status"] == "succeeded" && !state["truncated"]
reply, problems = check_reply(state["output"]["output"], body)
abort "reply disagrees with the engine: #{problems.join('; ')}" unless problems.empty?
puts reply["statement"]["expanded"]
if reply["verdict"] != "report_as_is"
reply["missing"].each { |m| puts " #{m['id']} [#{m['priority']}] #{m['what']}\n #{m['sheet_line']}" }
abort "verdict: #{reply['verdict']} - #{reply['verdict_reason']}"
end
$sheet = file_get_contents("measurement.sheet");
$facts = shell_exec("node facts.js measurement.sheet"); // the step-4 script
$body = ["task" => "budget", "sheet" => $sheet, "context_hint" => "CI check", "facts" => $facts];
$state = runAndWait($body);
if ($state["status"] !== "succeeded" || !empty($state["truncated"])) { fwrite(STDERR, "run did not complete\n"); exit(1); }
[$reply, $problems] = checkReply($state["output"]["output"], $body);
if ($problems) { fwrite(STDERR, "reply disagrees with the engine: " . implode("; ", $problems) . "\n"); exit(1); }
echo $reply["statement"]["expanded"], "\n";
if ($reply["verdict"] !== "report_as_is") {
foreach ($reply["missing"] as $m) echo " {$m['id']} [{$m['priority']}] {$m['what']}\n {$m['sheet_line']}\n";
fwrite(STDERR, "verdict: {$reply['verdict']} - {$reply['verdict_reason']}\n"); exit(2);
}
var sheet = File.ReadAllText("measurement.sheet");
var facts = RunNode("facts.js", "measurement.sheet"); // the step-4 script, via Process.Start
var body = new Dictionary<string, string> { ["task"] = "budget", ["sheet"] = sheet, ["context_hint"] = "CI check", ["facts"] = facts };
var state = await RunAndWait(body);
if (state.GetProperty("status").GetString() != "succeeded" || (state.TryGetProperty("truncated", out var tr) && tr.GetBoolean()))
{ Console.Error.WriteLine("run did not complete"); return 1; }
var (reply, problems) = CheckReply(state.GetProperty("output").GetProperty("output").GetString()!, body);
if (problems.Count > 0) { Console.Error.WriteLine("reply disagrees with the engine: " + string.Join("; ", problems)); return 1; }
Console.WriteLine(reply.GetProperty("statement").GetProperty("expanded").GetString());
if (reply.GetProperty("verdict").GetString() != "report_as_is")
{
foreach (var m in reply.GetProperty("missing").EnumerateArray())
Console.WriteLine($" {m.GetProperty("id")} [{m.GetProperty("priority")}] {m.GetProperty("what")}\n {m.GetProperty("sheet_line")}");
Console.Error.WriteLine($"verdict: {reply.GetProperty("verdict")} - {reply.GetProperty("verdict_reason")}"); return 2;
}
return 0;
Truncation and partial results
When the balance sits between min_credits and hold_credits, the run is not
refused: it executes with a reduced output cap and comes back with truncated: true on
the finished job and on the streaming done event. What you hold then is a prefix of the
reply. The envelope and coverage come first, so in the budget lane a truncated reply
may carry every flag's status and the judged contributors while missing,
improvements, statement and verdict — the fields a gate reads
— are absent or cut mid-string. In the report lane the paragraphs are the longest strings and come
early, so what gets lost is the budget table, the checklist and the caveats.
Check the flag before you treat a reply as complete. A truncated budget review with
two judged inputs parses cleanly and looks like a short review. The page's answer is to close the
JSON structure that arrived, render every section that parsed, and say "the stream ended early"
above it; a script should do the same or retry. The right retry is a top-up or a smaller sheet with
the attempt suffix on the Idempotency-Key incremented — never a repair that appends
braces and calls the result a review.
One more honest limit: the sheet is clipped from the middle at 12,000 characters, whole lines at a time, with a comment line marking the cut, and both the engine and the model work on what remains. A budget over a clipped sheet is a budget of the inputs that survived; the statement the engine computes is still exact for those inputs, but it is not the measurement's. Keep one measurement per sheet — at a few dozen inputs a sheet is nowhere near the budget.