REAX.docs Launching soon · not live yet
MenuVerification & calibration

Release 1.1 (pre-release)

Verification and calibration

A validator checks a miner by asking it fresh questions and recomputing the same answers itself on the same pinned model. The two sets of probabilities are compared in logit space, and the thresholds were set from measurements on three GPU types.

Status · 3 October 2026

Release candidate 1.1.0rc1, pre-release; calibrated on rented GPUs and tested on a local chain, not on testnet or mainnet. All numbers below are from docs/CALIBRATION.md and docs/DESIGN-v1.1.md in the private repository.

Fresh items, no answer key

Items come from a cryptographically seeded generator per miner and round, so there is no answer key to leak. Each is a typed decision: a short text or JSON state and one choice (2 to 8 options), noul or score (3 to 5 levels) question. Every round a miner gets 8 items. The validator draws which items it verifies with its own secret only after the answers arrive, so the miner cannot tell them apart.

The logit distance

From the miner's probabilities p and the reference's q the validator computes centred log-ratios and takes the largest disagreement between option logits, in the model's own units:

d(p, q) = T · max_i | (ln p_i − mean_j ln p_j) − (ln q_i − mean_j ln q_j) |p and q are floored at 1e-6. T is the production temperature (s1-fast 2.07, s1-pro 2.0).

The distance does not depend on the softmax temperature and does not saturate when the model is confident: a near-one-hot answer still carries the exact logit gaps, which are a fingerprint of the weights and the prompt. The earlier mechanism used a protocol temperature of 4; 1.1 keeps the production temperature because answers must be production answers.

Thresholds

D is the mean distance over the round's verified items.

ConditionResult for the round
Malformed, late, unsigned or inconsistent responsescores 0
D ≤ 0.20scores 0.8 + 0.2 × speed
D > 0.20scores 0; the round is lost
D > 0.50, or any verified item above 2.0strike: the round scores 0 and the moving average resets to 0

Speed is measured by the validator's own clock, within latency bounds of 0.5 to 8 s for s1-fast and 1 to 12 s for s1-pro, and is 20 % of a passing round's score. Weights are per pool with fixed shares.

What the calibration measured

The calibration used 871 decisions (640 from the validator generator plus 231 public JevBench items) with the exact production runtimes on RTX PRO 6000 Blackwell, H100 80 GB and RTX 4090 (s1-fast only).

s1-fast (Plumb-4B)

PairMedian dMax dTop choice agrees
Production path, new process, same GPU00100 %
Production path, H100 vs RTX PRO 60000.0710.70898.7 %
FP8 build of the same weights0.2503.0096.8 %
Closest sibling model, best temperature0.895.7983.2 %
Base Qwen3.5-4B1.4310.7074.4 %

Over 8 verified items, an honest miner on any pair of the three GPU types scores 0.9998 to 1.0 of its rounds and is never struck (worst round mean 0.24). An FP8 build scores about 0.095 to 0.098 of its rounds, which keeps its moving average near 0.1, below the 0.3 needed for any weight. Other models are struck in at least 0.997 of rounds. Raising the mean threshold to 0.25 would let FP8 score 28 % of rounds, which is why 0.20 was chosen.

s1-pro (Rune 26B-A4B v3)

Why a spot check is enough for s1-pro

If a cheater answers a fraction x of items with something cheaper and 3 of 8 are verified, a round catches it with probability 0.375 at x = 1/8, 0.64 at x = 1/4 and 0.93 at x = 1/2. A catch resets the moving average, and rebuilding takes about 22 rounds. Cheating therefore costs more rounds than it saves compute.

Attacks and defences

AttackDefence
Wrong or cheaper model, quantised buildLogit distance on fresh items; FP8 and other models measured explicitly
Cached answersItems are fresh per miner and round
Copying another minerHotkey-bound requests and responses; per-miner items
Latency spoofingValidator-side clock; only 20 % weight, and only on passing rounds
Answering only the verified items wellVerified set is drawn after the response
Production-looking fields over wrong probabilitiesValidator recomputes derived fields
Hopping between models to dodge strikesState keyed by hotkey, model and manifest

Validator load

The s1-fast reference verifies 8 items per miner and round. At 30 decisions/s (RTX 4090), 255 miners need about 68 s of the 120 s round; at 55 to 60 decisions/s (RTX PRO 6000, H100) about 35 s. A 24 GB validator card is fine for up to roughly 150 s1-fast miners.

Not measured

A100, L4, L40S, H200, RTX 5090 and other cards; FP8 or INT4 builds other than vLLM online FP8; and a model distilled on the public item generator (the closest relative averages seven times the threshold). Repository documents: docs/CALIBRATION.md, docs/DESIGN-v1.1.md, docs/VALIDATOR.md, docs/THREAT-MODEL.md.