Release 1.1 (pre-release)
Verification and calibration
A validator checks a miner by asking it fresh questions and recomputing the same answers itself on the same pinned model. The two sets of probabilities are compared in logit space, and the thresholds were set from measurements on three GPU types.
Status · 3 October 2026
Release candidate 1.1.0rc1, pre-release; calibrated on rented GPUs and tested on a local chain, not on testnet or mainnet. All numbers below are from docs/CALIBRATION.md and docs/DESIGN-v1.1.md in the private repository.
Fresh items, no answer key
Items come from a cryptographically seeded generator per miner and round, so there is no answer key to leak. Each is a typed decision: a short text or JSON state and one choice (2 to 8 options), noul or score (3 to 5 levels) question. Every round a miner gets 8 items. The validator draws which items it verifies with its own secret only after the answers arrive, so the miner cannot tell them apart.
- s1-fast: all 8 items are verified.
- s1-pro: a secret 3 of 8 items are verified (a spot check); the rest are real work the miner does anyway.
The logit distance
From the miner's probabilities p and the reference's q the validator computes centred log-ratios and takes the largest disagreement between option logits, in the model's own units:
The distance does not depend on the softmax temperature and does not saturate when the model is confident: a near-one-hot answer still carries the exact logit gaps, which are a fingerprint of the weights and the prompt. The earlier mechanism used a protocol temperature of 4; 1.1 keeps the production temperature because answers must be production answers.
Thresholds
D is the mean distance over the round's verified items.
| Condition | Result for the round |
|---|---|
| Malformed, late, unsigned or inconsistent response | scores 0 |
| D ≤ 0.20 | scores 0.8 + 0.2 × speed |
| D > 0.20 | scores 0; the round is lost |
| D > 0.50, or any verified item above 2.0 | strike: the round scores 0 and the moving average resets to 0 |
Speed is measured by the validator's own clock, within latency bounds of 0.5 to 8 s for s1-fast and 1 to 12 s for s1-pro, and is 20 % of a passing round's score. Weights are per pool with fixed shares.
What the calibration measured
The calibration used 871 decisions (640 from the validator generator plus 231 public JevBench items) with the exact production runtimes on RTX PRO 6000 Blackwell, H100 80 GB and RTX 4090 (s1-fast only).
s1-fast (Plumb-4B)
| Pair | Median d | Max d | Top choice agrees |
|---|---|---|---|
| Production path, new process, same GPU | 0 | 0 | 100 % |
| Production path, H100 vs RTX PRO 6000 | 0.071 | 0.708 | 98.7 % |
| FP8 build of the same weights | 0.250 | 3.00 | 96.8 % |
| Closest sibling model, best temperature | 0.89 | 5.79 | 83.2 % |
| Base Qwen3.5-4B | 1.43 | 10.70 | 74.4 % |
Over 8 verified items, an honest miner on any pair of the three GPU types scores 0.9998 to 1.0 of its rounds and is never struck (worst round mean 0.24). An FP8 build scores about 0.095 to 0.098 of its rounds, which keeps its moving average near 0.1, below the 0.3 needed for any weight. Other models are struck in at least 0.997 of rounds. Raising the mean threshold to 0.25 would let FP8 score 28 % of rounds, which is why 0.20 was chosen.
s1-pro (Rune 26B-A4B v3)
- The deterministic build is bit-identical (distance 0) across restarts, concurrency 1, 16 and 64, and between H100 and RTX PRO 6000.
- Stock vLLM without the deterministic flags: median 0.146, max 3.94; it passes 0.65 of rounds, so miners must run the pinned deterministic runtime.
- FP8 of the same weights: median 0.313, max 4.1 to 5.0; it passes 0.12 of rounds and is struck in 22 %.
Why a spot check is enough for s1-pro
If a cheater answers a fraction x of items with something cheaper and 3 of 8 are verified, a round catches it with probability 0.375 at x = 1/8, 0.64 at x = 1/4 and 0.93 at x = 1/2. A catch resets the moving average, and rebuilding takes about 22 rounds. Cheating therefore costs more rounds than it saves compute.
Attacks and defences
| Attack | Defence |
|---|---|
| Wrong or cheaper model, quantised build | Logit distance on fresh items; FP8 and other models measured explicitly |
| Cached answers | Items are fresh per miner and round |
| Copying another miner | Hotkey-bound requests and responses; per-miner items |
| Latency spoofing | Validator-side clock; only 20 % weight, and only on passing rounds |
| Answering only the verified items well | Verified set is drawn after the response |
| Production-looking fields over wrong probabilities | Validator recomputes derived fields |
| Hopping between models to dodge strikes | State keyed by hotkey, model and manifest |
Validator load
The s1-fast reference verifies 8 items per miner and round. At 30 decisions/s (RTX 4090), 255 miners need about 68 s of the 120 s round; at 55 to 60 decisions/s (RTX PRO 6000, H100) about 35 s. A 24 GB validator card is fine for up to roughly 150 s1-fast miners.
Not measured
A100, L4, L40S, H200, RTX 5090 and other cards; FP8 or INT4 builds other than vLLM online FP8; and a model distilled on the public item generator (the closest relative averages seven times the threshold). Repository documents: docs/CALIBRATION.md, docs/DESIGN-v1.1.md, docs/VALIDATOR.md, docs/THREAT-MODEL.md.