Release 1.1 (pre-release)
Inference mining (1.1)
In release 1.1 a miner serves exactly one of the System1 production models: the same weights, the same prompt, the same one-pass readout and the same answer schema as production. A validator can then check a miner by recomputing the answer on its own copy of the same model.
Status · 3 October 2026
Release candidate 1.1.0rc1, pre-release. It was built and tested on a local chain with the real production runtimes. It has not been run on testnet or mainnet, no REAX subnet is registered and there is no netuid. Numbers on this page come from the repository documents named in each section.
Roles
- Miner. One model per hotkey. The miner answers signed
reax-s1/1requests with production's answer objects, using a model runtime on loopback that reproduces production exactly. - Validator. Runs the reference runtime of every enabled model on its own GPU, sends fresh items to miners each round, and recomputes the answers to compare them. See Verification & calibration.
- Owner. Any pool share without an enrolled miner goes to the owner UID, so a subnet with no miners is safe.
The two models
| s1-fast | s1-pro | |
|---|---|---|
| Model | Plumb-4B | Surogate Rune 26B-A4B v3 |
| Weights | crh225/plumb-4b at revision 55de037, public | surogate/rune-26b-a4b-GGUF at revision c6b360d (bf16 safetensors despite the repository name), Hugging Face gated |
| Licence | Apache-2.0 | Apache-2.0 |
| GPU | one NVIDIA GPU with 24 GB or more | one GPU with 80 to 96 GB |
| Pool at first | on | switched off |
| Readout temperature | 2.07 | 2.0 |
| Options per question | 2 to 16 (the subnet generator uses 2 to 8) | 2 to 64 (the subnet generator uses 2 to 8) |
Commercial serving by third-party miners is allowed by both licences. Rune requires every operator to accept the Hugging Face gate with their own account; REAX never mirrors the weights. Limits for both models: state of at most 16 KiB and at most 4,096 input tokens.
On-chain commitment
A miner publishes a plaintext commitment through the Commitments pallet:
model is s1-fast or s1-pro. The manifest_id is the first 16 hex characters of the sha256 of the canonical manifest bytes (s1-fast ff6e01b983d48024, s1-pro 5b11440bf57b5434). A hotkey without a valid commitment for an enabled model is not queried and earns nothing. Changing the commitment resets that UID's score: state is keyed by hotkey, model and manifest.
Manifests and runtimes
A versioned manifest in the repository pins the Hugging Face revision and the sha256 of every weight file the runtime reads (miners verify them before serving), the runtime base image by digest, package pins, the readout settings and the limits. The s1-pro runtime is vLLM 0.30.0 with deterministic flags (VLLM_BATCH_INVARIANT=1, no chunked prefill, -O0); stock vLLM without them loses about a third of its rounds in the calibration. The s1-fast runtime uses the production CUDA-graph, batch-1 path.
Request and answer
Requests are POST /v1/s1/decide, signed in both directions as in the earlier mechanism. Each decision is a System1 body without model: one question, no images. An answer is production's normalized answer object: type, the probability of every option, plus choice and confidence, noul, or score, confidence and legend. A response whose derived fields disagree with its own probabilities is rejected.
Hardware measured
| GPU | s1-fast | s1-pro |
|---|---|---|
| RTX PRO 6000 Blackwell (96 GB) | measured, 55.6 decisions/s | measured, 17 decisions/s at concurrency 1 |
| H100 80 GB | measured, 60.5 decisions/s | measured, 15 decisions/s at concurrency 1 |
| RTX 4090 (24 GB) | measured, 30.4 decisions/s | not applicable (too small) |
Throughput is the production path, one decision at a time. The deterministic s1-pro build gave bit-identical answers on the RTX PRO 6000 and the H100. A100, L4, L40S, H200, RTX 5090 and other cards should work but are not measured: run scripts/calibration_s1 in the repository first.
Switches and defaults
| Setting | Default in 1.1.0rc1 |
|---|---|
REAX_MECHANISM | s1; mpm1 selects the earlier mechanism (Qwen3-4B) as a fallback |
| s1-fast pool | on; share 1.0 when alone |
s1-pro pool (REAX_FEATURE_S1_PRO_POOL) | off. Every validator would need an 80 to 96 GB reference GPU, and turning it on changes weights for all validators, so it is a release plus owner decision. When on: shares s1-fast 0.35 and s1-pro 0.65 |
| Images, multi-question requests, customer traffic, capacity audits | off; startup refuses them |
Scoring constants are release constants, and validators refuse overrides on finney. The pool shares are not a statement about the value of anything; they only split the weights.
Open points
- API relay. A miner without a GPU could relay validator items to the public System1 API and pass. This is an accepted risk pending an owner decision before miners are told about the subnet (repository
docs/THREAT-MODEL.md, S15). - Many hotkeys on one GPU is an accepted, documented risk.
- Peer-to-Peer tier routing is design only; see the roadmap.
Repository documents: READINESS.md, docs/DESIGN-v1.1.md, docs/MINER.md, docs/VALIDATOR.md, docs/CALIBRATION.md, docs/THREAT-MODEL.md. The repository is private.