Release 1.1 (pre-release)
Competition mining (planned)
A second way to mine, designed for the release after 1.1: 20 % of miner emission goes to the current number one of a public competition for faster or cheaper versions of the production models. It is switched off until it is proven robust.
Status · OFF
Designed, implemented on a separate branch behind the feature flag REAX_FEATURE_COMPETITION (default off), and not yet robust enough to switch on. While it is off, 100 % of miner emission stays with inference miners, as in 1.1. Nothing here runs on testnet or mainnet. Source: DESIGN-v1.2-competition.md and READINESS.md.
The split
- The owner share is taken first, unchanged.
- With no current winner, or with the flag off, the 20 % goes back to the inference pools pro rata, never to the owner.
- The winner is a competition hotkey that commits a submission, not an inference role. One hotkey cannot be both, because the Commitments pallet allows one commitment per hotkey and subnet; an operator who wants both registers two UIDs.
Tracks
| Track | Goal | Quality gate |
|---|---|---|
s1-pro-speed | Faster s1-pro with provably identical outputs | Weights byte-identical to the pinned Rune files, or a declared quantisation recipe the evaluator re-runs to byte-identical files; and on 2,000 fresh items every 3-item round mean stays within logit distance 0.20 with no item above 2.0, against the deterministic production runtime |
s1-fast-speed | Faster or cheaper s1-fast | JevBench Capability at least the incumbent's on sealed rotating items (the evaluator receives only the score), and agreement with s1-pro on fresh generated items not lower than the incumbent's. Both must pass |
In the 1.1 calibration, an online FP8 build of Rune did not pass the identical-output check; faster exact inference code can.
Submission format
A competition hotkey commits (at most 128 bytes):
The public Hugging Face repository at that commit must contain:
reax-submission.json: track, the base manifest it claims to serve, the runtime image recipe (Dockerfile with a base image pinned by digest and pinned packages), the weights list with sha256, and for quantised s1-pro weights the exact reproducible recipe (tool, pinned version, command) that derives them from the pinned Rune weights.- The runtime: a sidecar serving the existing 1.1 internal contract (
POST /internal/v1/decide,/health, loopback, token), so a winning runtime can later be adopted by inference miners through a new manifest. - A model card with the licence of everything inside (Apache or MIT compatible only).
The commit hash pins everything; a later push changes nothing for that submission. Submission time is the block of the commitment, and the earlier block wins a tie.
Review and sandbox
An evaluator runs each submission on isolated rented GPU pods, never on the project's own servers and never with secrets. In the first version the evaluator is run by the owner; results are signed, published with full logs and seeds, and anyone can rerun them. That is a stated trust assumption; a fully decentralised evaluation is a later step.
- Fetch the repository at the pinned commit; reject oversized files, symlinks leaving the tree and LFS pointers to other repositories.
- Automatic code review before anything runs. Static checks for network code paths, subprocess or eval of fetched data, obfuscated blobs, references to hosted LLM endpoints, reads of environment secrets and benchmark detection; then an LLM review with a verdict of PASS, FAIL or NEEDS-HUMAN. Anything but PASS stops the pipeline and is published.
- Sandboxed build and run: the image is built from the submitted Dockerfile and runs with
--network none, a read-only root, no secrets and CPU, RAM and GPU limits; weights are mounted read-only after their hashes match. - Equivalence on fresh items whose generator seeds are drawn after the submission block and revealed in the scoreboard.
- Speed and cost on reference hardware (RTX PRO 6000 96 GB for both tracks at first): sustained decisions per second at p95 of at most 300 ms (s1-fast) or 1 s (s1-pro) on a fixed public workload mix, three runs, median. Score is throughput at equal or better quality.
Winner rules
- Margin: a challenger replaces the incumbent only with at least 5 % higher throughput (for s1-fast, or at least 5 % lower cost at equal throughput), measured in the same batch as a fresh re-measurement of the incumbent.
- Holding period: a new number one is confirmed after a second evaluation at least 24 hours later reproduces the margin; until then the incumbent keeps the 20 %.
- Copy detection: identical weight hashes, or runtime code differing only cosmetically, credit the earlier submission and disqualify the copy. Work derived from another submission must beat it by the margin.
- Re-measurement: the incumbent is re-measured weekly; if it fails a gate (for example its repository disappears) the title moves to the next valid submission.
Main threats and handling
- Benchmark detection: fresh items with post-hoc seeds, sealed JevBench items that never leave the owner pipeline, static and LLM review.
- Calling a hosted model: no network during scored runs; static review flags client code.
- Weight copying: copy detection, earliest block, margin for derived work.
- Many submissions: a registration cost per competition hotkey and a rate-limited queue per coldkey.
- Evaluator bias or compromise: signed published receipts and reproducible reruns.
Not in scope yet
Decentralised evaluation and customer fine-tuning jobs through the competition. See the roadmap.