Reference
Scoring and incentive mechanism
REAX rewards one thing: serving the approved model faithfully, reliably, at reference speed, with real capacity. This page explains in plain words how validators turn evidence into weights, then gives the formulas from spec revision 0.6.
Provisional mechanism
Spec 0.6 is not frozen. Every constant marked (provisional) lives in signed policy and can change; tolerances are placeholders until honest-noise measurements exist. No weights are set on any network today. Emissions are an incentive mechanism only: they are not customer billing, not a service guarantee and not a promise of any return.
In plain words
- Serve the right model. Validators compare your probabilities with the reference model's on the exact same bytes. If you consistently drift, you fail a hard gate and earn nothing for the affected epochs. Beating the reference's accuracy earns nothing extra; falling measurably below it fails the quality gate.
- Be there. Availability multiplies everything. Below 50 % (provisional) you are not eligible.
- Be as fast as the reference. Latency within 10 % (provisional) of the reference manifest gets full credit. Faster earns nothing extra, so a cheaper, faster substitute gains nothing.
- Bring real capacity. Verified capacity, measured in synchronized audited bursts, scales your weight. In the spec's models, declaring more concurrency than you have does not help.
- One operator, one share. Hotkeys are grouped by operator, and no group can take more than 40 % (provisional) of the weight.
The gain formula
For each miner i and epoch, over the set Φi of question families its model alias covers:
There is no quality score, cost term or identity score inside g. Identity and quality act only as hard gates and through manifest admission. Two miners that both pass the gates differ only by availability, latency and capacity.
Ramp
Increases are gradual and decreases apply at once. A miner in fidelity_watch cannot increase. Other miners joining or leaving change only the normalization.
Fidelity gates
Two statistics compare a miner with the reference output p_ref for the same item:
At window close each test uses a one-sided empirical-Bernstein bound (Maurer and Pontil 2009) on the window mean:
| Parameter | Value |
|---|---|
Identity window size n_id,win | 2 000 (provisional) |
Labelled window size n_lab,win | 288 (provisional) |
Burst audit sample n_b,win | 2 016 |
Per-window error budget δ_w | 0.005 (provisional), split into a fixed δ_test = 0.000625 per test and a separate tripwire budget |
Tolerances τ_id, τ_item, ε | 0.05, 0.04, 0.02: placeholders until measured per hardware configuration |
Code vs spec
The current validator code uses 288-cluster windows for both statistics and δ = 0.01; the 2 000-item identity window is the spec 0.6 target.
Two strikes. The first failing window puts the miner in fidelity_watch: its applied gain is frozen and the next window uses only never-exposed items. A second consecutive failing window, on disjoint evidence, makes it fidelity_failed for the affected epochs: F = 0 and κ = 0. Exclusion from the network is disabled in this revision, because the independence assumption it would rest on is not yet proven; the miner is simply re-evaluated in the next window.
Support. A miner with too few items in its first window is unsupported: no weight and no penalty. A miner on a hardware configuration that has not been calibrated is also unsupported, never penalized.
Model admission
Before any miner is scored on a manifest, the manifest itself must pass two checks per question family:
- Absolute grade. On at least 1 000 (provisional) label-balanced items, the lower bound of its skill over the uniform baseline must be at least 0.02 (provisional). A constant uniform predictor scores 0 and is not admitted. If the grade fails, the family is dropped for that manifest for every miner equally.
- Power admission. On at least 5 000 (provisional) identity items, the identity test must provably catch a signed panel of cheap substitutes (uniform, climatology, label-frequency and at least one cheaper sibling). If any panel member would go undetected, the manifest is not approved for that family.
Substitutes that stay within tolerance on most items are output-equivalent: they earn like the reference, and the catalog publishes every known one. Output fingerprinting never proves which weights are running; only separately evaluated, nonce-bound hardware attestation could do that.
Availability
Counted as failures: timeouts after dispatch, connection refused or reset, invalid or late responses, wrong identity or signature. Excluded: burst dispatches, pre-dispatch expiry, validator or gateway faults and injection-API failures, windows where the gateway's egress health probe failed, chain or RPC faults, revoked_in_flight and payload_binding_mismatch.
Latency, capped at the reference
Every miner whose p50 and p95 are within (1 + η) of the reference receives L = 1, and nothing below that earns more. A substitute that is faster than the reference therefore gains nothing over an honest miner at reference latency. Latency is measured by the serving region's gateway on a monotonic clock, from dispatch to verified receipt.
Worked example from the spec's acceptance vectors (tref50 = 400 ms, tref95 = 1 000 ms, η = 0.1): p50 = 300 ms and p95 = 900 ms gives L = 1.0; p50 = 640 ms and p95 = 1 600 ms gives f50 = 0.8, f95 = 0.5, L = 0.65. These are test values, not measured reference latencies.
Code vs spec
The current validator code uses a simpler transform, L = 0.5·max(0, 1 − p50/deadline) + 0.5·max(0, 1 − p95/deadline) against a signed per-family deadline, which is not capped at the reference. The capped f_q above is the spec 0.6 target.
Hidden probes
Validators do not announce which requests are tests. On the GDPR tier, probes enter through the gateway with the same caller ID, TLS configuration, egress addresses and connection pooling as organic traffic; their deadlines, arrival times and field shapes follow privacy-reviewed organic profiles; and every receipt carries the same fields, so receipts do not reveal probe status. Identity items are used exactly once and labelled items are never shown twice to the same operator.
What is not assumed: that probe content is indistinguishable from real content. A generator's questions can differ from customers' questions, which is exactly why the organic identity shadow is mandatory before real traffic. On the permissionless route every dispatch is synthetic, so indistinguishability is not claimed there.
Organic identity shadow
Designed, disabled by default, and mandatory before any real traffic. When enabled for a tier, the gateway sends a seed-selected fraction r_sh = 0.02 (provisional) of organic dispatches synchronously to the reference processor, computes t = TV(p_i, p_ref) in memory, updates a per-miner, per-family, per-window accumulator, and drops payload and outputs. Nothing but the accumulator is persisted, and stored or asynchronous replay is forbidden.
At window close the validator receives only a gateway-signed verdict {miner, family, window, verdict ∈ pass | fail | unsupported}, with no counts. A miner fails if the mean identity test fails or if the share of items beyond per-item tolerance is credibly above 0.05 (provisional), which catches substitution on just part of the traffic. Fewer than 500 (provisional) shadowed items gives unsupported, which is never turned into pass.
Enabling it requires all of: the reference processor admitted for that tier, listed as a subprocessor in the DPA and admission workflow, covered by the System1 acceptance record, and a privacy review approving the aggregate outputs.
Verified capacity
At b_e = 4 (provisional) seed-derived times per epoch the gateway bursts every eligible miner at the same moment for T_b = 60 s (provisional), with fresh sealed items and a strict per-item deadline. Every dispatch is either a completion or a timeout. At window close one uniform audited sample of completions is compared with the reference, and
A dropped item costs as much as an out-of-tolerance answer, so skipping hard items does not help, and, under the spec's stated assumptions, a content-independent stream that is out of tolerance on at least half its items adds nothing in expectation. Faster honest hardware is genuine capacity and is not capped.
Not active. Until the prober, the pooled sample, the headroom budget and the adversarial capacity simulations exist, κ ≡ 1 and no externally operated miner receives weight on any non-local network. The current code fixes capacity at 1.
Operator groups and caps
- Groups. GDPR tier: the admitted
operator_id. Permissionless: the owning coldkey, merged only by proven linkage (same coldkey, same serving key, or a live challenge proving the same TLS key or origin). Correlations such as latency or duplicate payloads are monitoring signals only. Unknown off-chain identity is never by itself a reason for zero weight. - Group reward. With capacity proven, a group's reward is the sum of its members' applied gains; while
κ ≡ 1, it is the maximum member gain, so adding hotkeys adds nothing. - Water-filling. Shares
s_k = min(cap, λ*·G_k)with λ* chosen so the shares sum to 1 andcap = per_operator_cap = 0.40(provisional). Within a group, share is split in proportion to applied gain. - Abstain. With fewer than ⌈1/cap⌉ groups of positive reward, the validator abstains rather than concentrating weight.
- Linked groups. A newly proven link can never raise a member above what it would have received unlinked in the same epoch.
Any cap can be evaded by operators who split across unlinkable coldkeys. The spec states this residual openly rather than claiming Sybil resistance.
Fiat and emission separation
- The score ledger reads only validator evidence. It never reads prices, invoices, contract terms or settlements.
- Billing and settlement never read emissions, weights or balances.
- The customer router never reads the score, gain or weights. It picks among eligible miners using health, capacity, contract terms and a share cap.
- On chain there are only registration and weights: no customer identifiers, payloads, answers, labels or invoices.
What is still provisional
- All tolerances (
τ_id,τ_item,ε) until the honest-noise measurement per hardware configuration (M1) exists. - Reference latency, throughput and burst deadlines until M2; gateway burst headroom until M3; substitute-panel statistics until M4.
- Window sizes, error budgets, reuse limits, dispute budgets and the ramp until the S1–S4 simulations pass.
- Capacity activation and the split-neutrality target, which is explicitly not established.
- Exclusion after repeated failures, which stays disabled until the conditional-independence assumption has evidence.
- Multi-family scoring: any policy with two or more families abstains.
- Leaf-position privacy of the gateway's joint Merkle root, which has not been demonstrated.