REAX.docs Pre-release · not on testnet
MenuScoring & incentives

Reference

Scoring and incentive mechanism

REAX rewards one thing: serving the approved model faithfully, reliably, at reference speed, with real capacity. This page explains in plain words how validators turn evidence into weights, then gives the formulas from spec revision 0.6.

Provisional mechanism

Spec 0.6 is not frozen. Every constant marked (provisional) lives in signed policy and can change; tolerances are placeholders until honest-noise measurements exist. No weights are set on any network today. Emissions are an incentive mechanism only: they are not customer billing, not a service guarantee and not a promise of any return.

In plain words

  1. Serve the right model. Validators compare your probabilities with the reference model's on the exact same bytes. If you consistently drift, you fail a hard gate and earn nothing for the affected epochs. Beating the reference's accuracy earns nothing extra; falling measurably below it fails the quality gate.
  2. Be there. Availability multiplies everything. Below 50 % (provisional) you are not eligible.
  3. Be as fast as the reference. Latency within 10 % (provisional) of the reference manifest gets full credit. Faster earns nothing extra, so a cheaper, faster substitute gains nothing.
  4. Bring real capacity. Verified capacity, measured in synchronized audited bursts, scales your weight. In the spec's models, declaring more concurrency than you have does not help.
  5. One operator, one share. Hotkeys are grouped by operator, and no group can take more than 40 % (provisional) of the weight.

The gain formula

For each miner i and epoch, over the set Φi of question families its model alias covers:

g_i = F_i · A_i · (0.8 + 0.2 · L_i) · κ_iF_i = Π_{f ∈ Φ_i} 1[not fidelity_failed_{i,f}]L_i = Σ_{f ∈ Φ_i} w_f · L_{i,f}

There is no quality score, cost term or identity score inside g. Identity and quality act only as hard gates and through manifest admission. Two miners that both pass the gates differ only by availability, latency and capacity.

Ramp

g^app_i(e) = min( g_i(e), g^app_i(e−1) + ρ · g_i(e) ), ρ = 0.5 (provisional)

Increases are gradual and decreases apply at once. A miner in fidelity_watch cannot increase. Other miners joining or leaving change only the normalization.

Fidelity gates

Two statistics compare a miner with the reference output p_ref for the same item:

identity: t_i(x) = TV( p_i(x), p_ref(x) ) ∈ [0, 1] no label neededquality: d_i(x) = B( p_i(x), y ) − B( p_ref(x), y ) ∈ [−1, 1]Brier: B = 1 − 0.5 · Σ_k (p_k − y_k)²

At window close each test uses a one-sided empirical-Bernstein bound (Maurer and Pontil 2009) on the window mean:

r_EB(n, V, b, δ) = √(2·V·ln(2/δ)/n) + 7·b·ln(2/δ) / (3·(n − 1))identity_fail ⇔ mean t − r_EB(n_id, V, 1, δ_test) > τ_idquality_fail ⇔ mean d + r_EB(n_lab, V, 2, δ_test) < −εburst_audit_fail is the identity test on the pooled window capacity sample; organic_identity_fail comes from the shadow when it is enabled.
ParameterValue
Identity window size n_id,win2 000 (provisional)
Labelled window size n_lab,win288 (provisional)
Burst audit sample n_b,win2 016
Per-window error budget δ_w0.005 (provisional), split into a fixed δ_test = 0.000625 per test and a separate tripwire budget
Tolerances τ_id, τ_item, ε0.05, 0.04, 0.02: placeholders until measured per hardware configuration

Code vs spec

The current validator code uses 288-cluster windows for both statistics and δ = 0.01; the 2 000-item identity window is the spec 0.6 target.

Two strikes. The first failing window puts the miner in fidelity_watch: its applied gain is frozen and the next window uses only never-exposed items. A second consecutive failing window, on disjoint evidence, makes it fidelity_failed for the affected epochs: F = 0 and κ = 0. Exclusion from the network is disabled in this revision, because the independence assumption it would rest on is not yet proven; the miner is simply re-evaluated in the next window.

Support. A miner with too few items in its first window is unsupported: no weight and no penalty. A miner on a hardware configuration that has not been calibrated is also unsupported, never penalized.

Model admission

Before any miner is scored on a manifest, the manifest itself must pass two checks per question family:

Substitutes that stay within tolerance on most items are output-equivalent: they earn like the reference, and the catalog publishes every known one. Output fingerprinting never proves which weights are running; only separately evaluated, nonce-bound hardware attestation could do that.

Availability

A_i = miner-attributable successes / miner-attributable scheduled attemptsPooled over supported families in the window. Eligibility requires A_i ≥ 0.50 (provisional).

Counted as failures: timeouts after dispatch, connection refused or reset, invalid or late responses, wrong identity or signature. Excluded: burst dispatches, pre-dispatch expiry, validator or gateway faults and injection-API failures, windows where the gateway's egress health probe failed, chain or RPC faults, revoked_in_flight and payload_binding_mismatch.

Latency, capped at the reference

f_q(t) = clamp( 1 − max(0, t − (1 + η) · t_ref,q) / D_f , 0, 1 ) q ∈ {50, 95}L_{i,f} = 0.5 · f_50(p50_i) + 0.5 · f_95(p95_i)η = 0.10 (provisional); D_f = t_ref95 (provisional). A timeout enters as t = the request deadline.

Every miner whose p50 and p95 are within (1 + η) of the reference receives L = 1, and nothing below that earns more. A substitute that is faster than the reference therefore gains nothing over an honest miner at reference latency. Latency is measured by the serving region's gateway on a monotonic clock, from dispatch to verified receipt.

Worked example from the spec's acceptance vectors (tref50 = 400 ms, tref95 = 1 000 ms, η = 0.1): p50 = 300 ms and p95 = 900 ms gives L = 1.0; p50 = 640 ms and p95 = 1 600 ms gives f50 = 0.8, f95 = 0.5, L = 0.65. These are test values, not measured reference latencies.

Code vs spec

The current validator code uses a simpler transform, L = 0.5·max(0, 1 − p50/deadline) + 0.5·max(0, 1 − p95/deadline) against a signed per-family deadline, which is not capped at the reference. The capped f_q above is the spec 0.6 target.

Hidden probes

Validators do not announce which requests are tests. On the GDPR tier, probes enter through the gateway with the same caller ID, TLS configuration, egress addresses and connection pooling as organic traffic; their deadlines, arrival times and field shapes follow privacy-reviewed organic profiles; and every receipt carries the same fields, so receipts do not reveal probe status. Identity items are used exactly once and labelled items are never shown twice to the same operator.

What is not assumed: that probe content is indistinguishable from real content. A generator's questions can differ from customers' questions, which is exactly why the organic identity shadow is mandatory before real traffic. On the permissionless route every dispatch is synthetic, so indistinguishability is not claimed there.

Organic identity shadow

Designed, disabled by default, and mandatory before any real traffic. When enabled for a tier, the gateway sends a seed-selected fraction r_sh = 0.02 (provisional) of organic dispatches synchronously to the reference processor, computes t = TV(p_i, p_ref) in memory, updates a per-miner, per-family, per-window accumulator, and drops payload and outputs. Nothing but the accumulator is persisted, and stored or asynchronous replay is forbidden.

At window close the validator receives only a gateway-signed verdict {miner, family, window, verdict ∈ pass | fail | unsupported}, with no counts. A miner fails if the mean identity test fails or if the share of items beyond per-item tolerance is credibly above 0.05 (provisional), which catches substitution on just part of the traffic. Fewer than 500 (provisional) shadowed items gives unsupported, which is never turned into pass.

Enabling it requires all of: the reference processor admitted for that tier, listed as a subprocessor in the DPA and admission workflow, covered by the System1 acceptance record, and a privacy review approving the aggregate outputs.

Verified capacity

At b_e = 4 (provisional) seed-derived times per epoch the gateway bursts every eligible miner at the same moment for T_b = 60 s (provisional), with fresh sealed items and a strict per-item deadline. Every dispatch is either a completion or a timeout. At window close one uniform audited sample of completions is compared with the reference, and

net_i(w) = C_{i,w} · (1 − 2 · k_{i,w} / n_{i,w}) − T_{i,w}κ̃_i(w) = max(0, net_i(w)) / (T_{b,w} · thr_ref)C = completions, T = timeouts, k of n audited completions outside per-item tolerance, thr_ref = measured reference throughput. Defined only with at least 336 (provisional) audited completions; otherwise capacity is insufficient.

A dropped item costs as much as an out-of-tolerance answer, so skipping hard items does not help, and, under the spec's stated assumptions, a content-independent stream that is out of tolerance on at least half its items adds nothing in expectation. Faster honest hardware is genuine capacity and is not capped.

Not active. Until the prober, the pooled sample, the headroom budget and the adversarial capacity simulations exist, κ ≡ 1 and no externally operated miner receives weight on any non-local network. The current code fixes capacity at 1.

Operator groups and caps

Any cap can be evaded by operators who split across unlinkable coldkeys. The spec states this residual openly rather than claiming Sybil resistance.

Fiat and emission separation

What is still provisional