Same useful work · result loading
Result loading
The observatory will state whether adaptive switching preserved held-out learning and beat the frozen fixed comparator.
Infrastructure remains modeled unless the artifact says otherwise.E001-SC1 Observable Semantic Slack
Loading the held-out semantic-consistency result. No conclusion is shown until its evidence boundary is known.
Reading data/e001-semantic-consistency-v1.json…
E001-SC1 · held-out software experiment
Same useful work · result loading
The observatory will state whether adaptive switching preserved held-out learning and beat the frozen fixed comparator.
Infrastructure remains modeled unless the artifact says otherwise.Every policy receives the same training examples and must reach the same useful-work target. The result will show whether changing consistency mode helped, what it saved, and where the controller refused to guess.
The fixed comparator is selected on calibration only. Six untouched stress families pair adaptive and fixed policies against a hindsight whole-policy envelope; sampling intervals and infrastructure uncertainty remain separate.
Exact optimizer commits, mode transitions, membership, lineage, work, WAN accounting, assumptions, uncertainty, and missing evidence appear below.
Paired held-out effects
Epistemic ranking regions
Each held-out stress family is divided into labeled infrastructure-uncertainty regions where adaptive wins, a fixed policy wins, ranking reverses, or the controller abstains.
| Family | Uncertainty region | Ranking | Comparator | Reason |
|---|
Selected held-out family
Aligned adaptive and comparator mode intervals with membership, WAN, abstention, merge, and rejoin events.
Untouched family ledger
| Family | Learning delta | Completion ratio | WAN ratio | Work replayed | Energy | Ranking / abstention |
|---|
| Wall / logical tick | Mode / action | Stress / membership | Lineage | Update age / disagreement | Useful / attempted / replayed | WAN payload / time | Held-out NLL | Evidence / abstention |
|---|
not loaded
Raw optimizer-commit trace has not been loaded. Opening this disclosure in Full trace fetches the separately bound artifact.
Aligned experiment time
Matched recovery trace
The same failure hits all four policies. The shared clock shows what stops, what is restored, what must be redone, and when useful work catches up.
All policy tracks share the artifact time domain. Recovery completion is mechanical: each policy reaches the same durable frontier after explicit preemption, restore, replay, and membership transitions.
Every mark is projected from a persisted recovery episode. Event IDs, nanosecond bounds, work dispositions, checkpoint lineage, and matched-frontier hashes remain available in the structured trace.
Measured small-model learning
LC1 · paired interrupted calibration
The observatory will state the measured answer once the persisted artifact is available.
Lower held-out loss means the model predicts better. The cards show how much training each policy attempted to get there, and the lines show what it learned over the same failure clock.
Policy medians summarize matched strata. The paired effect subtracts fixed-local progress per FLOP from adaptive progress per FLOP within each stratum, so zero means no retained-efficiency difference.
Artifact, dataset, runtime, paired evaluation, and per-run records remain attached below.
Median trajectory
Paired causal contrast
Paired interval not loaded.
Preregistered decision
| Stratum | Fixed progress/FLOP | Adaptive progress/FLOP | Adaptive − fixed |
|---|
| Run | Stratum | Final NLL | Attempted tokens | Energy | Active time |
|---|
not loaded
LC3 · measured equal-work endpoint
Same 524,288 useful tokens · six held-out pairs
The observatory will state the result after loading the persisted LC3 artifact.
Useful work is the fair ruler here. Both policies finish with the same 524,288 tokens actually kept by the model. Adaptive avoided redo work and finished the schedule sooner without a meaningful learning loss, but its GPU consumed too much energy to pass.
LC3 removes LC1’s stopping-early denominator trap. Six paired evaluation strata share an equal canonical-work endpoint. Read learning as paired NLL noninferiority, physical work as attempted-FLOP savings, schedule time as opportunity ticks, and energy as a paired device-energy ratio.
Exact pairs, run records, protocol predecessors, hashes, and the modeled-mechanics boundary remain attached below.
Observed policy medians
Paired 90% intervals
Frozen LC3 decision
| Stratum | Fixed NLL | Adaptive NLL | Adaptive − fixed NLL | Attempted-FLOP saved | Ticks saved | Energy ratio |
|---|
| Run | Split | Policy | Interrupted | Final NLL | Attempted / canonical tokens | Opportunity ticks | Device energy | Checkpoints |
|---|
| Artifact path | Conclusion | SHA-256 |
|---|
not loaded
E002-PW1 · measured local power
32 real GPU runs · measurement rejected
The observatory will state the measurement boundary after loading the persisted artifact.
No policy claim is available yet.The experiment ran; the measuring stick failed. All 32 planned comparisons started from the exact same trained model. Asking the sensor every 20 milliseconds did not make it answer that fast: its reading effectively changed about every 495 milliseconds. That is too slow to decide whether sparse continuation saves energy.
The 2×2 cadence-by-continuation design remains intact, but its energy estimand is inadmissible. Both frozen update-count invalidators fired and the logger calibration hit its boundary. Arm medians and contrasts remain visible as diagnostics, never as mechanism evidence.
The 32-run ledger, per-phase metrics, bound source identities, and lazy raw telemetry are attached below.
Validity before effect size
Observed arm medians
No-failure controls
Raw diagnostic contrasts
| Run | Split / block | Arm | Attempted / canonical | Final NLL | Samples / effective updates | Raw board energy | Checkpoints | Raw trace SHA-256 |
|---|
Phase energy is retained for diagnosis, but remains inadmissible for a causal policy comparison because the logger invalidators fired.
| Phase | Duration | Idle-subtracted board energy | Energy / canonical token | Effective update equivalents |
|---|
not loaded
Raw telemetry has not been loaded. Opening this disclosure in Full trace fetches the separate point artifact.
| Point | Time | Phase at timestamp | Board power | GPU util. | Memory util. | SM / memory clock | Temperature | P-state |
|---|
E002-PW2 · valid cumulative energy
32 real GPU runs · measurement valid
The observatory will state the mechanism after loading the persisted artifact.
Mechanism estimate loading.Checkpointing less often fixed the measured energy problem without giving back the recovery win. Both policies kept exactly 524,288 useful tokens. Sparse continuation attempted less work, finished 40 scheduling ticks sooner, and stayed inside the frozen GPU-energy limit. An estimated idle-baseline subtraction was less certain and could not rule out zero interaction.
PW2 repeats PW1’s frozen 2×2 design with a supported cumulative-energy counter. No measurement invalidator fired. Positive total and checkpoint-group interactions pass the frozen primary attribution gates while sparse continuation preserves the learning, work, schedule, and energy gates. The sensitivity-only idle-subtracted interaction crossed zero, so the result is not insensitive to baseline treatment.
The 32-run ledger, selected-run counter phases, bound identities, and lazy raw counter points are attached below.
Measurement validity
Paired causal contrasts
Counter support by phase
Frozen decision
| Run | Split / block | Arm | Attempted / canonical | Final NLL | Polls / counter updates | Run energy | Checkpoints | Counter trace SHA-256 |
|---|
| Phase | Evidence class | Duration | Energy | Energy / canonical token | Counter updates | Effective update equivalents |
|---|
not loaded
Raw counter telemetry has not been loaded. Opening this disclosure in Full trace fetches the separate point artifact.
| Point | Time | Phase at timestamp | Cumulative board energy | Counter update | Ancillary board power | GPU util. | Temperature |
|---|
E002-PW3 · physical rack mechanism
Physical result loading
The observatory will explain what moved, what the rack meter saw, and whether learning stayed fixed.
No rack result is inferred from the earlier laptop experiment.The same training and recovery work runs several ways. The shaped run may move only operations that have real timing slack, then the rack meter tells us whether separating those operations actually reduced the electrical shock.
Paired blocks hold useful work, state generations, failures, and learning commitments fixed while changing only the legal release policy for checkpoint and rejoin flows.
Exact event intervals, clock alignment, sensor coverage, semantic obligations, and chunk hashes appear below.
Same clock · measured boundary
Paired rack-PDU power traces aligned above job-level checkpoint and recovery event rails.
Paired physical effects
Five frozen policies
| Policy | Rack ramp | 0.1–10 Hz energy | Useful-token rate | Rack J / token | p95 recovery | Held-out NLL |
|---|
Preregistered decision
| Arm / job | Operation | Generation | Ready → release | Actual interval | Bytes | Outcome |
|---|
not loaded
E001 v1 screening artifact
| Policy | Local steps / decision | Progress per FLOP | Inter-site bytes | Time to target | Base + compute energy | Falsifier status |
|---|
Waiting for a generated artifact. Result cells remain not run.
E001 v1 screening evidence chain
Source observationsUNMEASURED · UNAVAILABLE FROM ARTIFACT
Unfitted sensitivity priorPRIOR · NOT FITTED
Progress per FLOPUNMEASURED
Held-out time to targetUNMEASURED
Published source measurements are read from the observatory artifact. Publication-rounding intervals remain distinct from run-to-run variance.
| Observation | Value | Evidence | Action |
|---|
The attached literature records cover a narrow delay setting. They do not identify progress per FLOP, longer local-update intervals, frontier-scale transfer, multi-site interruption behavior, or an active-outage controller.
Repeated small-model delay calibrationMeasure multiple delay intervals and optimizers rather than extrapolating one step.
Held-out optimizer, model, and site combinationsEvaluate combinations excluded from prior construction.
Controlled 30B to 100B-plus multi-site runVary delay and cadence under a defined policy with identical evaluation accounting.
| After | Observed state | Decision | Applies to | Evidence |
|---|
not run