You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Can the protocol control false claims from identical implementations, sentinels, build layout, period, and host noise while detecting a materially useful effect with adequate power and precision?
Mechanism
Emits inside one window are not independent. Build-pair, day/order, player-launch, period, host, Unity-version, and binary-layout variation can dominate a micro-optimization. Paired cycles remove only approximately common-mode movement.
Contract
Preregister one primary endpoint, affected rows, at least two untouched sentinels, estimand, experimental-unit hierarchy, strata, randomization, period/carryover model, health invalidation, null margin, true alternative for power, fixed or maximum sample count, sequential rule, estimator, interval, multiplicity, and analysis hash.
Keep the existing warmed window and reducer as legacy screening until this task approves confirmatory inference. Keep SubUnsub on a separately calibrated lifecycle protocol. Retain unfavorable observations and arm-blind invalidated whole blocks.
Factors and workloads
Pilot with at least five independently clean identical-source build pairs, five player launches per binary, and at least three day/order blocks. Include:
same-implementation A/A controls;
an injected +3% boundary treatment;
target-only 5% and 10% positive controls;
randomized and balanced C/A/C versus A/C/A orientation;
complementary ABBABAAB and BAABABBA within-binary order;
GlobalToOne, StructNoBox, Filtered, SubUnsub, GlobalToMany, and KeyedToOne.
Primary response
Coverage-calibrated interval behavior for the sentinel-normalized log-rate effect under the selected paired hierarchical or randomization analysis.
Independent unit
One matched clean IL2CPP build pair. Player launches and cycles are nested subsamples; emits are not replicates. Host and Unity version remain fixed strata until enough levels exist to generalize.
Effect threshold
Target superiority: one-sided 95% lower effect bound greater than +3%.
Reachable affected row: one-sided 95% lower bound greater than -3%.
Secondary selected claims: preregistered gatekeeping or Holm correction.
Use the +3% boundary control to validate bias, interval coverage, and type-I behavior. Calculate at least 90% power at a preregistered true alternative strictly above +3%, not at the null boundary, plus a precision target. Power #414 parity at a true ratio of 1.0.
RED proof
Prove the analysis rejects manifest drift, pseudo-replication, duplicate builds, changed outer arms, source/tree mismatch, failed controls, moved sentinels, optional peeking, and outcome-driven run invalidation.
GREEN suites
Pass deterministic analysis fixtures, simulation/permutation coverage checks, reducer/manifest tests, selector/order tests, full pilot players, the separate SubUnsub protocol, and independent replay from #508.
Stop rule
Stop confirmatory runtime work when positive controls are not detectable, A/A produces false wins, sentinels do not behave as controls, interval coverage fails, or power/precision requires more top-level replication than the preregistered feasible maximum. Improve measurement first.
Immutable evidence
Retain raw cycles, hierarchy, randomized schedule, health telemetry, binaries/layout hashes, controls, analysis source/environment, variance components, power and precision calculation, intervals, diagnostics, and decision.
Dependencies
This task consumes #506, #507, and #508. Its approved rules become normative in #500, #414, and every runtime workstream.
Completion checklist
Preregister hierarchy, model, randomization, invalidation, and stopping.
Add A/A, +3% boundary, 5%, and 10% controls.
Collect the full pilot across build pairs, launches, and blocks.
Estimate variance, period, carryover, stationarity, and health effects.
Validate bias, coverage, false-reject rate, power, and precision.
Execute the power-derived confirmation sample size on fresh data.
Part of #500. Depends on #506, #507, and #508. Blocks confirmatory claims in #414 and #501 through #505.
Question
Can the protocol control false claims from identical implementations, sentinels, build layout, period, and host noise while detecting a materially useful effect with adequate power and precision?
Mechanism
Emits inside one window are not independent. Build-pair, day/order, player-launch, period, host, Unity-version, and binary-layout variation can dominate a micro-optimization. Paired cycles remove only approximately common-mode movement.
Contract
Preregister one primary endpoint, affected rows, at least two untouched sentinels, estimand, experimental-unit hierarchy, strata, randomization, period/carryover model, health invalidation, null margin, true alternative for power, fixed or maximum sample count, sequential rule, estimator, interval, multiplicity, and analysis hash.
Keep the existing warmed window and reducer as legacy screening until this task approves confirmatory inference. Keep SubUnsub on a separately calibrated lifecycle protocol. Retain unfavorable observations and arm-blind invalidated whole blocks.
Factors and workloads
Pilot with at least five independently clean identical-source build pairs, five player launches per binary, and at least three day/order blocks. Include:
ABBABAABandBAABABBAwithin-binary order;Primary response
Coverage-calibrated interval behavior for the sentinel-normalized log-rate effect under the selected paired hierarchical or randomization analysis.
Independent unit
One matched clean IL2CPP build pair. Player launches and cycles are nested subsamples; emits are not replicates. Host and Unity version remain fixed strata until enough levels exist to generalize.
Effect threshold
DxMessaging / MessagePipelower bounds at least 0.90.Use the +3% boundary control to validate bias, interval coverage, and type-I behavior. Calculate at least 90% power at a preregistered true alternative strictly above +3%, not at the null boundary, plus a precision target. Power #414 parity at a true ratio of 1.0.
RED proof
Prove the analysis rejects manifest drift, pseudo-replication, duplicate builds, changed outer arms, source/tree mismatch, failed controls, moved sentinels, optional peeking, and outcome-driven run invalidation.
GREEN suites
Pass deterministic analysis fixtures, simulation/permutation coverage checks, reducer/manifest tests, selector/order tests, full pilot players, the separate SubUnsub protocol, and independent replay from #508.
Stop rule
Stop confirmatory runtime work when positive controls are not detectable, A/A produces false wins, sentinels do not behave as controls, interval coverage fails, or power/precision requires more top-level replication than the preregistered feasible maximum. Improve measurement first.
Immutable evidence
Retain raw cycles, hierarchy, randomized schedule, health telemetry, binaries/layout hashes, controls, analysis source/environment, variance components, power and precision calculation, intervals, diagnostics, and decision.
Dependencies
This task consumes #506, #507, and #508. Its approved rules become normative in #500, #414, and every runtime workstream.
Completion checklist