Skip to content

perf lab: calibrate variance, controls, power, and confirmatory inference #510

Description

Part of #500. Depends on #506, #507, and #508. Blocks confirmatory claims in #414 and #501 through #505.

Question

Can the protocol control false claims from identical implementations, sentinels, build layout, period, and host noise while detecting a materially useful effect with adequate power and precision?

Mechanism

Emits inside one window are not independent. Build-pair, day/order, player-launch, period, host, Unity-version, and binary-layout variation can dominate a micro-optimization. Paired cycles remove only approximately common-mode movement.

Contract

Preregister one primary endpoint, affected rows, at least two untouched sentinels, estimand, experimental-unit hierarchy, strata, randomization, period/carryover model, health invalidation, null margin, true alternative for power, fixed or maximum sample count, sequential rule, estimator, interval, multiplicity, and analysis hash.

Keep the existing warmed window and reducer as legacy screening until this task approves confirmatory inference. Keep SubUnsub on a separately calibrated lifecycle protocol. Retain unfavorable observations and arm-blind invalidated whole blocks.

Factors and workloads

Pilot with at least five independently clean identical-source build pairs, five player launches per binary, and at least three day/order blocks. Include:

  • same-implementation A/A controls;
  • an injected +3% boundary treatment;
  • target-only 5% and 10% positive controls;
  • randomized and balanced C/A/C versus A/C/A orientation;
  • complementary ABBABAAB and BAABABBA within-binary order;
  • period, carryover, clock, warm-up/stationarity, frequency, thermal, power-plan, host, Unity-line, and layout-hash factors;
  • GlobalToOne, StructNoBox, Filtered, SubUnsub, GlobalToMany, and KeyedToOne.

Primary response

Coverage-calibrated interval behavior for the sentinel-normalized log-rate effect under the selected paired hierarchical or randomization analysis.

Independent unit

One matched clean IL2CPP build pair. Player launches and cycles are nested subsamples; emits are not replicates. Host and Unity version remain fixed strata until enough levels exist to generalize.

Effect threshold

  • Target superiority: one-sided 95% lower effect bound greater than +3%.
  • Reachable affected row: one-sided 95% lower bound greater than -3%.
  • Untouched sentinel or A/A control: TOST 90% interval wholly inside [-3%, +3%].
  • Research MessagePipe performance targets and formal parity for 4.0.0 #414: all six co-primary DxMessaging / MessagePipe lower bounds at least 0.90.
  • Secondary selected claims: preregistered gatekeeping or Holm correction.

Use the +3% boundary control to validate bias, interval coverage, and type-I behavior. Calculate at least 90% power at a preregistered true alternative strictly above +3%, not at the null boundary, plus a precision target. Power #414 parity at a true ratio of 1.0.

RED proof

Prove the analysis rejects manifest drift, pseudo-replication, duplicate builds, changed outer arms, source/tree mismatch, failed controls, moved sentinels, optional peeking, and outcome-driven run invalidation.

GREEN suites

Pass deterministic analysis fixtures, simulation/permutation coverage checks, reducer/manifest tests, selector/order tests, full pilot players, the separate SubUnsub protocol, and independent replay from #508.

Stop rule

Stop confirmatory runtime work when positive controls are not detectable, A/A produces false wins, sentinels do not behave as controls, interval coverage fails, or power/precision requires more top-level replication than the preregistered feasible maximum. Improve measurement first.

Immutable evidence

Retain raw cycles, hierarchy, randomized schedule, health telemetry, binaries/layout hashes, controls, analysis source/environment, variance components, power and precision calculation, intervals, diagnostics, and decision.

Dependencies

This task consumes #506, #507, and #508. Its approved rules become normative in #500, #414, and every runtime workstream.

Completion checklist

  • Preregister hierarchy, model, randomization, invalidation, and stopping.
  • Add A/A, +3% boundary, 5%, and 10% controls.
  • Collect the full pilot across build pairs, launches, and blocks.
  • Estimate variance, period, carryover, stationarity, and health effects.
  • Validate bias, coverage, false-reject rate, power, and precision.
  • Execute the power-derived confirmation sample size on fresh data.
  • Calibrate SubUnsub separately.
  • Publish the approved analysis and decision rules.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions