| case | <case name, e.g. bq-pipeline | gcp-deploy | sample-plan> | ||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| run_at | <UTC ISO timestamp, e.g. 2026-05-01T08:30:00Z> | ||||||||||||||||||||||||
| plan_file | <relative path to plan.md, e.g. examples/sample-plan.md> | ||||||||||||||||||||||||
| models |
|
||||||||||||||||||||||||
| timeout_sec | 180 | ||||||||||||||||||||||||
| notes | Output is non-deterministic. Findings below are from THIS specific run; your output may differ in wording but should match the pattern. |
Replace the metadata header above with your actual run values, then list the consolidated findings below.
| # | Finding | Teams | Severity |
|---|---|---|---|
| 1 | ... | Claude, Codex | must-fix |
| # | Finding | Team | Severity |
|---|---|---|---|
| 2 | ... | Gemini | should-fix |
- Claude says X; Codex says Y. Both valid; team chooses.
- Dimension N had < 2 concrete scenarios across all teams.
- ...