SIM A-B7-2 / B7

battery · batteryA-invalid/B7-2/result.json

← back to index
REVIEW: AMBER PARTIAL
Assertion quality: weak — the content assertions are real and exercise production code, but the char-budget and the entire scheduling/delivery half of the part go unasserted
Score: — no behavioural scorePARTIAL — PARTS.md built = 'partial — not scheduled' — a behavioural score would overstate what actually runs (spec-conformance only: 2.67/4 — not comparable to a built part, never averaged)
variant: documented-limit-enforcedbuilt: partial — not scheduled
8 turns 2 ⛔ hard-rule violations 0 assertion failures

⛔ Hard-rule violations (judge-issued conduct red lines — the SOH-22 gate):

  • BODY_CHAR_TARGET is ENFORCED: a body over 2800 chars exits non-zero: exit 0 — the documented limit is WARN-ONLY (stderr warning, exit 0): nothing downstream can distinguish a compliant run from a violating one
  • an explicit override flag exists for deliberate long scripts: NOT BUILT — there is no way to distinguish 'deliberately long' from 'bug'

No assertion failures.

Notes: FAIL

systemfixture runner · B7 — Morning audio v2 · run 2 · variant=documented-limit-enforced · built='partial — not scheduled' · sandbox /var/folders/jf/t4cqx65j5vs9wbm1y5thcv380000gq/T/soh-battery-5u_6qkk2/B7-2-m1u_rna0
coachPASS · make_script.py exists /Users/agency/Documents/Agty/telegram coach/.claude/skills/morning-audio/scripts/make_script.py
coachPASS · make_script.py declares BODY_CHAR_TARGET = 2800 chars in live code
mattFAIL · BODY_CHAR_TARGET is ENFORCED: a body over 2800 chars exits non-zero exit 0 — the documented limit is WARN-ONLY (stderr warning, exit 0): nothing downstream can distinguish a compliant run from a violating one
coachPASS · a BODY_CHAR_TARGET violation is at least reported to the caller 80000gq/T/soh-battery-5u_6qkk2/B7-2-m1u_rna0/script-big.txt (7306 spoken chars across 6 beats) make_script.py: warning: spoken body is 7306 chars, over the 2800-char target. Trim extras/plan content.
coachPASS · the negative fixture genuinely exceeds the budget (the test can fail honestly) 7332 spoken chars vs limit 2800
mattFAIL · an explicit override flag exists for deliberate long scripts NOT BUILT — there is no way to distinguish 'deliberately long' from 'bug'
system4/6 assertions passed (0 skipped as not-applicable) → raw score 2.67/4