A verified, tested personalisation skill

Prepare once, generate blindly, review the results yourself, and iterate only when your corrections show it is needed.

One main path, with bounded repair loops

1 · Freeze the benchmark
Lock 200 leads: 100 accepted unchanged and 100 pre-correction cases.
Success criteria
  • Every case has adequate website evidence and a stored hash.
  • The candidate renders inside “Láttam, hogy [p1] foglalkoztok.”
Constraints
  • No duplicates, leaked corrections, or Personalisation 2 scoring.
  • The set cannot change after testing starts.
Benchmark valid?200 cases · evidence · hashes · no leakage No
Repair the benchmark
Success: replace only invalid cases and rehash.
Constraint: do not inspect model outputs yet.
Yes 2 · Run both existing skills once
Run the current skill and recovered Claude skill on identical evidence.
Success criteria
  • Every case receives one output or a recorded provider failure.
  • Outputs are stored unchanged with model settings.
Constraints
  • No content retries, prompt edits, examples, or campaign uploads.
  • Retry infrastructure failures only, with the same inputs.
Run valid?all outputs or recorded provider failures No
Rerun failed cases
Success: provider failures are resolved or retained.
Constraint: inputs and prompt stay frozen.
Yes 3 · Publish the HTML table for Matt to review
Columns: Website URL · Scraped-content path · Personalisation · Greeting · Notes.
Success criteria
  • Every tested row appears; URLs and scrape paths are usable.
  • Personalisation and greeting are editable; Notes starts empty.
  • Greeting is “Szép napot” or “Kedves [firstname]”.
Constraints
  • No model judge or automatic approval; Matt is the reviewer.
  • Original values stay preserved for before/after scoring.
unfinished
Continue the same review
Success: every row is checked or corrected.
Constraint: do not replace the table mid-review.
Matt approves?≥190/200 unchanged · each group ≥90 · zero wrong claims Yes · no iteration No 4 · Improve the stronger skill
Fix observed error classes on separate development cases, then publish one new terminal review table.
Success criteria
  • Matt finishes the new table with ≥190/200 unchanged, each group ≥90, zero wrong claims.
  • The winning prompt and files are frozen.
Constraints
  • Maximum five development rounds or stop after three without improvement.
  • Locked cases cannot guide edits; only one terminal table.
Matt approves terminal table?same locked thresholds · one terminal review No
Stop with evidence
Success: preserve the best version, score, errors, and next hypothesis.
Constraint: do not promote an unverified skill.
Yes 5 · Release the exact tested skill
Install the frozen winner as the canonical personalisation skill.
Success criteria
  • Installed files match tested hashes; a clean invocation succeeds.
  • Benchmark, settings, verdicts, and adjudication are retained.
Constraints
  • No untested rewrite during installation.
  • No lead upload, campaign modification, or email sending.
Exact tested release?hash match · clean invocation · receipt saved No
Repair the installation
Success: restore exact files and rerun clean check.
Constraint: tested content cannot change.
Yes Done · verified skill released
Success: Matt-approved thresholds passed and installed hashes match.
Constraint: campaign activation remains separate.

SEND means true, representative, grammatical, and natural. REWRITE means the fact is usable but the wording is weak. WRONG means false or materially misleading. NO_EVIDENCE means the source cannot support a safe claim.