A verified, tested personalisation skill

One locked test, one independent review, and iteration only if the evidence says it is needed.

The simplest complete verification process

dataset ready valid run confirmed result candidate ready 1 · Freeze the definition and test set
Build a locked set of 200 leads: 100 accepted unchanged and 100 corrected, using only cases with enough website evidence.
Constraints
  • The model cannot see corrections, labels, or expected answers.
  • Personalisation 2 is outside this test because it was not reviewed.
  • Do not alter the set after generation starts.
  • Do not require exact wording when several claims are equally good.
READY · 200-case benchmark is locked
Success verification criteria
  • Exactly 100 accepted and 100 pre-correction cases.
  • Every case renders as “Láttam, hogy [p1] foglalkoztok.”
  • Every source and record has a stored hash.
  • Allowed verdicts are SEND, REWRITE, WRONG, and NO_EVIDENCE.
NOT READY · repair the dataset first
Success verification criteria
  • Missing evidence, duplicate companies, leaked labels, or an incorrect group count is detected.
  • The affected cases are replaced before any model is scored.
  • The corrected 200-case manifest is hashed again.
2 · Run both existing skills once
Run the current skill and the recovered Claude skill on the same locked evidence. Render each candidate inside the real Hungarian sentence before judging.
Constraints
  • No retries for weak content; retry infrastructure failures only.
  • No prompt edits, examples, or tuning during the test.
  • No skill identity or expected answer shown to the judge.
  • No campaign upload or live sending.
PASS · at least one skill clears the locked test
Success: ≥190/200 SEND, ≥90/100 in each group, zero WRONG.
BELOW TARGET · neither skill clears the test
Success: all failures retain evidence, output, verdict, and reason for adjudication.
INVALID RUN · rerun only affected cases
Success: provider errors are separated from content failures and corrected without changing inputs.
3 · Independently adjudicate the score
A fresh judge reviews every non-SEND case plus a random sample of SEND cases. It sees the website evidence and candidate only.
Constraints
  • The judge cannot see the first verdict, correction, skill name, or total score.
  • Alternative truthful, representative phrasing must be accepted.
  • Style preference alone cannot turn SEND into REWRITE.
  • Disagreements must cite the failed property: truth, relevance, grammar, or naturalness.
CONFIRMED PASS · release candidate selected
Success: adjudicated score still meets all three pass thresholds.
CONFIRMED BELOW TARGET · improve the stronger skill
Success: the stronger baseline and its real error categories are identified.
UNRELIABLE EVALUATION · fix rubric, then rejudge
Success: inconsistent judgments are reconciled before any prompt is changed.
4 · Improve only when below target
Use separate development cases to repair the strongest skill. Run one terminal evaluation on the still-locked 200 cases when development is done.
Constraints
  • The locked test cases cannot train or guide prompt edits.
  • Maximum five development rounds, or stop after three rounds without improvement.
  • Do not weaken truth or evidence rules to increase the score.
  • Do not rerun the terminal test to hunt for a lucky score.
TARGET REACHED · terminal test passes
Success verification criteria
  • ≥190/200 SEND after independent adjudication.
  • Each 100-case group remains ≥90% SEND.
  • Zero WRONG outputs.
  • Prompt and supporting files are frozen for release.
TARGET NOT REACHED · stop with evidence
Success verification criteria
  • The best measured version, score, and remaining error classes are preserved.
  • No unverified version is promoted.
  • The exact next hypothesis is documented if another cycle is justified.
5 · Release the exact tested skill
Install the winning frozen version as the canonical personalisation skill and retain its benchmark receipt for future regressions.
Constraints
  • The released files must exactly match the tested hashes.
  • No untested cleanup or prompt rewrite during installation.
  • The skill cannot upload leads, modify campaigns, or send email during verification.
  • Future changes require a new version and a fresh locked test.
RELEASED · verified skill is discoverable and reproducible
Success verification criteria
  • Canonical skill files match the winning test hashes.
  • A clean invocation succeeds on a held example.
  • The benchmark manifest, model settings, verdicts, and adjudication are saved.
NOT RELEASED · keep the current production skill
Success verification criteria
  • A hash mismatch, failed clean invocation, or incomplete receipt blocks promotion.
  • The tested candidate remains preserved for repair.
  • The existing installed skill remains unchanged.
released threshold not met Stop · evidence says not ready
Success verification criteria
  • No weak or unverified version reaches production.
  • Best score and exact failure evidence are retained.
Constraint: no campaign use.
Done · verified skill released
Success verification criteria
  • Adjudicated ≥95%, both groups ≥90%, and zero WRONG.
  • Installed skill exactly matches the tested winner.
Constraint: campaign activation remains a separate action.

SEND means the statement is true, representative of the company, grammatically correct in the complete sentence, and natural enough to send. REWRITE means the core fact is usable but the wording is weak. WRONG means false or materially misleading. NO_EVIDENCE means the source cannot support a safe claim.