End result and current gate
- Invoke /get-leads in a fresh task. It understands the conversation, project files, existing inventory and prior runs, then opens a useful method-review page. After review, it sources when needed, scrapes, checks fit and personalises. Results and edits remain available across Matt's devices in one hosted run page and history index.
- Now: review this plan only. No builder/trial launch, skill archive, lead processing or paid work starts until Matt explicitly acknowledges that he has reviewed the plan. An annotation, page view, silence or absence never authorizes execution.
- Done later: the installed skill works independently in a pinned Sol Medium task on the first 100 source candidates. Its demonstrated friction is fixed and retested. Templates and screenshots alone do not prove completion.
Simple execution phases, after Matt's review
- Build and prove usability. The parent opens this outcome contract as a localhost dashboard in this chat's in-app browser/sidebar, then gives a new visible Astra High task
/goaland the exact contract. Original contract text is black. Runner-owned progress, evidence captions and results are blue. Cards preserve scroll and edits while refreshing and include the screenshot placeholders below. Astra owns implementation, recovery and verification, choosing tools and bounded subagents, usually Luna Max. - Check the result, then trial it independently. The parent checks Astra's delivered evidence and returns only failed criteria to the same owner. Once usable, it creates and pins a fresh Sol Medium task with
/goal, a source-backed brief and the installed/get-leadsentry. Sol uses the ChatGPT browser to fill/review its generated page and run scrape → fit → personalise on the frozen first 100. No implementation coaching is needed. - Close demonstrated friction. Sol reports run/version, affected identity, expected/actual behavior and evidence. Astra owns fixes in a safe revision after active ownership is coordinated. Sol reruns the affected behavior against the same manifest and returns evidence. The parent accepts outcomes. Unresolved failures remain explicit, with their recovery/resume point, rather than being called complete.
The parent uses passive completion/blocker waits. It makes one status check 20 minutes after Sol launches, if Sol is still running, and may give one bounded intervention. Otherwise it does not actively supervise. A reported blocker can receive a scoped alternative. Silence is not proof of a stall. Runner-produced updates keep the dashboard useful without repeated parent polling.
Six deliverables
D1 · One usable entry point
- Outcome:
/get-leadssupports new sourcing, existing-leads-only enrichment and resumption. Keeppersonalise-eg-tetofedessel-a-tetofedoas the proven helper. Recoverably archive onlypersonalisation-eg-tetofedessel-a-tetofedoandpersonalisation-eg-tetofedessel-a-tetofedo-codex, retaining precise examples and valid callers. Prepare a simple reusable HTML template as inspiration, not rigid extra machinery. - Verify: fresh-task discovery and realistic invocations resolve the helper and
/scrape, reuse prior records and validate the installed skill. Record original paths, backups/checksums, caller updates and tested restore instructions before retirement. - Must not: break unrelated skills, erase originals, build on unresolved skill pointers or multiply user entry points.
- Evidence to fill: [pending: installation/restore receipt, fresh invocation and dependency checks].
D2 · A review page that controls the approved method
- Outcome: automatically open a concise HTML review with situation, existing and already-scraped counts, desired output, missing inputs, decisions, editable fields/toggles/annotations, scraper/source/model, all proposed Apify filters, exact extraction/generation prompts and clearly labelled examples. The future results table starts empty.
- Gate: saved changes remain pending. On explicit review acknowledgement, the agent reads the centrally saved decision revision and records actor, time, revision/hash and method it will use. Revisions stay pending until acknowledged. Rejection/stopping preserves the run without paid work. Material method changes need a new acknowledgement.
- Normal workflow: human method review → sample of 100, or the smaller requested total → stop for human feedback → revise results as requested → proceed with the remainder only when authorized. Sol's reviewer exception applies only to this first-100 test run, after Matt's plan review. Sol cannot authorize remainder work or future runs.
- Verify: browser edits, save, reload and an agent readback use the same revision. Test that annotation, timeout and stale approval cannot start work.
- Must not: mistake examples/unknown counts for real leads, treat saved comments as approval or spend before the relevant gate.
- Evidence to fill: [pending screenshot: method page], [pending: saved-decision/agent-consumption and gate tests].
D3 · Evidence-backed Python pipeline
- Outcome: reuse sufficient readable company evidence, scrape missing/inadequate pages, deduplicate and assess fit before personalisation. Keep source identity, evidence URLs/quotes, original values and every candidate's outcome.
- Model: OpenRouter
openai/gpt-5.6-luna, Medium, is the default. Preserve explicitgoogle/gemini-2.5-flashselection and validate its availability/model-specific parameters before use. The latter currently has an OpenRouter retirement date of 2026-10-20, so future invocation must recheck it. No automatic model replacement. Same-model provider failover is the practical default, recorded visibly, with required parameters enforced. - Verify: scripts reconcile every input to output/exclusion/failure, stable identities, evidence, exports and cost. Each generation records requested/returned model, reasoning settings, actual provider, request ID, usage/cost and available retry receipts. Reject mismatched models or invalid required fields. Test malformed output, weak evidence, duplicates, rate limits, budget exhaustion and interruption/resume.
- Must not: invent activity/name/fit, accept website copy as instructions, hide substitutions or charge completed work again on resume.
- Evidence to fill: [pending: all-row validation, provider receipts and recovery tests].
D4 · Live, editable results across devices
- Outcome: hosted rows and statistics update while processing, with a practical default refresh of about five seconds. Notes begin empty. Saved edits survive reload, a second authorized browser/device, concurrent progress and restart. Failures show reason, retry state and terminal outcome.
- State: a durable central run record is authoritative. Stable run/row IDs and revisions prevent stale writes from overwriting human edits. The UI distinguishes saving/saved/failed and the agent records which committed revision it consumed. Localhost is a progress dashboard, not a second lead store.
- Verify: real before/during/after screenshots plus persistence readbacks. Edit during progress on one client, inspect it on another, force a stale-write conflict, resume and export the same saved values. Check protected access and all existing shared comment/image/asset routes after deployment and after a later publish.
- Must not: claim static HTML, browser localStorage or worker memory provides central persistence, erase prior output, silently lose edits or expose API credentials to the browser.
- Evidence to fill: [pending screenshots: before/during/after and second client], [pending: conflict, persistence, access and publish-compatibility tests].
D5 · Permanent, agent-readable history
- Outcome: every invocation immediately creates a run in one hosted HTML index, including pending, failed and abandoned runs. Structured records retain inputs, source/order/filters, scraper/model/prompt versions, costs/counts, outputs/evidence, decision revisions, feedback and lessons. One focused landing page links skill guidance, trial review, history and dashboard.
- Verify: a second invocation finds and uses the first record, avoids duplicated processing and preserves earlier revisions. A resume reuses its run ID and durable stage state. A new run records its relationship to reused inputs/results.
- Must not: confuse generated/reserved/contacted status, present stale verification as current or disclose private leads/raw transcripts unrestrictedly.
- Evidence to fill: [pending: two-run history/reuse and revision/export checks], [pending screenshot: index].
D6 · Independent first-100 trial and friction closure
- Outcome: Sol independently invokes the installed skill with the source-task brief. Freeze positions 1–100 in an ordered manifest before processing, with source IDs, source/query/order/time and manifest hash. Every position ends as personalised, nonfit/excluded or failed. Retry the same identity, never substitute position 101 to claim success.
- Review table: website URL, scraped-content link,
personalisation1, greeting and initially empty notes, plus fit/reason, original/casual company name,cég/vállalkozás, evidence and rendered opening/full-copy preview needed for this trial. Preserve original/generated/human-edited values.personalisation2is optional and unreviewed unless separately reviewed. - Verify: reconcile all 100 identities and statuses across table/export/history. Sol checks each proposed row's truth, representative activity, evidence, grammar and natural Hungarian in the rendered sentence. Deliver real screenshots and a friction report, with fixes and affected retest receipts. If fewer than 100 inputs can be recovered, report the exact shortfall and continue authorized recovery rather than fabricate a completed trial.
- Must not: promise 100 accepted leads, silently replace rejects, assert 95% reliability from this trial, expand to the historical 1,200 leads or create/edit/upload/activate campaigns or send messages.
- Evidence to fill: [pending: frozen manifest, 100-row reconciliation, Sol review, screenshots and friction closure].
Operational defaults and boundaries
- Language and identity: use concise evidence-backed Hungarian fragments. Render them with the approved sentence and objective/formal trial copy. Use a verified named greeting where suitable, otherwise
Jó napot kívánok. Show two impersonal alternatives where relevant to the source copy without altering campaign content. Normal runs inherit their reviewed copy/greeting. - Business noun: Python regex uses
cégonly whenKft.orKorlátolt felelősségű társaságbelongs to the target business. Footer vendors and other companies do not qualify. Otherwise usevállalkozás, without inferring sole-trader status. Test these cases. - Casual name: company branding/logo is the strongest clue, domain second, and legal text before Kft. another candidate. Preserve original/evidence and flag ambiguity. Do not blindly strip legal names into unnatural names. Distinguish manufacturers, retailers and installers using evidence and retain proven human corrections.
- Spend and recovery: one runner-owned durable ledger covers all agents, builder checks, tests, scraping, retries and trial under USD 15 total unless Matt raises it. Reserve bounded expected maximum cost before calls, reconcile provider cost/unknown outcomes, retain retry counters and stop paid dispatch when remaining budget is unknown or insufficient. Use provider ceilings where available. Cache successes and reconcile uncertain requests before retrying. Evidence changes invalidate only affected stages. Exhausted retries/budget preserve an honest failed/paused state and exact resume point.
- Reuse and privacy: use the smallest maintainable extension of existing scraper/helper/review infrastructure. Reuse existing authorized-user access across devices, protecting page, API, evidence and exports. Keep private transcripts local. Verify live configuration and selected provider's applicable privacy/parameter support before paid processing. Public documentation is not proof of deployed access or persistence.
- Outreach eligibility: current email verification, prior-used-company/reservation and suppression checks are required before any row is labelled outreach-ready. This trial prepares reviewable data and makes no campaign writes or sends.
- Ownership: preserve unrelated work and active writers. Astra coordinates fixes in a safe revision rather than editing code another runner is using. Fix and retest within scope and budget. An optional check is not a new gate, but material missing proof remains incomplete.
Required source brief
- Planning root:
/Users/agency/Documents/Agty/ColdEmail/.tmp/get-leads-plan-20260921. Readuser_prompts_and_goal_relevant_info.md,evidence/report.mdand both session gists. The clean transcripts support quotes, not historical provider/tool claims. - Parent task:
01a07ec5-9b94-7c23-95f5-8a89ee994b47. Trial source task:01a0c37e-068b-7333-be1e-9d673b91c121. - Read the edited source plan:
/Users/agency/.codex/plans/01a0c37e-068b-7333-be1e-9d673b91c121/01a0c380-e62e-7433-8a59-8f459a0b8dc6/PLAN.md. Before Sol launch, a fresh Luna Max usesextract-gist-of-chat-jsonson the trial source and this edited plan to produce its self-contained brief, reusing the current verified extraction when unchanged. Refresh changing source/inventory facts only as needed. Historical campaign operations remain outside this contract. - Canonical skills:
/Users/agency/.agents/skills. Verify current files/callers before editing. Oldlead-bank-campaignandlead-bank-enrichmentreferences are unresolved and cannot be assumed usable.
Review now: the six outcomes, first-100 boundary, normal human gates and the trial-only Sol exception. All evidence slots are pending execution.