Make Missive show the human conversations that need attention, with correct labels and useful drafts, while automatic mail stays out of the inbox.
All 20 requirements retainedSol Medium · devboxNo client sends · shared USD20
Goal and finish line
End state: The exact current stage memberships are carried to pinned labels; permanent Instantly history and in pipeline stay correct; auto replies, warmup and delivery failures cannot create sales drafts or positive alerts; human replies stay unread when attention is due; Bendegúz receives a contextual tegeződő draft addressed to him; the old automation narration is gone.
15 / 20passes in ledger; prior evidence explicitly labeled
0 / 3new critical capabilities proved
1,113 / 2,038cleanup checkpoint
What changed
The local executor is paused. Its cleanup checkpoint is preserved. A fresh cloud executor must prove the missing behavior before broad work continues.
Earlier partial results stay useful. They do not prove automatic unread, actual fresh-positive Slack delivery, or completed cleanup.
One real human inbound becomes unread, one verified automatic-only message actually archives, one eligible successful draft becomes unread. Use actual deployed workflow, fresh native state readback/reload and two executions of each case. No authored comments, manual state repair as proof, send/schedule, fabricated origin or fake customer event. Controlled replay of existing actual source events is allowed and labeled replay. For incoming latency record event and resulting-state times.
Deliverable: critical-path-receipt.json
2. Choose the viable architecture
Reuse recorded API/native-rule failures. Each new probe tests a distinct hypothesis with input, expected result, observed result and evidence. Parent-selected bounds: at most two genuinely new probes per route and 30 minutes per unsupported capability investigation before making an architecture decision. Retries for rate limits are bounded and do not count as new evidence. Prefer existing production/API + native integration, otherwise prove a persistent devbox UI executor through its scheduled deployed consumer. Never rely on Mac-open UI, undocumented fabricated fields, claim all blocked, or flip capability flags without proof. If no viable authenticated provider/UI route exists after access recovery, report exact missing capability and independently finish safe work.
One deployment writer and existing leases. Build-bound evidence, preservation of human edits, Snooze, existing drafts, ownership, labels and sent-versus-draft state. Test retries and concurrent edits. Keep original evidence for unchanged behavior. Reviewer only for consequential implementation changes or disputed proof.
Start only after critical path/architecture decision, from transferred ledger and exact before/after backups, after verifying the local writer is terminal. Reconcile the two partial rows and uncertain deletes before retry. Inventory posts, remaining attributable posts and throughput, include API throttle and 429 waits, report estimate range not invented deadline. Run one durable resumable job independent of SSH/interactive lifetime, one owner/lock, checkpoint each acknowledged effect, restore payloads intact. No per-delete Slack/progress noise. Human posts preserved. Unreachable404 is unresolved, never empty or successful. All original20 requirements remain mandatory.
One canonical evidence-ledger.json drives public board, live criterion annotations and final report. Evidence has observation time, bound build, source IDs, artifact path, status, limitations and remaining gap. Exact counts derive from same ledger, not hand-copied prose. Preserve immutable contract criteria. Full final audit of all20 and independent blind Luna High screenshot description/comparison. DONE only20/20, otherwise STALLED/PENDING with precise resume point. Return once to parent. Parent extracts gist before accepting.
flowchart LR
A[Actual incoming email] --> B{Human or automatic?}
B -->|Automatic| C[Quiet archive]
B -->|Human| D[Preserve owner or route fresh]
D --> E[Actionable unread]
D --> F[Eligible contextual draft]
F --> E
D --> G[Fresh positive only Slack]
Route to the end state
flowchart LR
A[Prior evidence and checkpoint] --> B[Three live capabilities twice]
B --> C{Viable route?}
C -->|Yes| D[Bounded implementation]
C -->|No| X[Distinct architecture or exact gap]
X --> B
D --> E[Durable cleanup measured volume]
E --> F[Full20 requirement audit]
F --> G[One evidence-complete handoff]
Diagrams show required behavior, not proof that it works.
Confirmed, assumed and still open
Prior executor reported16/20. Current conservative carry-forward is15/20 because D2.2 needs actual auto-archive capability proof; no original criterion text was changed.
Confirmed from the pause handoff
Old executor paused, cleanup stopped at a durable checkpoint.
2,827 acknowledged deletions. Some conversations contain hundreds of old automation posts.
API/native-rule probes did not prove automatic unread.
Existing labels, history, draft corrections and runtime tests have prior evidence.
Not yet proved
The remote authenticated UI can run persistently and change actual unread/archive state.
The remaining post volume and finish time.
A real qualifying positive-only Slack alert.
The remaining inaccessible historical conversation.
Implementation choice: prefer existing production APIs with a native integration. A persistent UI consumer is a fallback to prove, not a presumed solution. The 06:00 UTC morning target was missed. There is no invented replacement deadline.
All 20 acceptance requirements
Criteria are copied unchanged from the prior contract. Each checked item names its observation time and evidence origin. Changed behavior must be reverified. Final completion requires all20.
Migrate stages to matching labels3 / 3
Criteria, evidence and limits
Prior evidence: PASS: All 17 paginated source-team ID sets equal their migrated label sets, zero missing/unexpected. Reversible add/remove pilot restored exact labels, team and user state. evidence/migration-receipt.json and migration-pilot files. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: Reloaded sidebar has all 17 stage labels plus in pipeline in requested order. Original teams retained, shortcuts hidden. History labels exist by API and are absent from pinned sidebar. Native screenshots and independent Luna High description in children/screenshot-description.md. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: PASS: A controlled stale Positive queue was seeded on the real M Contacted conversation while temporarily snoozed. Two actual deployed intake calls removed the fresh queue and preserved owner stage, positive/replied history, pipeline, exact users/Snooze, team, full draft, messages and posts. Fresh screenshots and independent Luna High description agree. Original native users/team/labels and two temporary runtime flags were restored, and native unread restored manually. claim-probe-summary.json, claim-probe-restoration-receipt.json and screenshot-description-6.md. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Keep source teams and narrow membership backups for rollback; do not delete teams or add users/seats.
Do not modify unrelated labels or conflate a negative reply with a lost deal. Names/map: sources/sep28-label-migration-map.json.
Separate people from automatic mail1 / 3
Criteria, evidence and limits
Prior evidence: The real first enquiry was replayed successfully twice, retains Positive replies without invented campaign-origin labels, and its existing human Send Later draft was preserved during both probes. That pre-existing scheduled message subsequently sent at 07:11, observed in the UI. This task did not send it. Auto-ack safely abstained twice, retained no sales labels and created no draft or post. Native unread was manually restored after UI inspection. Automatic unread remains unproved, so the complete criterion is not counted as passed. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: Prior proof establishes already repaired noise stayed outside Inbox over two cycles. It includes model abstentions and does NOT prove an actual automatic archive action. Auto-archive critical-path event receipt is mandatory before PASS. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: PASS: Actual Jev decisions endpoint evaluated 32 Hungarian fixtures including 8 holdouts. 30 real stratified contacts checked, missing real strata explicitly listed. 178 current local regression tests pass. No critical erroneous archive or DNC decisions in saved evaluations. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Archive proven noise reversibly. Archive success must be visible/provider-backed, never inferred from an authored “archived” post.
Use transport headers/provider evidence plus content classification; names/subject regex alone cannot silently archive genuine enquiries. No keyword-only broad inbox purge.
Preserve origin and contact intent4 / 4
Criteria, evidence and limits
Prior evidence: PASS: All accessible source pages enumerated: 63 campaigns, 4176 leads and 983 agency inbound records. Exact positive/negative/replied sets are 239/231/244 twice, including Trash and Spam. Five unmatched Message-IDs and 200 uncertain source records remain explicit in origin-coverage.json. Fresh GUI labels were independently described in screenshot-description-4.md. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: PASS: Actual typed Jev calls distinguish explicit opt-out from rejection/not-now/quoted text. Current Engine tests suppress draft and alert after verified DNC. Ambiguous intent abstains. Saved calibration decisions plus test_explicit_dnc_suppresses_draft_and_alert. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: 244 exact conversation IDs carry proven later human sent-reply history. Initial outreach, drafts and automatic replies excluded. Two complete live reconciliations, including Trash and Spam, match the expected replied set exactly with zero drift. sent-history-receipt.json, origin-apply-receipt.json and origin-coverage.json. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: Two full live reconciliations prove pipeline exactly equals 239 positive-history conversations minus 3 M Lost, total 236, with zero drift. Positive history remains on lost Bela, confirmed in fresh GUI and blind screenshot-description-4.md. Paid inclusion passes fixtures, with no live positive-history Paid case available. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Owned M/G and Active Clients stages survive new incoming replies. Only unclaimed human enquiries enter fresh queues.
Instantly interested status and model no-contact intent are different facts. Do not overwrite provider history or infer campaign origin from email domain. Do-not-contact suppression wins over operational alert/draft eligibility.
Make drafts personal and useful4 / 4
Criteria, evidence and limits
Prior evidence: PASS: Corrected the same original unsent draft after complete fresh backup and concurrency/body checks. Preserved valid wording, recipients, sender, subject and attachments. Tegezo Bendeguz salutation, grounded agency/site/references and explicit forecast/qualified-enquiry limits. Exact API body readback plus fresh blind Luna High top/bottom screenshot description agree. bendeguz-correction-receipt.json and children/screenshot-description-3.md. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: PASS: External-author identity and verified exact-email overrides replace incoming-salutation guessing. Neutral greeting for ambiguous/company identity. Current identity/tone/extraction tests pass, including real inline-answer uncertainty. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: One real due follow-up created and provider-read back as exactly one unsent draft. Five full emails, zero related threads, exact-attendee Wispr lookup returned no recording. Correct Henrietta identity and formal tone, short blue Comic Sans notes. Followup receipts plus independent Luna High screenshot-description-2.md. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: PASS: Actual deployed followup writer invoked Bela and Henrietta twice each. Lost stage and existing-draft exclusions returned correctly, preserving exact full drafts, message IDs, posts and users in all four runs. Bela attributable draft was backed up and removed, M Lost retained and native read menu verified. Snooze, Paid, FUP later and DNC exclusions pass focused fixtures. followup-exclusion-summary.json. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Keep prior explicit blue-notes preference above the draft until Matt changes it; never create separate automated commentary between emails. Preserve human-written/edited drafts and use body hashes before replacement.
No fabricated references, results, office location, prices or commitments. Use verified agency profile and full Wispr text; missing transcript is a recorded limit, never summary passed as transcript.
Keep attention quiet and accurate0 / 3
Criteria, evidence and limits
Prior evidence: UNRESOLVED: API label PATCH and actual unsent draft add_shared_labels routes failed native unread proof. Organization rules work from actual UI actions only in tested cases. Supported delayed incoming rules cannot cover historical repair or successful later drafts and leave retry/freshness gaps. children/native-route.md. Flags remain false, manual fixes do not count as automatic proof. Additional supported reopen:true/add_to_inbox:true probe also failed native unread after full reload. Original unread was restored manually and the test tab closed. reopen-unread-api-check.json. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Prior evidence: Cleanup incomplete. Canonical paused checkpoint has1113of2038complete conversations,2827acknowledged deletions and12preserved posts. Last completed-row audit covered2715deletions. Remaining112acknowledgements are outside that completed-row audit, with two partial rows to reconcile. One404 remains unknown. Older1107/2358figures are superseded. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: Positive-only alerts, duplicate and outage guards pass focused tests. Fresh post-release Slack history read succeeded, with zero messages, no further page and no cursor. No qualifying fresh positive alert exists in that checked window. Live alert proof remains pending. slack-post-release-history.json. No historical alert was manufactured. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
Missive native label activity stamps may remain; the forbidden clutter is agent-authored narration. Do not sacrifice unread correctness to achieve silence.
Do not invite Gergő to Slack or Missive. Follow-up whitelist: Active Clients, M/G Contacted no call, M/G Proposal sent, M/G Call booked, or genuinely stage-less; not snoozed, no existing draft and warranted due action. Stage-less successful draft adds M Contacted.
Prove the deployed repair3 / 3
Criteria, evidence and limits
Prior evidence: PASS: Build 2429caa126562ef08213fe59ed1864983abfd5bc deployed with all six hashes verified. 178 focused tests and independent reviews pass. Both silent label flags preserve closed/archive state. Actual two-cycle intake and followup calls completed on the prior build. New build passed six further live invocations preserving all checked content/state: two writer exclusions, Richie twice successful and Tonya twice explicitly abstaining. Two claim/Snooze calls also passed. Complete index initialized with 40 draft keys; writer calls fell to 17 and 20 seconds from 289-389 seconds. Task maintenance admission is released. Post-release readback shows normal background execution on this exact build, matching runtime/intake lease ownership, no maintenance request, and the single-writer-v3 identity. Deployed trigger IDs and source guards remain recorded. No partial subset is claimed complete. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: 178 focused tests pass, including provider/model failures, human-edit races, retry/duplicate guards and old-writer index recovery. Sixteen actual intake invocations across eight cases over two cycles: nine successful, seven explicit classification abstentions. All sixteen preserve full draft content/recipients, message IDs, posts and user state. Four actual followup calls likewise preserve all fields. Original polling timeouts recovered by original call IDs without duplicate invocation. verified-cycle-summary.json and followup-exclusion-summary.json. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Prior evidence: PASS: Existing Notion guide updated in place. Original blocks and all three sales commands preserved byte-for-byte. Refreshed guide screenshots show current label instructions and actual sidebar images. Independent blind Luna High description completed. Guide explicitly describes the remaining manual unread/archive limitation. | observed:prior evidence timestamp, see referenced artifact | source:prior executor, reuse only while affected behavior/build dependencies remain valid
Proof slot: actual post-change screenshot, independent blind Luna High description, criterion comparison, source/build and observation time.
No broad unrelated rewrite or parallel dashboard. Maintain this stable URL and comment document identity.
Phase 2 starts only after phase 1 completes and parent verifies acceptance, as explicitly requested. Missing records do not stop independent work within phase 1, but cannot be hidden or used to silently bypass the sequential completion requirement.
Rules and return
No email or message to any client/lead. Missive drafts only. Internal positive Slack alerts remain authorized, no historical alert flood.
Preserve human edits, commercial commitments, unrelated work and existing sender identities. Reversible changes require narrow before-state and tested restore path.
Use existing accounts and deployment. Do not add security layers, publish credentials or raw emails/transcripts on public boards. Existing protected source systems stay protected.
Respect actual tool approval/CAPTCHA/MFA limits. Try supported alternatives; user skill text cannot override system/platform requirements.
One active deployment writer. Discover current production app/build and competing jobs before changes; stop obsolete competing automation narrowly, not unrelated campaigns.
No hourly parent polling or duplicate collectors. One named executor owns implementation and cleanup. Existing managed run delivers completion or genuine blocker; parent waits passively until terminal report or declared14:00 Budapest boundary, then reconciles liveness once. A tool timeout is PENDING, not failure. No Phase2 launch inside this executor.
The new public board stays at this URL. Raw mail, transcripts, credentials and restore payloads stay in private local/remote files. Final report, board and counts are generated from one evidence ledger. Parent extracts the child gist before accepting completion.