name: pup-bridge-test description: "The standard pup test scenarios (source-only, maintainer skill). Run these after ANY pup change instead of asking John to spell out a test. Covers: the up-gate, smoke, the 5-thread simulation (the John test), the heavy-load matrix, the 3-phase open/reload/fg-sweep, the parallel-open concurrency proof, one-flash, strand-rescue, login persistence, and cross-thread pollution. Every scenario states its EXPECTED output — a run without expectations is not a test."

pup-bridge-test — the standard test scenarios

STANDING DIRECTIVE (John, 2026-07-27): the AI-THREAD NAME is the identifier — sessionId is legacy. Tables and reports key on the thread name. Opens already default to the thread window (1.9.147); NEXT IMPLEMENTATION ITEM: every remaining verb (eval/screenshot/navigate/close/raise/ login/...) must accept a missing sessionId and resolve to the caller thread's own window via callerOwnWindow(); then retire sessionId from examples, hints, and describe. Test any verb sweep by calling it with NO sessionId from a thread that owns exactly one window. ALSO QUEUED — RECOVERY-BINDING, root cause now FULLY narrowed (2026-07-27, two failed proofs): all four title/window.name injectors are stamped (1.9.151+1.9.153) and the RESCAN matcher reads window.name — but the FROM-DISK recovery loop ('Recovering N session(s) from disk', ~line 1160) binds each session to a page with its OWN logic that never consults window.name; it bound docs-writer to pages[0] (ESPN) and the real rust-lang page, left unclaimed, was closed as a leftover. THE ONE REMAINING EDIT: in that loop, before binding, read window.name from candidate pages and bind ONLY the matching page (fall back to title tag, then lastKnownUrl); never close an unclaimed page that carries a pup: name. Prove with: open 3 (one SPA), restart, binding table 3/3. Prior scenario text: the recovery-binding scenario (open 3 incl. an SPA, restart, binding table 3/3) — the launch-path title injector still needs the window.name stamp (the third injector; two of three were patched in 1.9.151 and docs-writer still mis-bound).

Rules that apply to EVERY scenario (each one was paid for):

  • Up-gate first. After any install/restart, wait until a REAL verb answers — poll pup_list_windows until count is non-null. bridge_list showing the version only proves the file landed; the bridge may still be in its npm-reconstruct window and every test will false-fail (it torched two full runs on 2026-07-26).

  • Never >/dev/null an open. Capture ok + error + errorCode per call. Blind opens cost three re-runs in one afternoon.

  • Prove on glass. A rect read (desktop_find_windowrect) beats any log line. Logs cap at ~200 lines per read and scroll under load — absence of a log line is NOT absence of the event.

  • Clean up. Close every window the test opened (except ones a scenario deliberately leaves for John to inspect). Leaving video sites open cooks the laptop (348% CPU incident).

  • Use throwaway thread names (sim-*, thread-*) so ownership machinery is exercised and no real thread's window is touched.

  • ALWAYS end a window-opening scenario with the GEOMETRY TABLE (John's standing rule, 2026-07-27): one row per opened window — sessionId, position, width x height — shown to John in the report. Every background window must read the standard inset frame (position 10,10 CSS = work-area origin + 10px; size = work area - 20 in each dimension; e.g. 15,15 2531x1499 on the 2560-wide laptop at its DPI). One deviant row = a geometry regression; find the sibling path that wrote it (park, reveal, heal, or a legacy maximize) before anything else.

  • ALWAYS show the FOREGROUND-GATE TABLE for any reveal phase (John's standing rule, 2026-07-27): one row per raised window — sessionId, the BARE raise result (MUST be DENIED), the raise WITH foregroundReason (MUST be GRANTED), and the exact reason string used. The reason must quote or closely paraphrase what the USER said (it is logged and auditable) — the table proves to John that pup made the caller justify itself before touching his screen. Follow it with the post-raise geometry table: revealing must NOT change the standard frame.

  • EVERY test table includes the OVERLAY column (John's standing rule, 2026-07-27): per window, report the overlay field from pup_list_windowsfavicon:<host> (the active tab's icon, the standard), count:N (no favicon; tab-count badge fallback), or none with its reason. After ANY reveal phase, re-read it: raising redraws the taskbar button and v1.9.150 re-wears the overlay at +1.2s — a none after a raise on a favicon-serving page is the overlay-stripped regression John caught live.

  • The post-open table's full column set (John, 2026-07-27): sessionId | position | WxH | overlay | flashCount | progressBarCount | progressBarMs — all read straight from pup_list_windows (v1.9.152+). Expected for a plain background open: flashCount exactly 1, progressBar count/ms reflecting only real agent driving (a giant progressBarMs on an idle window = a stuck busy-state). After a reveal phase, re-read: overlay must survive, flashCount must NOT have grown unless a new operation ran.

Target boilerplate: T="pup bridge maintenance" and every call is adom-bridge-cli --target AdomLapper --ai-thread "<thread>" <verb> '<json>'.

1. Smoke (after every ship)

Open (no sessionId, thread-identified) → eval 1+1 → screenshot → close. Expect: open ok:true with sessionId: ai-<thread-slug>; eval 2; close chromeProcsRemaining: 0.

2. The 5-thread simulation ("the John test", 2026-07-27)

Five sim-* threads open windows with NO sessionIds and differing tab counts (extra opens from the same thread must return openedAsTab:true). Then: bare pup_raise_os_window on three of them (MUST be denied), reasoned raise (MUST be granted), eval sweep on all five (no crashes), wait 10s, close two as their own threads, verify the survivors' exact tab counts and that the closed pair is gone. Expect: every thread its own ai-sim-* window; tab counts exactly as opened; gate DENIED-then-GRANTED ×3; 5/5 evals ok; closes report chromeProcsRemaining: 0; survivors intact.

3. Heavy-load matrix

12 background windows across the heaviest sites (youtube, maps, twitch, cnn, espn, amazon, reddit, github, earth+wiki isolated, a 4-tab news window, docs). Then the verification matrix per window: rect / eval / flash-count / park-count, plus total chrome procs. Expect: every rect at the inset frame (10 CSS px ⇒ 15,15 physical at 150% DPI); every eval ok; exactly 1 [flash] per session (pull the FULL log; see log-cap rule); park + re-assert logged per window.

4. The 3-phase test (open → reload-all → fg-sweep)

10 heavy PARALLEL opens (this is what found the sibling-kill + Chrome-exit chain), wait 10s, reload ALL in parallel, wait 20s, pup_raise_os_window sweep with a quoted foregroundReason, 1s apart. Expect: 30/30 OKs. Any 'Navigating frame was detached' or 'Session not found' = the serialization regressed (_profileOpenChain).

5. Parallel-open concurrency proof

Two simultaneous no-id opens from DIFFERENT threads + ten simultaneous opens from one thread. Expect: zero failures; same-profile opens serialize (arrival order in the log), different profiles parallel.

6. One-flash proof

Open a slow site (youtube). Full-depth log grep for [flash] "<sid>". Expect: exactly ONE (launch) line, no (auto) after page load. Two = the _openFlashDone consume broke.

7. Strand-rescue (placement sweep)

Open a window, SetWindowPos it to -30000,-30000 via shell_execute, wait ~135s. Expect: rect back at the inset frame ("RESCUED"), and the repair reported via _cleanupReport on a later verb. (Report delivery was unverified on 2026-07-26 — check it.)

8. Login persistence

With a stored test cred: open the login page → auto-filled on load (_autoLogin present) → close window → kill the profile's Chrome → reopen → still signed in. Expect: cookie survives close, process kill, and bridge restart. (Google login proven 2026-07-26; use the-internet.herokuapp.com for repeatable runs.)

9. Cross-thread pollution

ALPHA opens a window; BETA calls pup_open_tab with NO sessionId, then WITH ALPHA's id. Expect: first refused no_target_window; second refused session_owned_by_another_thread; ALPHA's tab count unchanged both times.

10. Jam-resilience under heavy load (the "pup looks frozen" test, 2026-07-28)

Root cause established live: John's "pup seems frozen" is USUALLY a CDP jam, not a crash. Under heavy load (many heavy SPAs on the one shared Chrome + memory pressure) Chrome stays ALIVE (check Get-Process chrome count — it's ~45, not 0) but every CDP verb times out. The pre-1.9.172 pup_list_windows/pup_status did uncapped CDP reads (page.title() + a focus evaluate) SERIALLY per window, so ONE busy window hung the whole handler and the CLI returned count: null — the frozen symptom. Fixed 1.9.172: both verbs gather sessions in PARALLEL with a 2.5s per-session timeout and an in-memory getSessionInfoFast fallback; busy windows come back marked degraded. Test: open 5-6 heavy windows (maps, earth.google, youtube, figma, diagrams), force memory pressure (pup_eval pushing new Uint8Array(1MB) x50 per window), and hammer pup_list_windows x15 WHILE the allocations run. Expect: 15/15 return the full list (never null/timeout), each < ~2.5s; _degraded appears only if a window is genuinely mid-jam. Proven 1.9.173: 15/15 live, 37-226ms, 0 null.

11. Crash self-heal + the reclaim race (2026-07-28)

Open 3 windows on the shared profile → kill EVERY adom-you chrome (simulates the real crash) → do nothing 55s. Expect: the health/heal loop (15s tick) relaunches every crashed session at its last URL, signed-in, in the background — final window count == pre-crash count, with 0 "session not found" in the [heal] log. The not-found failure (seen on 1.9.171 under stress) was a RACE: reclaimUnmanagedWindows reaped the crashed session + deleted its file before the heal relaunched it. Fixed 1.9.172-173: disconnect stamps _lastUrl; reclaim skips _lostBrowser sessions for a 2-min heal window; recovery falls back to the cached URL. (The [heal] relaunched OK lines scroll off a 50-line log read under load — the count-based check is the reliable proof.)

12. TORTURE (the "never crash" push, 2026-07-28)

The full brutal load, with an INDEPENDENT monitor thread proving pup never goes frozen mid-storm. From thread pup-monitor, sample pup_list_windows every 3s for the whole run. From thread pup-torture, open 6 heavy windows and run several rounds of {50MB alloc x6 + navigate churn to earth/youtube}. Then settle 40s and check recovery. Expect: the monitor's null/timeout count is 0 across every sample (pup is NEVER frozen to a second, independent thread even while another thread hammers it); final count recovers; any Chrome crash shows [heal] recovery, not stranding. This is the standing "push it to the limits" run — re-run it after any change to the CDP/verb/heal paths.

When a scenario fails

Get the RAW error before theorizing (the parse-wrapper hides fields). Then check, in order: the up-gate (was the bridge actually ready?), the log for [park]/[flash]/[janitor]/[heal] lines, a single sequential repro of the failing call, and only then the code. Two of the three "total failures" on 2026-07-26 were the harness (npm window, outer timeout), not the bridge. NOTE on distinguishing jam vs crash: shell_execute still answers when only the CDP path is jammed (the Node event loop is fine), so Get-Process chrome | Measure tells you alive-but-jammed (~45 procs) vs actually-crashed (0) in one call — do this BEFORE assuming a crash.

⛔ Testing rule: a TEST may never fire native OS input without warning the user

Every scenario here runs on John's real machine while he is working. desktop_click / desktop_type / desktop_press_key go to whatever window has OS focus, so a test that fires them blind types into whatever he is doing. This happened. Before any native step in a test: toast + caption + an explicit "don't type or click until I say done" + wait until the target window is genuinely foreground + re-verify before each event + caption when done (ensureSafeForNativeInput / pup_assisted_click). If you cannot warn him, use CDP input (pup_click / pup_type / pup_hover) — it needs no foreground and cannot disturb him.

Also: a green API result is not a passing test. ok:true from a native verb means the event went to the OS, not that it reached your window; the badge API reported success on an invisible badge; and pup_type reports success while keystrokes never land in a window whose renderer is not accepting input. Verify the EFFECT (read the field back, screenshot the taskbar, probe the DOM).

AUDIT YOURSELF — the standard procedure when John says "audit your windows" (2026-07-30)

"Audit your windows" ALWAYS includes the ICONS. John had to tell me this twice: I reported sessions, health, geometry and logs and left the taskbar out, and the icon state was where the real regressions were hiding (a window that silently lost its badge after a restart; 8 stale AUMID keys for 3 windows). An audit without the icon section is not an audit — and it is not finished until you have looked at the taskbar ON GLASS, because the API reports a badge as applied while it is invisible.

Run ALL SEVEN sections, in this order, and report them as a table:

  1. Sessions (bridge view)pup_list_windows: count, sessionId, owner, tabCount, profile, ageMinutes, _degraded. Flag: owner:null (ownership protection is off), duplicate sessions pointing at one window, a tabCount that does not match what was opened.
  2. ICONS / TASKBAR (never skip)
    • per window: the overlay field — favicon+adom:<host> (correct), count:N, or none plus its reason. A none on a favicon-serving page is a regression; a none on local PDFs is legitimate (no favicon exists).
    • AUMID keys: reg query HKCU\Software\Classes\AppUserModelId filtered to Adom.Pup*. Expect exactly 1 (Adom.Pup) under minimalTouch. More = stale keys from closed windows, or per-window identity churn that destroys/recreates taskbar buttons.
    • ON GLASS: desktop_screenshot_screen, crop the taskbar strip, and LOOK. Confirm each pup button wears the composite badge (page favicon + Adom mark inset flush in the lower-right). Report how many of N windows are actually badged.
  3. OS ground truthdesktop_find_window {titleContains:"(session:"}: one real window per session (no ghosts, no untracked strays), and each rect at the standard inset frame (15,15 2531x1499 on this laptop). A maximized window reads ~-11,-11 2582x1550 — that is maximized, not stranded. -32000,-32000 = minimized, confirm with IsIconic before calling it lost.
  4. Healthpup_eval per window: readyState, visibilityState, a 1+1 drivability check. visibilityState:hidden is NORMAL for a parked/minimized window; only frontmost-yet-hidden is the broken state. Note any crashed tabs.
  5. Logbridge_log_read, count [crash] [heal] [health] [unstick] [disconnect] [reclaim] [tabs] park-failed Handler error. Explicitly report park-failed and [reclaim] adopted lines: a park-failed window is sitting wherever Chrome left it, and an adoption means pup LOST a window and re-bound it (which produces duplicate sessions for one window).
  6. Chrome footprint — procs + total MB, and pup's share vs the user's own Chrome. Never report a big total as if it were all pup's.
  7. Cleanup state — leftover test windows, phantom sessions, stale session files.

Report rules: state what is WRONG first, name which findings are your own regressions, and never present an API-reported success (badge applied, verb ok:true) as proof — prove it on glass or in the page.