skill
DEPRECATED: Adom Bridge Control (moved into Adom Bridge)
Public Made by Adomby adom
Drive the user's real desktop from the cloud: their Chrome/Edge, headless pup, native UI, the Hydrogen editor, sibling Claude tabs, plus screen recording and video post.
← Commit history
Publish 1.0.0
12 files changed
+637
SKILL.md+72install.sh+15package.json+44screenshots/hero.pngskills/claude-control/SKILL.md+50skills/desktop-ui-control/SKILL.md+72skills/hydrogen-web-control/SKILL.md+65skills/native-browser-control/SKILL.md+89skills/pup-browser-control/SKILL.md+54skills/screen-recording/SKILL.md+69skills/video-post-production/SKILL.md+100uninstall.sh+7SKILL.mdadded+72@@ -0,0 +1,72 @@+---+name: adom-desktop-control+description: >+ Drive an Adom user's REAL machine from the cloud through Adom Desktop: their+ logged-in Chrome/Edge (browser extension / CDP), headless pup browsers, native+ windows (UIA + SendInput), the Hydrogen web editor, and sibling Claude tabs,+ plus a screen-recording + video post-production pipeline. Use when you need to+ click/type/upload in the user's browser, drive a native app window, log a user+ in, record the screen, or produce a demo video. Trigger words: drive the+ browser, control chrome/edge, nbrowser, pup, automate the desktop, click a+ native window, UIA, record the screen, make a demo, speed up a video, type into+ the Claude tab, drive an OAuth login, upload a file via the browser.+---++# Adom Desktop Control++Everything here is about reaching OFF the cloud container and acting on the user's+**actual desktop** through Adom Desktop (the `adom-desktop` CLI + relay). It is a+pack of tested patterns, each one distilled from a real session, each labeling+**what works, what fails, and why** so you don't relearn it the hard way.++## Pick the right surface (the #1 decision)++| You want to drive... | Use | Skill |+|---|---|---|+| The user's REAL logged-in Chrome/Edge (their cookies/sessions) | `nbrowser_*` (browser extension, CDP) | [native-browser-control](skills/native-browser-control/SKILL.md) |+| A throwaway/anonymous browser, or one you'll record clean | `browser_*` (pup / Puppeteer) | [pup-browser-control](skills/pup-browser-control/SKILL.md) |+| A native app window (Fusion, dialogs, non-web UI) | `desktop_*` (UIA + SendInput) | [desktop-ui-control](skills/desktop-ui-control/SKILL.md) |+| The Hydrogen web editor / its Claude panel | (mostly can't, read why first) | [hydrogen-web-control](skills/hydrogen-web-control/SKILL.md) |+| A sibling Claude tab to run a task | depends on where it runs | [claude-control](skills/claude-control/SKILL.md) |+| Capture the screen to a video | `desktop_record_*` + `desktop_caption` | [screen-recording](skills/screen-recording/SKILL.md) |+| Turn raw clips into a finished demo | ffmpeg + `video-post` | [video-post-production](skills/video-post-production/SKILL.md) |++## The cross-cutting rules (true on every surface)++- **Background by default. Never steal the user's focus** unless you must. CDP eval+ (`nbrowser_eval`/`browser_eval`), UIA Invoke (`desktop_ui_click`), and+ `desktop_screenshot_window` all run on a BACKGROUND window. Only OS SendInput+ (`desktop_click`/`desktop_type`/`desktop_press_key`) needs the window foregrounded,+ so avoid it when a CDP/UIA path exists. (John: "browser windows should always be in+ the background so as not to disrupt the user's other work.")+- **Drive it yourself, including login. Never ask the user to click or log in.** You+ can drive their real browser end to end, including OAuth, see native-browser-control.+- **Upload files via DataTransfer, not the OS file picker.** Inject the bytes into the+ page and set the `<input type=file>`, works in pup and nbrowser. The OS picker is a+ last resort (needs foreground SendInput). See native-browser-control / pup-browser-control.+- **Recording captures whatever is topmost.** If you drive via background CDP while+ recording, the recording shows the WRONG window unless you re-foreground the target+ before every visible step (the "foreground re-bump"). See screen-recording.+- **Verify by screenshot, don't assume.** After any click/type into a webview, screenshot+ and confirm the text/state actually landed, OS input silently no-ops on many webview editors.++## Install++```bash+adompkg install adom/adom-desktop-control+```++Drops the pack skill plus all sub-skills into `~/.claude/skills/`. Pairs with the+`adom-desktop` app and the `adom-browser-extension`.++## Sub-skills++- **[native-browser-control](skills/native-browser-control/SKILL.md)** — drive the user's real Chrome/Edge via the extension (CDP).+- **[pup-browser-control](skills/pup-browser-control/SKILL.md)** — drive headless Puppeteer browsers.+- **[desktop-ui-control](skills/desktop-ui-control/SKILL.md)** — click/type native windows via UIA + SendInput.+- **[hydrogen-web-control](skills/hydrogen-web-control/SKILL.md)** — the Hydrogen editor's nested-webview limits.+- **[claude-control](skills/claude-control/SKILL.md)** — drive sibling Claude tabs (and when you can't).+- **[screen-recording](skills/screen-recording/SKILL.md)** — capture the desktop to video with on-screen captions.+- **[video-post-production](skills/video-post-production/SKILL.md)** — cut, speed up, caption, narrate, publish.++Narration for demos uses [`adom-tts`](https://wiki.adom.inc/adom/adom-tts) (house voice + pronunciation cache).
install.shadded+15@@ -0,0 +1,15 @@+#!/usr/bin/env bash+# Install Adom Desktop Control: the pack skill + every bundled sub-skill.+# Each skill named explicitly (the publish linter requires it, no globs).+set -e+DEST="$HOME/.claude/skills"+install_skill() { mkdir -p "$DEST/$1"; cp "$2" "$DEST/$1/SKILL.md"; }+install_skill adom-desktop-control SKILL.md+install_skill native-browser-control skills/native-browser-control/SKILL.md+install_skill pup-browser-control skills/pup-browser-control/SKILL.md+install_skill desktop-ui-control skills/desktop-ui-control/SKILL.md+install_skill hydrogen-web-control skills/hydrogen-web-control/SKILL.md+install_skill claude-control skills/claude-control/SKILL.md+install_skill screen-recording skills/screen-recording/SKILL.md+install_skill video-post-production skills/video-post-production/SKILL.md+echo "Installed adom-desktop-control + 7 sub-skills."
package.jsonadded+44@@ -0,0 +1,44 @@+{+ "schema_version": 1,+ "slug": "adom-desktop-control",+ "type": "skill",+ "title": "Adom Desktop Control",+ "brief": "Drive the user's real desktop from the cloud: their Chrome/Edge, headless pup, native UI, the Hydrogen editor, sibling Claude tabs, plus screen recording and video post.",+ "version": "1.0.0",+ "description": "A skill pack of hard-won, tested patterns for driving an Adom user's actual machine from the cloud AI through Adom Desktop: their real logged-in Chrome/Edge via the browser extension (CDP), headless pup browsers, native UI via UIA and SendInput, the Hydrogen web editor, sibling Claude tabs, and a full screen-recording + video post-production pipeline. Every skill is distilled from real sessions and labels what works, what fails, and why.",+ "org": "adom",+ "dependencies": {},+ "tags": ["adom-desktop", "automation", "browser", "nbrowser", "pup", "uia", "recording", "video", "hydrogen", "claude"],+ "discovery_triggers": [+ "drive the user's browser",+ "control chrome or edge",+ "native browser extension",+ "nbrowser",+ "drive a pup browser",+ "automate the desktop",+ "click a native window",+ "uia desktop control",+ "record the screen",+ "make a demo video",+ "screen recording",+ "desktop captions",+ "speed up a video",+ "drive the hydrogen editor",+ "type into the claude tab",+ "upload a file to a website via the browser",+ "drive an oauth login"+ ],+ "scripts": { "install": "./install.sh", "uninstall": "./uninstall.sh" },+ "files": [+ "SKILL.md", "README.md", "page.json", "install.sh", "uninstall.sh",+ "screenshots/hero.png",+ "skills/native-browser-control/SKILL.md",+ "skills/pup-browser-control/SKILL.md",+ "skills/desktop-ui-control/SKILL.md",+ "skills/hydrogen-web-control/SKILL.md",+ "skills/claude-control/SKILL.md",+ "skills/screen-recording/SKILL.md",+ "skills/video-post-production/SKILL.md"+ ],+ "hero": { "type": "image", "path": "screenshots/hero.png" }+}
screenshots/hero.pngadded⋯ 1 unchanged line ⋯
skills/claude-control/SKILL.mdadded+50@@ -0,0 +1,50 @@+---+name: claude-control+description: >+ When to delegate work to a sibling Claude tab on the user's machine vs do it+ yourself, and the access differences that decide it (Hydrogen/galliaApril bypass ++ relay access vs claude.ai cloud sandbox). Trigger words: drive another claude,+ sibling claude, claude code tab, delegate to claude, run a prompt in the claude tab.+---++# Claude control (driving a sibling Claude)++Sometimes the user wants a recording or task that shows "the AI" being asked a prompt+and then doing the work. Before you wire that up, know which sibling can actually DO+the work, and that you usually can't type into it anyway.++## Which sibling can drive the user's machine?++| Sibling | Relay/bridge access? | Can drive local Fusion/browser? |+|---|---|---|+| **Hydrogen Claude** in the user's cloud env (e.g. galliaApril), bypass-permissions mode | YES (same container/relay you use) | YES, it can run the full flow autonomously |+| **claude.ai/code** web tab | NO, runs in an isolated cloud sandbox | NO, can't reach the user's local relay |++So only a Hydrogen/galliaApril-env Claude can actually open Fusion, export, drive the+browser. A claude.ai/code tab will "chug" but can't touch the user's apps.++## The catch: you usually can't type the prompt into it++The Hydrogen Claude panel input is unreachable by automation (cross-origin nested+webview, OS input no-ops), see [hydrogen-web-control](../hydrogen-web-control/SKILL.md).+So "type a 1-shot prompt into the sibling and let it run" needs the **user** to paste+the prompt (the one legitimate ask). Don't promise an autonomous sibling-driven run you+can't actually kick off.++## Decision rule++- **You have relay access too.** For almost everything, just DO the work yourself+ (drive Fusion + the browser directly), it's reliable and you control timing/captions.+ Delegating to a sibling adds a coordination + reliability tax for no gain.+- **Delegate to the Hydrogen sibling only when** the deliverable specifically needs to+ SHOW the AI being prompted (e.g. a demo of "type a prompt, AI does it"), AND the user+ pastes the prompt (or it's already running). Then you record + caption + post-produce.+- **Never run two drivers at once.** If a sibling is driving the apps via the relay,+ don't also drive them yourself, you'll collide. Pick one actor.++## If you ARE recording a sibling-driven run++CDP-driven browser actions happen in the BACKGROUND, so the recording won't show them+unless you foreground the right window as the sibling progresses. Practically, driving ++recording + captioning yourself (you control the timing) beats chasing an autonomous+sibling. See [screen-recording](../screen-recording/SKILL.md).
skills/desktop-ui-control/SKILL.mdadded+72@@ -0,0 +1,72 @@+---+name: desktop-ui-control+description: >+ Click, type, and read NATIVE app windows (Fusion, dialogs, non-web UI) on the+ user's desktop via Adom Desktop UIA and SendInput. Covers find_control/ui_click,+ click/type/press_key, screenshot-to-coordinate mapping, and window foreground.+ Trigger words: UIA, desktop_click, desktop_type, click a native window, drive a+ dialog, find_control, press_key, bring window to front, coordinate click.+---++# Desktop UI control (native windows)++For NON-web app windows (Fusion, OS dialogs, OAuth account choosers rendered outside+a reachable page), use the `desktop_*` verbs. Two mechanisms, pick by what works:++## 1. UIA, programmatic, background (prefer this)++```bash+adom-desktop desktop_list_windows '{}'+adom-desktop desktop_find_window '{"titleContains":"Fusion"}' # -> hwnd, rect+adom-desktop desktop_find_control '{"hwnd":<h>,"nameContains":"Continue"}' # -> best.{invokable,settable,rect}+adom-desktop desktop_ui_click '{"hwnd":<h>,"nameContains":"Continue"}' # InvokePattern, NO foreground/focus steal+adom-desktop desktop_ui_set '{"hwnd":<h>,"nameContains":"Name","value":"X"}' # SetValue+adom-desktop desktop_screenshot_window '{"hwnd":<h>}' # captures a window even in the BACKGROUND+```++`desktop_find_control` returns `best.invokable` / `best.settable`. If `invokable:true`,+`desktop_ui_click` works in the background (Invoke/SetValue are programmatic, no cursor+move, no focus steal). **Chromium/Edge expose their a11y tree, so page buttons are+reachable by accessible name**, but NOT shadow-DOM or webview-internal controls.++## 2. SendInput, foreground only (fallback)++When a control isn't in the UIA tree (shadow DOM, canvas, OAuth pages that block eval):++```bash+adom-desktop desktop_bring_to_front '{"hwnd":<h>}' # SetForegroundWindow+adom-desktop desktop_click '{"x":1622,"y":671}' # SendInput at SCREEN pixels+adom-desktop desktop_type '{"text":"hello","hwnd":<h>}' # WM_CHAR unicode into the focused control+adom-desktop desktop_press_key '{"hwnd":<h>,"keys":["ctrl","l"]}' # chords + named keys+```++### Screenshot -> screen-coordinate mapping (for desktop_click)++`desktop_screenshot_window` returns a scaled image (e.g. 1500x900) of a window whose+real rect (from `desktop_find_window`) is e.g. 2560x1528 at left=2,top=0. To click an+element you see at image `(ix,iy)`:++```+scale = rect.width / image.width+screen_x = rect.left + ix*scale+screen_y = rect.top + iy*scale+```++Overlay a coordinate grid on the screenshot to read `(ix,iy)` precisely, then click.+Proven driving a Google OAuth account chooser this way.++## What SendInput CANNOT do (hard-won)++- **OS input no-ops on many webview editors.** SendInput reaches a browser's NATIVE UI+ (address bar, search, Edge's command palette) but **does not register** in some webview+ text editors (Hydrogen's Claude input, VS-Code-in-browser editors). `desktop_type`,+ `desktop_press_key`, and click+type all silently fail there. ALWAYS screenshot to verify+ the text landed; if it didn't, the editor only accepts CDP-injected input.+- See [hydrogen-web-control](../hydrogen-web-control/SKILL.md) for that specific dead end.++## Foreground hygiene++`desktop_bring_to_front` / SetForegroundWindow steals focus, the user notices. Use it+ONLY for SendInput, never for UIA or screenshots (both work in the background). Don't+foreground a window the user might be using. The exception is RECORDING, where you must+re-foreground the target before each step (see [screen-recording](../screen-recording/SKILL.md)).
skills/hydrogen-web-control/SKILL.mdadded+65@@ -0,0 +1,65 @@+---+name: hydrogen-web-control+description: >+ What you CAN and (mostly) CANNOT drive in the Hydrogen web editor+ (hydrogen.adom.inc/<user>/<env>/edit) from the cloud, and why its embedded Claude+ Code panel input is unreachable by automation. Read before trying to type into a+ Hydrogen Claude tab. Trigger words: hydrogen editor, hydrogen web control, type+ into the claude tab, drive hydrogen, galliaApril editor, claude panel input.+---++# Hydrogen web control++The Hydrogen editor (`https://hydrogen.adom.inc/<user>/<env>/edit`) is a web app the+user runs in Chrome or Edge. You can reach its **top frame** via the browser extension+([native-browser-control](../native-browser-control/SKILL.md), remember+`"browser":"edge"` if it's in Edge). But its embedded **Claude Code panel input is+buried in nested cross-origin webviews and is NOT reachable by any automation** we have.+This skill exists so you don't burn a session rediscovering that.++## The structure (why the Claude input is unreachable)++```+hydrogen.adom.inc/<user>/<env>/edit <- top page (browser-extension reachable)+ └─ iframe <user>-<env>-....adom.cloud <- CROSS-ORIGIN (accessible:false)+ └─ iframe (VS Code / code-server) <- another layer+ └─ webview (Claude Code panel) <- the "ctrl esc to focus" input lives here+```++What the top frame exposes: a single `textarea` with placeholder **"Plan/Build your+project"** (Hydrogen's own launcher), NOT the Claude chat input. The Claude input is+several cross-origin frames down.++## What FAILS (all tested, all dead ends)++- **nbrowser/CDP eval** (Chrome AND Edge): top-frame only. It can't cross into the+ cross-origin iframe, and nbrowser has **no frame-targeting** (allFrames/frameUrl/frame+ params are ignored). `iframe.contentDocument` -> `accessible:false`.+- **OS SendInput** (`desktop_click` + `desktop_type` + `desktop_press_key`, and the+ "ctrl esc to focus Claude" shortcut): reaches the browser and Edge's NATIVE UI (your+ text lands in Edge's *search box*, not Claude) but does NOT register in the webview+ editor. Always screenshot, you'll see the placeholder unchanged.+- **UIA** (`desktop_find_control`): the webview's shadow DOM isn't exposed; returns the+ document fallback, not the input.++## What WORKS++- Reading/driving the **top frame** (the "Plan/Build your project" textarea, page chrome)+ via `nbrowser_eval`. Set a textarea value with the native setter:+ `Object.getOwnPropertyDescriptor(HTMLTextAreaElement.prototype,'value').set.call(ta, text)`+ then dispatch `input`.+- **Screenshotting** the Hydrogen window (`desktop_screenshot_window`, background).+- Navigating the top page.++## If you need the Hydrogen Claude to run a prompt++You can't type it in programmatically. Options:+1. **Ask the user to paste the prompt** into that Claude tab (it's the one case where+ asking is correct, you genuinely cannot do it). If it's a galliaApril-env Claude in+ bypass mode, it has relay access and will drive Fusion/the browser itself, see+ [claude-control](../claude-control/SKILL.md).+2. **Do the work yourself** instead of delegating to that sibling, you have the same+ relay access it does.++The real unlock (flag it to whoever owns `adom-browser-extension`): **frame-targeted+eval** so CDP can address a nested webview's execution context.
skills/native-browser-control/SKILL.mdadded+89@@ -0,0 +1,89 @@+---+name: native-browser-control+description: >+ Drive the user's REAL logged-in Chrome or Edge from the cloud via the Adom+ browser extension (CDP), in the background, no focus steal. Use to click/scroll/+ read a page in their actual browser, upload a file, or log them in through OAuth+ using their existing sessions. Trigger words: nbrowser, drive my chrome, drive+ edge, real browser, logged-in browser, upload to a website, sign in with google,+ oauth login, CDP eval browser.+---++# Native browser control (nbrowser)++`nbrowser_*` drives the user's **real, logged-in Chrome/Edge** through the+`adom-browser-extension` over CDP. Because it's their actual browser, their cookies+and logins are already there, so a JLCPCB/Autodesk/Google flow "just works" with no+credentials handed over. It runs in the **background** (no focus steal).++## Core verbs++```bash+adom-desktop nbrowser_status '{}' # {ok, ext, browser:"chrome"|"edge"}+adom-desktop nbrowser_list_tabs '{}' # all tabs (tabId, url, windowId, active)+adom-desktop nbrowser_open_tab '{"url":"https://..."}' # opens a DEDICATED tab, returns tabId+adom-desktop nbrowser_eval '{"tabId":"<id>","expr":"<js>"}' # run JS in that tab (background)+adom-desktop nbrowser_screenshot '{"tabId":"<id>"}' # page PNG as base64 (background)+adom-desktop nbrowser_navigate '{"tabId":"<id>","url":"..."}'+```++## Chrome vs Edge, this trips everyone++`nbrowser` defaults to **Chrome**. To drive Edge, pass `"browser":"edge"` on EVERY+verb. `nbrowser_list_tabs` without it shows only Chrome tabs.++```bash+adom-desktop nbrowser_list_tabs '{"browser":"edge"}'+adom-desktop nbrowser_eval '{"browser":"edge","tabId":"<id>","expr":"location.href"}'+```++## Gotchas (each cost a real debugging session)++- **It refuses to drive PWA / installed-app tabs** (e.g. a taskbar Gmail/Chat PWA) to+ avoid hijacking a standalone app. Open a **dedicated tab** with `nbrowser_open_tab`.+- **A first `eval` can return `null`** right after opening a tab (timing). Retry; it works.+- **No frame-targeting.** `eval` runs in the TOP frame only. Params like `allFrames`,+ `frameUrl`, `frame` are **ignored** (they silently return the top frame). You CANNOT+ reach into a **cross-origin iframe** (`iframe.contentDocument` throws / `accessible:false`).+ This is why deeply-nested webviews (the Hydrogen Claude panel) are unreachable, see+ [hydrogen-web-control](../hydrogen-web-control/SKILL.md).+- `nbrowser_eval` runs in the PAGE context, so JS can click, scroll, read, and set+ inputs, but only same-origin.++## Upload a file (DataTransfer, no OS picker)++CDP can't touch the native file dialog, so inject the bytes and set the+`<input type=file>` directly. Base64 the file, **stream it into the page in <40k+chunks** (the CLI arg cap), then build a `File` and set it via `DataTransfer`:++```bash+# per chunk: comma operator = ONE expression (nbrowser_eval requires a single expression)+adom-desktop nbrowser_eval '{"tabId":"<id>","expr":"(window.__b += \"<chunk>\", window.__b.length)"}'+# then decode + set the input:+adom-desktop nbrowser_eval '{"tabId":"<id>","expr":"(function(){var s=atob(window.__b),u=new Uint8Array(s.length);for(var i=0;i<s.length;i++)u[i]=s.charCodeAt(i);var f=new File([u],\"board.zip\",{type:\"application/zip\"});var inp=document.querySelectorAll(\"input[type=file]\")[0];var dt=new DataTransfer();dt.items.add(f);inp.files=dt.files;inp.dispatchEvent(new Event(\"change\",{bubbles:true}));return inp.files[0].name})()"}'+```++The input often **clears itself right after** the change event (the site read the file+already), that is NORMAL, don't treat the empty input as failure. Verify by screenshot.++## Drive an OAuth login yourself (NEVER ask the user)++You can log the user in end to end. Proven on JLCPCB "Sign in with Google":++1. Click the page's "Sign in with Google" button (`nbrowser_eval` `.click()`).+2. On `accounts.google.com`, **CDP eval WORKS** (it's not blocked). Click the user's+ account row, then the OAuth **Continue** button, both via `eval` `.click()`.+3. It redirects back logged in; any file you already uploaded persists in the session.++This is one-time per site; afterwards the session stays logged in and every run is+fully background. Pick the user's **primary work** Google account if there's a chooser.+NEVER ask the user to click or log in, you can do all of it.++## When CDP eval can't reach a control++If a page genuinely blocks eval on an element, screenshot the window, map the pixel,+and fall back to a foregrounded OS click, see [desktop-ui-control](../desktop-ui-control/SKILL.md).+But try eval first, it's background and doesn't disturb the user.++Related: [pup-browser-control](../pup-browser-control/SKILL.md) (anonymous browsers),+[screen-recording](../screen-recording/SKILL.md) (recording a live drive).
skills/pup-browser-control/SKILL.mdadded+54@@ -0,0 +1,54 @@+---+name: pup-browser-control+description: >+ Drive a headless/throwaway Puppeteer (pup) Chrome window via Adom Desktop, for+ tasks that need no login or that you want to record cleanly without touching the+ user's real browser. Trigger words: pup, puppeteer browser, headless chrome,+ anonymous browser, browser_eval, browser_open_window, record a browser tab.+---++# Pup browser control++`browser_*` (alias `pup`) drives a **separate Puppeteer-launched Chrome** on the+user's desktop, orthogonal to their real browser. Use it when **no login is needed**+(public pages, anonymous renders) or when you want a clean window to record.++## Core verbs++```bash+adom-desktop browser_open_window '{"sessionId":"x","profile":"x","url":"https://..."}'+adom-desktop browser_navigate '{"sessionId":"x","url":"..."}'+adom-desktop browser_eval '{"sessionId":"x","expr":"document.title"}'+adom-desktop browser_wait '{"sessionId":"x","ms":4000}'+adom-desktop browser_screenshot '{"sessionId":"x"}' # PNG saved locally on the CLI host (/tmp)+adom-desktop browser_record_start/_stop # CDP screencast (LOW fps, see below)+adom-desktop browser_close '{"sessionId":"x"}'+```++Sessions are keyed by `sessionId`; reuse it across calls.++## Upload via DataTransfer (same trick as nbrowser)++Pup is sandboxed from the local file input, so inject the bytes. Base64 the file,+stream into `window.__b` in <40k chunks via `(window.__b += "<chunk>", window.__b.length)`+(comma op = single expression), then decode to a `File` and set the input via+`DataTransfer` + dispatch `change`. Verified getting JLCPCB to render a board ZIP+anonymously, no login.++## Recording a pup tab, read this first++- `browser_record_start` uses **CDP `Page.startScreencast`**, which is **low-fps, ugly+ JPEG**. For a polished demo, prefer whole-desktop `desktop_record_*` (see+ [screen-recording](../screen-recording/SKILL.md)) over this.+- CDP screencast is **paint-throttled** when the tab isn't OS-foreground. If you must use+ it, `browser_focus_window` (Win32 SetForegroundWindow) **before every action**, calling+ it once isn't enough; anything that grabs focus re-throttles it. Sniff-test:+ `frameCount / (durationMs/1000)` should be near your requested fps; below ~12 fps means+ the window lost foreground.++## When to choose pup vs the real browser++- **No login needed / want it clean / don't touch the user's session** -> pup.+- **Needs the user's cookies/logins, or you'll push it to checkout** -> the real browser+ via [native-browser-control](../native-browser-control/SKILL.md).+- Common combo: anonymous board render/price in pup, the full logged-in order in nbrowser.
skills/screen-recording/SKILL.mdadded+69@@ -0,0 +1,69 @@+---+name: screen-recording+description: >+ Record the user's desktop to a video clip via Adom Desktop, with large on-screen+ captions, and the gotchas that make or break it (foreground re-bump, recorder+ reset, pulling the file). Trigger words: record the screen, screen recording,+ desktop_record, record a demo, capture the screen, on-screen captions,+ desktop_caption, record the browser flow.+---++# Screen recording (desktop capture + captions)++Capture the whole desktop to a `.webm` while you drive an app, with big AD-rendered+captions burned into the capture. Hard-won, every rule below comes from a failure.++## The verbs++```bash+# whole-desktop capture (getDisplayMedia, 30fps). recordingId is at the TOP LEVEL of the response.+adom-desktop desktop_record_start '{"confirmDesktopNotTabRecording":true,"reason":"<why>","fps":30,"audio":false,"monitor":"primary"}'+adom-desktop desktop_record_stop '{"recordingId":"rec-N"}' # -> filePath, durationMs, sizeKB (top level)+adom-desktop desktop_recorder_close '{}' # reset stale recorder state (see below)++# large always-on-top click-through caption, captured by the recording:+adom-desktop desktop_caption '{"text":"...","size":"large","position":"bottom","id":"d","persist":true}'+adom-desktop desktop_caption '{"action":"force-clear"}' # nuke ALL captions+```++For a single browser tab you can also `browser_record_*` (pup), but it's low-fps CDP+screencast, prefer desktop capture for quality. See+[pup-browser-control](../pup-browser-control/SKILL.md).++## The five rules that make it work++1. **Parse `recordingId` from the TOP LEVEL** of `desktop_record_start`'s response, NOT+ from a nested `output` field. Getting this wrong = empty RID = no recording.+2. **Reset the recorder first.** It goes stale (accumulated clips) and silently returns+ no recordingId / a 439-byte empty file. Call `desktop_recorder_close` + `sleep 2`+ before `desktop_record_start`. Run it inline / via the background tool, NOT detached+ with `&` (detached runs produced empty files).+3. **Foreground re-bump before EVERY visible step.** The capture shows whatever is+ topmost. If you drive via background CDP, the recording shows the WRONG window unless+ you `desktop_bring_to_front` the target before each click/scroll. This was THE fix for+ "my recording captured the wrong app." Re-bump, then act.+4. **Force-clear captions before each recording.** `persist:true` captions survive across+ recordings and leak a stale caption into the next clip. `desktop_caption {action:"force-clear"}`+ first. (Or caption in post instead, see below.)+5. **Live captions drift on autonomous flows.** If the page advances faster than your+ sleeps, the live caption lands on the wrong screen. For anything whose timing you don't+ control tightly, record CLEAN (no live captions) and add them in post, see+ [video-post-production](../video-post-production/SKILL.md).++## Pull the file to post-produce++```bash+adom-desktop pull_file '{"filePaths":["C:/Users/<u>/AppData/Local/Adom Desktop/plugins/puppeteer/recordings/rec-desktop-<ts>.webm"]}'+# NOTE: it lands in /tmp on the CLI host (your container), NOT your dest arg. Find it there.+```++## Privacy++Whole-desktop capture records EVERYTHING on the primary monitor (the user's other tabs/+apps). Maximize the target app so it dominates, and crop the taskbar / private tab strip+in post. The user must have authorized the recording (e.g. "I'm at the gym, go").++## Reliability note++The recorder is flaky. Verify after stop (`durationMs`, `sizeKB`), pull, and extract a+frame to CONFIRM the right window was captured before you build on it. Don't assume.
skills/video-post-production/SKILL.mdadded+100@@ -0,0 +1,100 @@+---+name: video-post-production+description: >+ Turn raw screen recordings + screenshots into a finished, narrated, captioned demo+ video with ffmpeg and the video-post tool: crop, speed up boring parts, brand-font+ captions, Ken Burns on stills, TTS voiceover, concat, and wiki-ready keyframes.+ Trigger words: speed up a video, edit a recording, make a demo video, add captions,+ voiceover, ffmpeg demo, video-post, concat clips, publish a demo to the wiki.+---++# Video post-production++Raw recordings + screenshots -> a polished demo. Pipeline: per-section clips ->+narration -> concat -> publish. Use [`adom-tts`](https://wiki.adom.inc/adom/adom-tts)+for voiceover and the `video-post` tool for review/manifest.++## Captions: brand font, burned in++Use the Adom brand font **Familjen Grotesk**, NOT a generic system font (DejaVu looks+cheap and got a demo rejected). A `.ttf` lives in the EDA hero assets; render captions+with ImageMagick or ffmpeg `drawtext`:++```bash+FONT=.../FamiljenGrotesk-Bold.ttf+DT="fontfile=$FONT:fontsize=42:fontcolor=0xeafdfb:box=1:[email protected]:boxborderw=22:x=(w-tw)/2:y=h-95"+# time a caption to a window of the clip:+-vf "drawtext=$DT:text='Every part matched to stock':enable='between(t,19,32)'"+```++Brand tokens: text `#eafdfb`, teal accent `#00e6dc`, caption band `#0d3b39`, dark+canvas `#0a0e14`.++## Real recording: crop, speed up, caption++```bash+# crop out the taskbar + any stuck caption / private tab strip, scale to 16:9, speed up,+# and burn in correctly-timed captions:+ffmpeg -ss 5 -t 51 -i raw.webm -vf "\+crop=2560:1410:0:0,scale=1920:-2,pad=1920:1080:0:(oh-ih)/2:color=0x0a0e14,\+setpts=PTS/1.6,\+drawtext=$DT:text='...':enable='between(t,0,8)',drawtext=$DT:text='...':enable='between(t,8,13)'\+" -an -r 30 -c:v libx264 -crf 21 -pix_fmt yuv420p -g 30 jlc-live.mp4+```++`setpts=PTS/1.6` speeds video 1.6x (the "speed up boring parts"); drawtext `enable`+times are in the OUTPUT (sped) timeline. Recordings have no audio if you captured+`audio:false`.++## Stills -> clips (Ken Burns + caption + narration)++For steps you have as screenshots (or that live recording couldn't reach cleanly):+composite the shot on the dark canvas with a caption band (ImageMagick), then animate+with a gentle zoompan and mux the TTS:++```bash+# zoompan on a still, pre-upscale to avoid jitter:+[0:v]scale=2560:1440,zoompan=z='min(zoom+0.0005,1.06)':d=<frames>:fps=30:s=1920x1080,format=yuv420p+# clip length = TTS duration + 0.6s; mux: -map "[v]" -map 1:a -t <dur> ... -c:a aac+```++## Voiceover++```bash+adom-tts say "Spoken sentence, numbers spelled out like four layer." --out tts/step.mp3+```++House voice + pronunciation cache (see [adom-tts](https://wiki.adom.inc/adom/adom-tts) /+`tts-pronunciation`). Mux per-clip; if narration is shorter than the clip, that's fine+(audio ends, video continues).++## Concat + wiki keyframes++Build all clips with IDENTICAL params (1920x1080, libx264, yuv420p, 30fps, aac+44100/2), then stream-copy concat (fast, lossless):++```bash+printf "file '%s'\n" clips/*.mp4 > list.txt+ffmpeg -f concat -safe 0 -i list.txt -c copy final.mp4+```++**Every per-clip encode MUST include `-g 30 -keyint_min 30`** (1s keyframes) or the wiki+`<video>` can't scrub. Verify: `ffprobe ... -show_entries packet=pts_time,flags | grep ,K`+should show keyframes ~1s apart.++## Review + publish++- `video-post manifest add <manifest> --id X --raw clip.mp4 --description "<intent>" --narration "<spoken>"`+- `video-post app <manifest> --port <p>` opens a Hydrogen review UI (it prints a proxy URL,+ open THAT in pup, not localhost). **Don't launch it with a short `timeout`, the server+ dies and the player 500s.** Skip the review only if the user said "just ship it".+- Publish to a wiki page: push the `.mp4` into the page repo (`adom-wiki repo push`) and+ embed it in the README with an HTML5 `<video><source ...></video>` pointing at+ `https://wiki.adom.inc/api/v1/pages/<slug>/files/<path>`. (`video-post publish` currently+ wraps a removed `adom-wiki asset` command, push the file directly instead.)++## Building from stills vs real recording++If clean live recording isn't feasible, a video built from REAL screenshots with motion ++captions + voiceover is a legitimate fallback, just be honest about which beats are stills.+Real screen footage is better when you can get it; see [screen-recording](../screen-recording/SKILL.md).
uninstall.shadded+7@@ -0,0 +1,7 @@+#!/usr/bin/env bash+set -e+DEST="$HOME/.claude/skills"+for s in adom-desktop-control native-browser-control pup-browser-control desktop-ui-control hydrogen-web-control claude-control screen-recording video-post-production; do+ rm -rf "$DEST/$s"+done+echo "Removed adom-desktop-control + sub-skills."