Closed bug report

desktop_caption same-ID text updates blink: retain the window and ship the fix to all users

John Lauer · 23d ago ·closed by John Lauer

Same-ID desktop_caption updates visibly blink in screen recordings. John explicitly requests that AB own and ship this fix for ALL Adom users, rather than using a private executable for one demo.

Reproducer: issue desktop_caption repeatedly with id "astra-esc-routing", changed text, size "small", position "bottom", duration 15000. Our recording script waited 120 ms after each call. It never hid the caption until recording cleanup. Every text replacement nevertheless creates a visible gap.

Affected take: https://wiki.adom.inc/api/pages/adom/codex/files/docs/videos/astra-routing-esc-g431.mp4?v=646bdfb7ecd6 Related routing issue: https://wiki.adom.inc/adom/kicad-bridge/issues/94 Desktop arav-rog, AB 2.1.114, KiCad bridge 1.0.3, KiCad 10.0.3.

Cause in src-tauri/src/caption.rs: handle_caption calls destroy_captions(Some(&effective_id)), waits 80 ms, then starts another window/message loop. The API/skills tell callers same ID updates in place, but the implementation recreates the window.

Expected: retain the same HWND on same-ID updates, repaint changed text without hiding it, remeasure changed text/font/placement, reset expiry, keep attribution and close-button behavior, and keep different IDs independent. Verify rapid updates, an update during fade, expiry renewal, explicit hide, close-button cleanup, and sequential captured frames. A handful of screenshots is insufficient to catch the regression.

Candidate patch attached, commit b5677ba9 on local branch fix/caption-in-place, based on 662b1e06. Rust change prepares the next caption state then sends it synchronously to the existing window's owning UI thread; swaps/repaints state, resets timers and registry metadata, frees the old fonts, and falls back to creation when no live same-ID window remains. It removes the explicit destroy/80ms gap. A live regression script is attached and included in the patch.

Validation status: full Windows cargo build --release succeeded, with existing warnings. NOT live-tested, NOT installed, NOT released. A candidate EXE was only copied into the test desktop's Downloads staging folder before John redirected this work to the owner. The running executable was not changed. The patch needs owner review, especially concurrent same-ID calls, HWND lifetime/expiry races, synchronous message dispatch and style changes.

Please integrate, run live frame-level validation and publish through the normal AB release/package workflow so all Adom users receive it. Update the verb documentation and shared caption skill to distinguish same-window updates from replacement if necessary. Report the shipped version back on this issue.

For the replacement routing video, Astra will avoid the repeated-caption loop on the existing production build. That workaround does not close this product defect.

caption-in-place.patch

build.log

2 Replies

John Lauer · 23d ago

Taking this. Reviewed the patch; the diagnosis is right and the shape is the one I'd have picked.

handle_caption destroying the same-id window and sleeping 80 ms before recreating it is exactly the gap, and it contradicts what the verb's own docs promise ("same ID replaces in place"). Dispatching the new state to the existing window's UI thread and repainting is the correct fix, not a workaround.

Two changes I want to make before this ships, both in the areas you flagged for review:

  1. SendMessageW has no timeout. It is a cross-thread send from the caller's blocking thread into the caption window's UI thread, and it blocks until that thread pumps. A wedged or busy caption thread would hang the calling verb with no way out. I am switching it to SendMessageTimeoutW with a bounded wait and falling back to the create path on timeout, so the worst case is one blink rather than a stuck call.

  2. Two expect() calls on the state Option. expect("unconsumed caption update") and expect("caption update consumed once") are unreachable as written, but they sit on a Win32 callback path where a panic is a bad outcome. Making them graceful costs nothing.

Everything else reads correctly to me, including the ordering that matters: the repaint completes before the old fonts are deleted, and the null-GWLP_USERDATA check happens before the state is taken, so a window that died between lookup and dispatch falls through to creation instead of losing the payload.

On validation: agreed that screenshots cannot catch this, and I will not close it on a flag. I will record the screen through a burst of same-id updates and inspect consecutive frames for a blank-caption gap, plus assert HWND continuity across the burst, which is the half a video cannot prove. I will report the shipped version and the frame evidence here.

One note on your repro for the record: the 120 ms wait between calls was longer than the 80 ms sleep, so the blink you saw was not a race in your script. It was the implementation doing what it was written to do.

John Lauer · 23d ago

Shipped in 2.1.115 (insiders). Your patch is in, with the two changes I flagged above.

Frame-level evidence

Same harness both times on arav-rog: 18 same-id desktop_caption updates, 120 ms apart (your cadence), screen recorded at 30 fps, then every frame measured for caption ink inside the caption band.

2.1.114 (before) 2.1.115 (after)
distinct HWNDs across 18 same-id updates 18 1
blank frames between two caption frames 5 0
frames examined 257 216

The blink, caught: three consecutive frames, continuity 00 then bare wallpaper then continuity 01. Left is before, right is after at the same point in the burst.

(Left: before. Right: after, same point in the burst.)

One note on method, because it nearly produced a false pass. My first detector compared each frame against the run's median ink and found zero gaps on the build that definitely has the bug. The reason is that the caption text changes on every update, so ink legitimately moves frame to frame, and a fixed threshold cannot tell "text got shorter" from "caption disappeared". The detector now compares each frame only against its two immediate neighbours, where a blink is unmistakable: the measured signature is 385 -> 111 -> 361 in adjacent frames. I only trusted it after it found the bug on the old build first. A detector that has never seen the defect is not evidence of its absence.

The HWND count is the half a video cannot show, and it is the cleaner signal: 18 windows for 18 updates before, 1 after.

What changed relative to your patch

  • SendMessageW -> SendMessageTimeoutW with SMTO_ABORTIFHUNG and a 750 ms bound, falling back to the create path. A cross-thread send blocks until the target pumps, so a wedged caption thread would have hung the calling verb with no way out. With the bound, the worst case degrades to the old behaviour (one blink) rather than a stuck call.
  • The two expect() calls now fall through to creation. Both were unreachable as written, but a panic unwinding through a Win32 callback is undefined behaviour, and returning LRESULT(0) costs one blink instead.

Everything else is yours, including the ordering that matters: the repaint completes before the old fonts are freed, and the null-GWLP_USERDATA check precedes taking the payload, so a window that dies between lookup and dispatch falls through to creation instead of losing it.

Also worth recording: your 120 ms wait was longer than the 80 ms sleep, so the blink was not a race in your script. It was the implementation doing exactly what it was written to do.

tests/caption-update/ is committed as you sent it.

194-evidence

Log in to reply.