main
History Download
Ray Publish 0.1.0 4cd2f07 15d ago

name: adom-bluegreen description: Zero-failed-request deploys for an Adom service container. nginx front door, blue and green slots as git worktrees, a flip only when the new build reports ready at the new commit, and the old slot stopped only after it has served every request it accepted. Use when a service restarts in place on every push (and is down while it warms), when someone asks for blue-green, zero downtime, rolling or hitless deploys, or when a service needs a /ready endpoint. Trigger words - blue-green, blue green deploy, zero downtime, zero failed requests, hitless deploy, service down while it updates, watchdog restart, /ready endpoint, readiness probe, drain in-flight requests, adom-bluegreen.

adom-bluegreen

A service container that pulls origin/main and restarts in place is down for as long as the new process takes to warm up. Component Ocean took 60-120 s to load its data, so every push was a minute or two of outage for every caller. Ray, 2026-09-22: "for anything production-dependent I consider this unacceptable ... it has to be 0 failed requests."

This kit is what took CO to zero. Measured on the live service: a cron-driven deploy had 0 failures in 1,372 requests and a manual one 0 in 2,365. The kit's own self-test measured 0 in 5,709 across a deploy and 0 in 7,659 while a broken build was refused.

How it works

 public URL ──> :PUBLIC_PORT  nginx (front door; reloads never drop a request)
                     ├── :BLUE_PORT    blue slot   (git worktree at commit A)
                     └── :GREEN_PORT   green slot  (git worktree at commit B)

Every 2 minutes cron runs bluegreen.sh <service.conf>. When origin/BRANCH has a new commit:

  1. It is checked out and started in the idle slot, while the live slot keeps serving.
  2. The front door flips only when the new slot's /ready answers 200 with that exact commit. Anything else (it crashed, it never warmed, it runs the wrong commit) means no flip. The commit is recorded in failed_sha and never retried. The old slot keeps serving.
  3. nginx reloads gracefully. New requests go to the new slot and requests already in progress finish where they are.
  4. The old slot is stopped only once it has finished every request it accepted (its /ready reports inflight 0 for 5 s).

If the live slot dies while the other is warm, traffic fails over to it. The only remaining outage is a true cold start (container restart), when nothing is running anyway.

Adopting it on a service (about 15 minutes)

  1. Add the /ready contract to the service. Copy the pattern from examples/ (Bun, Node, Python):
    GET /ready  ->  200 {"ready":true,  "git_sha":"<$BLUEGREEN_SHA>", "inflight":<n>}   when warm
                ->  503 {"ready":false, ...}                                          while warming
    
    • Listen on $PORT; the kit sets it per slot.
    • git_sha echoes $BLUEGREEN_SHA; the kit sets it too, so the app never has to read .git.
    • ready must mean everything callers need is loaded, not just "the port is open". CO's first version flipped before its passive value index was attached, and value filters would have answered 503.
    • inflight counts requests in progress, excluding /ready. It's optional; without it the old slot gets a timed 90 s drain.
    • Push this first and let the old deploy mechanism ship it.
  2. Write the config from examples/service.conf.example and keep it under /home: NAME, REPO_DIR, BRANCH, the three ports, START_CMD (one plain command, run in the slot directory), INSTALL_CMD, and SHARED (gitignored state such as data, symlinked into both slots).
  3. Stop the old auto-deploy (its cron line or watchdog) so two deployers don't fight.
  4. Install: bash scripts/install.sh ~/my-service.conf, adding --boot-hook on images whose entrypoint only starts sshd. It checks the config, installs nginx and schedules the watchdog.
  5. Adopt the port:
    • If the service already listens on PUBLIC_PORT, set LEGACY_STOP_CMD to how to stop it.
    • The first pass warms a slot, stops the legacy process and starts nginx in its place.
    • That handoff is the one moment two processes would need the same port, so there is a gap of milliseconds. Do it during a container restart (nothing is listening then anyway) or at a quiet time. Every deploy after it is zero-failure.
  6. Prove it: start scripts/probe.sh http://127.0.0.1:PUBLIC_PORT/ 300 4 on the container, push a commit, and read the last line. It must say 0 failed.

Check before you trust it on a new container: bash tests/selftest.sh 3490 runs a real deploy of a throwaway app on spare ports (cold start, deploy under load with slow requests, a never-ready build) and must end 10 passed, 0 failed. It touches nothing else.

Operating it

where traffic goes bash scripts/bluegreen.sh <conf> status, or curl :PUBLIC_PORT/__bluegreen
is my push live? /__bluegreen shows its sha, usually 2-4 min after the push
logs ~/.bluegreen/<NAME>/bluegreen.log, blue.log, green.log, nginx-error.log
a deploy failed ~/.bluegreen/<NAME>/failed_sha; read <slot>.log; push a fix (a new sha retries)
roll back git revert and push: a rollback is just another zero-failure deploy
stop deploying remove the cron line; both slots and nginx keep serving

Traps this kit already handles (do not reintroduce them)

  • Never pkill -f. It matches the watchdog's own command line and kills both slots. Processes are found through /proc by working directory and exact command.
  • $! after cd dir && cmd & is a wrapper shell, not the server. Killing it leaves the server holding its port, and the next deploy into that slot fails.
  • Every launched process must drop the lock fd (exec 9>&- first in its subshell). A child that inherits it holds the lock for life, every later pass exits at flock, and deploys silently stop with no error anywhere. CO hit this twice.
  • git reset rewrites the running watchdog script under bash, which reads it by offset. The script re-runs itself from a private copy first.
  • Never replace a front door that is serving. Two processes cannot share a port, so any handoff is a gap. nginx changes its config by reload, keeping the socket. A new front door is adopted only at a cold start.
  • Don't name a bash array PORT. Bash cannot export an array, so export PORT=... silently does nothing and the app starts without its port. The self-test caught this.
  • What the container cannot do: there's no NET_ADMIN and /proc/sys is read-only, so there's no iptables and no tcp_migrate_req. The edge port mapping can be added or removed but never re-pointed. That is why the front door is a long-lived nginx on the mapped port.
  • nginx lives in /usr, which a container rebuild resets. The watchdog reinstalls it at the next cold start. Keep the config and the service's state under /home.