Fix silent watcher death: self-healing, observability, review fixes (v0.6.0) #1

Merged
michalvankodev merged 7 commits from fix/silent-watcher-death into main 2026-09-25 10:19:41 +02:00

Summary

Fixes the "silent watcher death" incident (2026-08-07, katelyn): a notify event with empty paths panicked notify-debouncer-full <= 0.5.0 while the debouncer mutex was held, killing both watcher threads. Commits fired only from notify events and pull refuses dirty trees, so the daemon deadlocked silently for 6 weeks — with push failures hidden at warn! level on top.

Beyond the dep pin, this PR removes the entire class of failure: the daemon now detects stuck states and recovers, git operations can't hang the loop, and health is observable on demand.

Commits

Commit What
fix: self-heal dead watchers and surface git failures dep pin documented (notify-debouncer-full >= 0.6.0), dirty-tree recovery (2 dirty-refused pulls → forced commit + watcher rebuild, 10 min cooldown), path-less event guard, testable process_debounce_result, PullOutcome, error-level failure logging, info default verbosity, commit sha logging
chore(release): v0.6.0 version bump
fix: harden config reload, config watch and push races reload parses new state before clearing watchers (broken mid-write config keeps old watchers — verified live), config watcher watches parent dir (survives atomic-rename writes), push retry after pull --rebase on non-fast-forward (katelyn/pheobe race on foambubble)
feat: status health subcommand, daemon heartbeat, recurring-log dedup git_afk status (daemon PID, watcher generation, heartbeat freshness, per-repo live git facts; non-zero exit when daemon missing/stale), heartbeat at $XDG_RUNTIME_DIR/git_afk/status.json every 10 s tick, identical repeating failures log once per hour
fix: address independent review — gitignore filter, timeouts, git edge cases fresh-context review-agent findings: .gitignore filter was completely dead (GitignoreBuilder::new never reads the file — now added + matched_path_or_any_parents), 120 s timeout + kill_on_drop on every git subprocess, mid-commit changes preserved, failed add/commit retries next tick, push heuristic no longer matches bare "rejected" + conflicted auto-rebase aborted, detached-HEAD guard, failed watches retried every 60 s, per-repo config skip on broken entries, CLI path errors instead of panics, per-uid /tmp status fallback

Plus AGENTS.md, .gitignore, README ops notes (self-healing, log levels, trust boundary, Restart=on-failure).

Verification

  • 27 tests (unit + real-git integration): incident-shaped empty-paths events through the debounce handler, dirty-tree recovery state machine and full wiring on a real repo, two-server push race with a bare origin, gitignore filtering regression, .git/ events, detached HEAD, config reload keeps watchers on broken TOML, heartbeat freshness/roundtrip
  • clippy clean; live smoke tests: broken-config survival, hot config reload, ignored churn not resetting debounce, status fresh/stale/missing
  • Deployed on pheobe since the first fix commit (systemd unit -vvv removed — journal is info-level now); katelyn still runs the vulnerable Jan-2025 build and needs the musl artifact from the v0.6.0 release after merge

Review focus

  • src/watcher.rs — state machine (tick_action/note_*), recovery + rebuild cooldown, watch-retry loop, heartbeat write
  • src/git.rs — run_git timeout wrapper, CommitOutcome/PullOutcome, push-retry heuristic + rebase-abort
  • src/status.rs — new module, snapshot format

Deferred backlog (documented, not blocking): per-repo rebuild cooldown, reload serialization, config-deletion semantics, spawn_blocking profiling.

— Michal's Pi Agent

## Summary Fixes the "silent watcher death" incident (2026-08-07, katelyn): a notify event with empty `paths` panicked notify-debouncer-full <= 0.5.0 while the debouncer mutex was held, killing both watcher threads. Commits fired only from notify events and pull refuses dirty trees, so the daemon deadlocked silently for 6 weeks — with push failures hidden at `warn!` level on top. Beyond the dep pin, this PR removes the entire *class* of failure: the daemon now detects stuck states and recovers, git operations can't hang the loop, and health is observable on demand. ## Commits | Commit | What | |---|---| | `fix: self-heal dead watchers and surface git failures` | dep pin documented (notify-debouncer-full >= 0.6.0), dirty-tree recovery (2 dirty-refused pulls → forced commit + watcher rebuild, 10 min cooldown), path-less event guard, testable `process_debounce_result`, `PullOutcome`, error-level failure logging, info default verbosity, commit sha logging | | `chore(release): v0.6.0` | version bump | | `fix: harden config reload, config watch and push races` | reload parses new state **before** clearing watchers (broken mid-write config keeps old watchers — verified live), config watcher watches parent dir (survives atomic-rename writes), push retry after `pull --rebase` on non-fast-forward (katelyn/pheobe race on foambubble) | | `feat: status health subcommand, daemon heartbeat, recurring-log dedup` | `git_afk status` (daemon PID, watcher generation, heartbeat freshness, per-repo live git facts; non-zero exit when daemon missing/stale), heartbeat at `$XDG_RUNTIME_DIR/git_afk/status.json` every 10 s tick, identical repeating failures log once per hour | | `fix: address independent review — gitignore filter, timeouts, git edge cases` | fresh-context review-agent findings: **.gitignore filter was completely dead** (`GitignoreBuilder::new` never reads the file — now added + `matched_path_or_any_parents`), 120 s timeout + `kill_on_drop` on every git subprocess, mid-commit changes preserved, failed add/commit retries next tick, push heuristic no longer matches bare "rejected" + conflicted auto-rebase aborted, detached-HEAD guard, failed watches retried every 60 s, per-repo config skip on broken entries, CLI path errors instead of panics, per-uid /tmp status fallback | Plus `AGENTS.md`, `.gitignore`, README ops notes (self-healing, log levels, trust boundary, `Restart=on-failure`). ## Verification - **27 tests** (unit + real-git integration): incident-shaped empty-paths events through the debounce handler, dirty-tree recovery state machine **and** full wiring on a real repo, two-server push race with a bare origin, gitignore filtering regression, `.git/` events, detached HEAD, config reload keeps watchers on broken TOML, heartbeat freshness/roundtrip - clippy clean; live smoke tests: broken-config survival, hot config reload, ignored churn not resetting debounce, `status` fresh/stale/missing - **Deployed on pheobe** since the first fix commit (systemd unit `-vvv` removed — journal is info-level now); katelyn still runs the vulnerable Jan-2025 build and needs the musl artifact from the v0.6.0 release after merge ## Review focus - `src/watcher.rs` — state machine (`tick_action`/`note_*`), recovery + rebuild cooldown, watch-retry loop, heartbeat write - `src/git.rs` — `run_git` timeout wrapper, `CommitOutcome`/`PullOutcome`, push-retry heuristic + rebase-abort - `src/status.rs` — new module, snapshot format Deferred backlog (documented, not blocking): per-repo rebuild cooldown, reload serialization, config-deletion semantics, spawn_blocking profiling. — Michal's Pi Agent
Fixes the "silent watcher death" outage (2026-08-07, server katelyn):
a notify event with empty paths panicked notify-debouncer-full <= 0.5.0
while the debouncer mutex was held, poisoning it and silently killing
both watcher threads. No events were delivered afterwards, and since
commits only fire on notify events while pull() refuses dirty trees,
the daemon deadlocked forever - 6 weeks of uncommitted changes with
only a repeating "Uncommitted changes present" error in the journal.

- deps: document the notify-debouncer-full >= 0.6.0 requirement in
  Cargo.toml (0.6.0 skips path-less events; do not downgrade)
- watcher: skip path-less events in handle_watch_event as defense in
  depth; extract process_debounce_result() so the debouncer callback is
  testable and never panics on unconfigured repositories
- watcher: dirty-tree recovery - after DIRTY_PULL_RECOVERY_THRESHOLD
  consecutive pulls refused by a dirty tree with no commit in between,
  force commit_and_push and rebuild all watchers (rate-limited by a
  10-minute cooldown). Breaks the dirty-tree/pull deadlock independent
  of notify internals.
- watcher: watch failures are logged and skipped instead of panicking
  (a panic during config reload left the daemon with zero watchers)
- git: pull() now reports PullOutcome {Clean, Dirty, UnusableState} so
  the state machine can detect stuck trees
- logging: push/pull/commit failures at error! (were warn!, invisible
  at the old default level - this masked a rejected non-fast-forward
  push during the incident); successful commits log repo + short sha;
  default verbosity raised to info so daemon health is answerable from
  the journal; notify errors use error! instead of println!
- tests: regression tests for path-less events through the debounce
  handler, the dirty-tree recovery state machine, and git outcomes
  (dirty/clean/unusable pull, commit creates a commit) on real temp
  repositories
Includes the notify-debouncer-full 0.6.0 pin (empty-paths panic fix) and the watcher self-healing changes. Deployed binaries on katelyn and phoebe run a vulnerable Jan-2025 build and must be upgraded to this release.
Found during the post-incident review, all verified with live daemon
smoke tests:

- reload_watchers cleared the existing debouncers BEFORE parsing the new
  configuration. A config reload triggered while the file was mid-write
  (invalid TOML) panicked the reload task and left the daemon with zero
  watchers - permanently deaf, same failure shape as the 2026-08
  incident, different trigger. The new state is now built first; on any
  parse/gitignore error the running watchers are kept and the failure is
  logged. Verified end-to-end: the old binary panicked on a mid-write
  config, the new one logs "Could not reload configuration, keeping
  existing watchers" and keeps committing.
- parse_config is fallible and skips non-UTF-8 paths instead of
  panicking; startup failures propagate as a clean process error.
- the config watcher now watches the parent directory and filters events
  by file name: an atomic-rename write (editors, tools) would silently
  detach a watch bound to the old inode.
- git_push recovers from rejected non-fast-forward pushes (katelyn and
  pheobe both watch the Syncthing-synced foambubble repo and race each
  other): on a rejection that looks like the remote moved ahead, pull
  --rebase and retry the push once instead of leaving a diverged branch
  until the next pull cycle. Covered by an integration test with a bare
  origin and two racing clones.
- README: recommend Restart=on-failure in the systemd unit.
- `git_afk status`: prints daemon health (pid, watcher generation,
  heartbeat freshness) and per-repo facts merged live from git (dirty
  tree + branch + last commit, rebase/merge state) with the daemon's
  internal view (pending change, next pull, dirty-pull streak). Exits
  non-zero when the daemon is missing or its heartbeat is stale - the
  "silently deaf daemon" failure of the 2026-08 incident is now
  detectable on demand.
- the daemon publishes a status heartbeat
  ($XDG_RUNTIME_DIR/git_afk/status.json) every 10s tick; a snapshot
  older than 60s means several missed ticks -> stuck loop.
- watcher generation counter increments on every watcher rebuild
  (config reload or dirty-tree recovery) and is published in the
  heartbeat, so repeated recovery is visible.
- recurring-log dedup (log_recurring): identical repeating failures
  (repo without upstream, stuck rebase, unreachable remote, uncleanable
  tree) are logged once per hour with a debug-level heartbeat instead
  of every debounce cycle; a changed message always logs and resets the
  timer. Applied to pull/push/commit failures, dirty-pull refusals,
  notify errors, watcher and config reload failures.

Verified live: status against a running daemon (fresh/stale/missing),
19 unit tests, clippy clean.
Findings from a fresh-context review agent (triaged: 1 blocker, 8
should-fix, cheap backlog items; all fixed here, verified live):

- BLOCKER: the .gitignore filter was completely non-functional —
  GitignoreBuilder::new() builds an EMPTY matcher and never reads the
  repo's .gitignore without an explicit add(). Ignored churn (build
  output, syncthing temp files) kept resetting the debounce and
  masqueraded as "watcher probably dead" via the recovery path. The
  root .gitignore is now added (missing file tolerated) and matched
  with matched_path_or_any_parents so directory patterns cover their
  contents.
- every git subprocess now has a 120s timeout with kill_on_drop: one
  hung push (stalled ssh, pinentry on a tty-less daemon) can no longer
  stall the whole maintenance loop and stale the heartbeat.
- commit bookkeeping: commit_and_push reports CommitOutcome; failed
  add/commit keeps the pending change so the next tick retries instead
  of waiting for dirty-tree recovery; a change observed DURING a commit
  is preserved (note_commit_finished(started_at)).
- push-retry heuristic no longer matches a bare "rejected" (permission/
  protected-branch rejections must not trigger a rebase); a conflicted
  auto-rebase is aborted so the repository stays usable.
- detached HEAD repositories are never auto-committed (commits would be
  unreachable and collectable).
- watches that fail to install are remembered and retried every 60s —
  a repo missing at startup with a clean tree would otherwise stay deaf
  forever (dirty-tree recovery never fires on clean trees).
- one broken config entry (unreadable .gitignore) no longer aborts
  daemon start/reload: per-repo log+skip.
- CLI: `git_afk add` with an invalid/nonexistent path returns an error
  instead of panicking.
- status file: per-uid /tmp fallback when XDG_RUNTIME_DIR is unset
  (avoids predictable shared /tmp paths); README notes the git-hooks
  trust boundary; .gitignore added (target/, heaptrack*); assorted
  clippy hygiene.

Tests: 26 passing — new coverage for gitignored-path filtering, .git/
events, the full dirty-recovery wiring on a real repository, detached
HEAD, the push-rejection heuristic, and the recurring-log window.
The journal now answers 'what is this daemon actually watching?' at info
level, including repositories whose watch setup failed (with the retry
cadence) — previously the scope was only visible in the config file.
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
michalvankodev/git_afk!1
No description provided.