Fix silent watcher death: self-healing, observability, review fixes (v0.6.0) #1
Loading…
Reference in a new issue
No description provided.
Delete branch "fix/silent-watcher-death"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Summary
Fixes the "silent watcher death" incident (2026-08-07, katelyn): a notify event with empty
pathspanicked notify-debouncer-full <= 0.5.0 while the debouncer mutex was held, killing both watcher threads. Commits fired only from notify events and pull refuses dirty trees, so the daemon deadlocked silently for 6 weeks — with push failures hidden atwarn!level on top.Beyond the dep pin, this PR removes the entire class of failure: the daemon now detects stuck states and recovers, git operations can't hang the loop, and health is observable on demand.
Commits
fix: self-heal dead watchers and surface git failuresprocess_debounce_result,PullOutcome, error-level failure logging, info default verbosity, commit sha loggingchore(release): v0.6.0fix: harden config reload, config watch and push racespull --rebaseon non-fast-forward (katelyn/pheobe race on foambubble)feat: status health subcommand, daemon heartbeat, recurring-log dedupgit_afk status(daemon PID, watcher generation, heartbeat freshness, per-repo live git facts; non-zero exit when daemon missing/stale), heartbeat at$XDG_RUNTIME_DIR/git_afk/status.jsonevery 10 s tick, identical repeating failures log once per hourfix: address independent review — gitignore filter, timeouts, git edge casesGitignoreBuilder::newnever reads the file — now added +matched_path_or_any_parents), 120 s timeout +kill_on_dropon every git subprocess, mid-commit changes preserved, failed add/commit retries next tick, push heuristic no longer matches bare "rejected" + conflicted auto-rebase aborted, detached-HEAD guard, failed watches retried every 60 s, per-repo config skip on broken entries, CLI path errors instead of panics, per-uid /tmp status fallbackPlus
AGENTS.md,.gitignore, README ops notes (self-healing, log levels, trust boundary,Restart=on-failure).Verification
.git/events, detached HEAD, config reload keeps watchers on broken TOML, heartbeat freshness/roundtripstatusfresh/stale/missing-vvvremoved — journal is info-level now); katelyn still runs the vulnerable Jan-2025 build and needs the musl artifact from the v0.6.0 release after mergeReview focus
src/watcher.rs— state machine (tick_action/note_*), recovery + rebuild cooldown, watch-retry loop, heartbeat writesrc/git.rs—run_gittimeout wrapper,CommitOutcome/PullOutcome, push-retry heuristic + rebase-abortsrc/status.rs— new module, snapshot formatDeferred backlog (documented, not blocking): per-repo rebuild cooldown, reload serialization, config-deletion semantics, spawn_blocking profiling.
— Michal's Pi Agent
Fixes the "silent watcher death" outage (2026-08-07, server katelyn): a notify event with empty paths panicked notify-debouncer-full <= 0.5.0 while the debouncer mutex was held, poisoning it and silently killing both watcher threads. No events were delivered afterwards, and since commits only fire on notify events while pull() refuses dirty trees, the daemon deadlocked forever - 6 weeks of uncommitted changes with only a repeating "Uncommitted changes present" error in the journal. - deps: document the notify-debouncer-full >= 0.6.0 requirement in Cargo.toml (0.6.0 skips path-less events; do not downgrade) - watcher: skip path-less events in handle_watch_event as defense in depth; extract process_debounce_result() so the debouncer callback is testable and never panics on unconfigured repositories - watcher: dirty-tree recovery - after DIRTY_PULL_RECOVERY_THRESHOLD consecutive pulls refused by a dirty tree with no commit in between, force commit_and_push and rebuild all watchers (rate-limited by a 10-minute cooldown). Breaks the dirty-tree/pull deadlock independent of notify internals. - watcher: watch failures are logged and skipped instead of panicking (a panic during config reload left the daemon with zero watchers) - git: pull() now reports PullOutcome {Clean, Dirty, UnusableState} so the state machine can detect stuck trees - logging: push/pull/commit failures at error! (were warn!, invisible at the old default level - this masked a rejected non-fast-forward push during the incident); successful commits log repo + short sha; default verbosity raised to info so daemon health is answerable from the journal; notify errors use error! instead of println! - tests: regression tests for path-less events through the debounce handler, the dirty-tree recovery state machine, and git outcomes (dirty/clean/unusable pull, commit creates a commit) on real temp repositories