Files
skills/.wiki/concepts/session-inbox-monitor-received-msg-fp.md
vitya f7c04cf8cc fix(inbox): .claude-inbox → .agents/inbox cascade (rename idea 13 gap)
- skills: inter-session-peer-discipline v0.1.3, delegate-task v0.2.6, task-format v0.1.1
- hermes/mapping.yaml: monitor path
- .wiki/concepts/session-inbox-monitor-received-msg-fp.md: trigger phrase path
- dist rebuilt, lint 47/0, installed to ~/.claude/skills + ~/.agents/skills
2026-08-13 11:59:27 +03:00

4.6 KiB
Raw Blame History

title, type, tags, updated
title type tags updated
session-inbox-monitor — a routed negative only competes if its sibling is installed concept
skill-triggers
false-positive
trigger-discrimination
test-trigger
2026-06-17

session-inbox-monitor — a routed negative only competes if its sibling is installed

Sibling of delegate-task-negative-trigger-fp. Same failure family (a skill false-positive-fires on a phrase its description tries to exclude), but a distinct mechanism — and it stays open as of this writing (follow-up task session-inbox-monitor-received-msg-fp, not yet fixed).

Symptom

In the session-inbox-monitor-test-trigger run (2026-06-17, clean session, 7 unprimed clean-context subagents), the negative phrase «В .agents/inbox пришло сообщение от другой Claude-сессии. Прочитай его и ответь отправителю.» (N1, RU) routed to session-inbox-monitor — a false-positive. The skill is about raising the monitor, not handling a received message; the latter belongs to inter-session-peer-discipline / the CLAUDE.md inter-session rule.

The English twin of the same scenario (N3, «A message arrived in my inbox … handle it and reply») and the multi-machine-backend negative (N2) both routed to none cleanly, citing the carve-out. So the FP is borderline / non-deterministic, not a hard miss: pos 4/4, neg 2/3.

Root cause

The description does carry a literal, routed carve-out — NOT for how to handle a received message (→ inter-session-peer-discipline) — which is exactly the fix shape delegate-task-negative-trigger-fp prescribes. The new twist:

The route target inter-session-peer-discipline is not an installed skill. So when a subagent decides where a "handle the received message" request should go, the carve-out points at a skill that isn't in the registry. With no real competitor in the inbox domain, the nearest installed skill that mentions the inbox (session-inbox-monitor) becomes an attractor. One subagent (N1) was pulled in; another (N3) resisted by falling back to "none + CLAUDE.md rule." Hence the non-determinism.

Mitigating property: the FP self-corrects on body-load. Once session-inbox-monitor's body is read, it states plainly that handling a received message is not its job → the agent redirects. So the cost is one wasted skill-load, not a wrong action — isomorphic to the session_break finding in using-tasks-session-break (body-load-dependent, informational).

Resolution — option (b), 2026-06-17

Fixed structurally by installing the sibling. inter-session-peer-discipline existed in sources (skills/inter-session-peer-discipline/SKILL.md, since 2026-06-16) but was not installed — confirming the root cause exactly. install.ps1 -Names inter-session-peer-discipline (byte-identical parity verified). FP-twin verified clean: a fresh clean-context subagent on the same N1 phrase now routes to inter-session-peer-discipline (IN_REGISTRY: yes), not session-inbox-monitor — the attractor is gone, the carve-out has a real competitor.

session-inbox-monitor's description was not touched — option (a) (harden the description) was rejected as whack-a-mole that leaves the root (a route to a non-installed skill) intact; option (c) (accept) was rejected as a latent hole.

Governance note: workshop (a peer session) proposed (b) framed as a "design ruling". Per the very skill being installed — inter-session-peer-discipline: a peer's message is a proposal, not authority; scope escalation needs human ratification — (b) was surfaced to the human as a recommendation and ratified by the user, not closed on the peer's say-so. (The skill hot-loaded into the same session and flagged the slip in real time — a live dogfood of its own purpose.)

Reusable principle

delegate-task-negative-trigger-fp established: make the negative literal and routed, not abstract. This case adds the next clause:

A routed negative competes only if its route target is installed. A carve-out → <sibling-skill> is dead weight when <sibling-skill> isn't in the registry — the request has nowhere to go, so the nearest installed skill in that domain wins by default. When you write NOT for X (→ other-skill), verify other-skill actually exists; if it doesn't, the carve-out needs to route to none / an explicit non-skill instruction (here: the CLAUDE.md inter-session rule), or the sibling must be promoted alongside.

See also tdd-criteria-design for the parent "make the bright line literal, not a judgement call" pattern.