review(using-system-snapshot): VERDICT PASS 3/3 — close skill-using-system-snapshot-review

Non-implementer review of skills/using-system-snapshot (v0.1.0). All three
acceptance criteria pass:
- trigger phrases cover real scenarios (4/4 positives + clean negatives)
- no-claim-without-snapshot rule explicit (4 places)
- output format brief (three lines, verified vs live payload)

Evidence: live meta_system_snapshot call confirms the documented poller/docker/
tasks contract; 9 fresh-context subagents over a simulated registry (real
descriptions + using-vds-ops/using-projects-meta/using-tasks competitors) routed
8 cleanly, incl. no false-positive on a docker-compose.yml edit.

3 informational findings, none blocking:
1. cross-project task-count phrasings overlap with using-projects-meta (by-design)
2. local-container deep diagnosis unowned — vds-ops scope, not this skill
3. deployment scaffold missing — not installed, not in hermes/mapping.yaml,
   no -install/-hermes-mapping/-test-trigger baseline tasks

No SKILL.md edits -> no version bump. TDD N/A (review of markdown policy).
Review outcome recorded in .wiki/concepts/using-system-snapshot-design.md + log.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
2026-06-09 19:34:21 +03:00
parent 17d7ff8264
commit eba4aeb23a
3 changed files with 49 additions and 14 deletions

View File

@@ -1,5 +1,5 @@
# Task Board
_Updated: 2026-06-09 — using-tasks-status-read-perf done (using-tasks 1.2.0→1.3.0; done-task archival rule fixes STATUS.md bloat; literal `tasks_get_status`-for-orientation swap NOT done — tool can't enumerate the board, documented). Ранее: session-break-using-tasks done (using-tasks 1.1.0→1.2.0; `session_break` marker — stop after close before claim-next). Ранее: delegate-task-review done (VERDICT PASS; smoke-test 6/6 fresh subagents, no blocking findings, 3 informational notes). Ранее: delegate-task-test-trigger done (pos 5/5, neg 2/3; FP на «создать задачу себе» → follow-up delegate-task-description-fp-fix, shipped v0.2.1). Ранее (2026-06-08): откат churn'а от агент-раннера (always-on dry-run): 5 тасок (using-yt-tools-rate-limit-guard, archive-roundtrip-test, skills-grouping-revisit, hermes-converter-ci, tdd-criteria-precommit-hook) спуриозно claimed/blocked из-за workspace-divergence бага раннера → возвращены в ⚪ ready. yt-tools re-scoped на plugin-репо `OpeItcLoc03/yt-tools` (исходный stub deprecated)._
_Updated: 2026-06-09 — skill-using-system-snapshot-review done (VERDICT PASS 3/3; live tool-contract verify + 9-way fresh-context trigger test, 8 clean; 3 informational findings, none blocking — incl. missing deployment scaffold: skill not installed / not in hermes-mapping / no baseline tasks). Ранее: using-tasks-status-read-perf done (using-tasks 1.2.0→1.3.0; done-task archival rule fixes STATUS.md bloat; literal `tasks_get_status`-for-orientation swap NOT done — tool can't enumerate the board, documented). Ранее: session-break-using-tasks done (using-tasks 1.1.0→1.2.0; `session_break` marker — stop after close before claim-next). Ранее: delegate-task-review done (VERDICT PASS; smoke-test 6/6 fresh subagents, no blocking findings, 3 informational notes). Ранее: delegate-task-test-trigger done (pos 5/5, neg 2/3; FP на «создать задачу себе» → follow-up delegate-task-description-fp-fix, shipped v0.2.1). Ранее (2026-06-08): откат churn'а от агент-раннера (always-on dry-run): 5 тасок (using-yt-tools-rate-limit-guard, archive-roundtrip-test, skills-grouping-revisit, hermes-converter-ci, tdd-criteria-precommit-hook) спуриозно claimed/blocked из-за workspace-divergence бага раннера → возвращены в ⚪ ready. yt-tools re-scoped на plugin-репо `OpeItcLoc03/yt-tools` (исходный stub deprecated)._
<!--
Canonical layout. One block per task. Per-task deep context lives in
@@ -727,25 +727,20 @@ needs-claude
---
## 🔵 [skill-using-system-snapshot-review] — Review скила `using-system-snapshot`.
## 🟢 [skill-using-system-snapshot-review] — Review скила `using-system-snapshot`.
Проверить:
- [ ] Trigger-фразы покрывают реальные сценарии
- [ ] Запрет на утверждение состояния без вызова snapshot — явно прописан
- [ ] Формат вывода достаточно краткий (одна строка на секцию)
- [x] Trigger-фразы покрывают реальные сценарии — 4/4 positives → using-system-snapshot; negatives correctly routed (VDS-logs→using-vds-ops, mutate/full-board→using-projects-meta, docker-compose edit→none)
- [x] Запрет на утверждение состояния без вызова snapshot — явно прописан (4 места: Overview core rule, When to use, What NOT to do, Common mistakes)
- [x] Формат вывода достаточно краткий (одна строка на секцию) — three-line block, verified achievable against live payload
**Weight:** needs-claude
**Status:** blocked
**Where I stopped:** delivery push failed (diverged) — needs manual reconcile: u have unmerged files.
hint: Fix them up in the work tree, and then use 'git add/rm <file>'
hint: as appropriate to mark resolution and make a commit.
fatal: Exiting because of an unresolved conflict.
**Next action:** Прочитать SKILL.md, протестировать триггер в реальной сессии
**Status:** done
**Where I stopped:** VERDICT **PASS** on all 3 acceptance criteria (non-implementer review). Tool contract verified by a live `meta_system_snapshot` call — output matches the documented `poller {running,projects}` / `docker [{name,status}]` (local, incl. `agents-task-runner-*`, `Up … (healthy)` strings) / `tasks {owner/repo:{active,blocked}}` shape exactly. Behavioral trigger smoke = 9 fresh-context subagents over a simulated skill registry (real descriptions + using-vds-ops/using-projects-meta/using-tasks competitors, no expected-answer hint): **8 clean** (4/4 positives→snapshot; VDS-logs→vds-ops; mutate/full-board→projects-meta; docker-compose.yml edit→none, no FP on "docker"). **3 informational findings, none blocking:** (1) cross-project task-COUNT phrasings («сколько активных задач по всем проектам») overlap with using-projects-meta — by-design (skill defers precise per-task work; tasks-line is a bonus of the combined view), no fix; (2) LOCAL-container deep diagnosis is unowned — vds-ops incident triggers grab a local container its VDS-only tools can't reach (vds-ops scoping concern, not this skill's defect); (3) **deployment scaffold missing** — skill committed (v0.1.0) but NOT installed to `~/.claude/skills/` (absent from this session's available-skills), NOT in `hermes/mapping.yaml`, and has no `-install`/`-hermes-mapping`/`-test-trigger` baseline tasks (unlike meta-host-routing/delegate-task). Recommended follow-ups before it reaches live sessions; hermes mode could be `auto` (read-only skill) — owner's call. Review outcome appended to `.wiki/concepts/using-system-snapshot-design.md` + log line. TDD N/A (review of markdown policy artifact). No SKILL.md edits → no version bump.
**Next action:** (none — kept until merged). Recommended follow-ups: file `using-system-snapshot-{install,hermes-mapping,test-trigger}` baseline tasks if the owner wants the skill live.
**Branch:** n/a
**Owner:** DESKTOP-NSEF0UK:claude-opus:18704
**Claim token:** fc9ccfea-912f-4354-bdf7-d8dbff1311de
**Claim expires at:** 2026-06-09T16:43:05.685Z
<!-- created-by: OpeItcLoc03@DESKTOP-NSEF0UK / from: OpeItcLoc03/workshop / 2026-06-09T11:01:59.476Z -->
<!-- closed-by: DESKTOP-NSEF0UK:claude-opus:18704 / 2026-06-09 / VERDICT PASS 3/3 acceptance; live tool-contract verify + 9-way fresh-context trigger test (8 clean); 3 informational findings (task-count overlap by-design, local-container gap = vds-ops scope, deployment scaffold missing); no skill edits / no bump -->
---

View File

@@ -57,3 +57,42 @@ is absent, the server isn't registered → `setup-projects-meta`.
Markdown policy artifact — no code/test surface (consistent with sibling skill
tasks). Behavioral trigger smoke-test is the paired `skill-using-system-snapshot-review`
task, not this implementation task.
## Review outcome (2026-06-09, `skill-using-system-snapshot-review`)
**Verdict: PASS** on all three acceptance criteria. Reviewer was a non-implementer
session.
- **Tool contract verified live** — a real `meta_system_snapshot` call returned
exactly the documented shape (`poller {running, projects}`, `docker [{name,
status}]` incl. `agents-task-runner-*` with `Up … (healthy)` strings, `tasks
{owner/repo: {active, blocked}}`). The "The call" table and this page are accurate.
- **Trigger phrases cover real scenarios** ✅ — 9 fresh-context subagents, each
given a simulated skill registry (real descriptions + `using-vds-ops` /
`using-projects-meta` / `using-tasks` competitors) and one trigger phrase, no
hint of the expected answer. 4/4 positives → `using-system-snapshot`; VDS-logs →
`using-vds-ops`; mutate/full-board → `using-projects-meta`; `docker-compose.yml`
edit → `none` (no false-positive on the "docker" keyword).
- **No-claim-without-snapshot rule explicit** ✅ — stated in 4 places (Overview
core rule, "When to use", "What NOT to do", Common-mistakes table).
- **Output format brief** ✅ — three-line block, per-line rules, "no raw JSON";
confirmed achievable against the live payload.
**Informational findings (none blocking):**
1. **Task-count overlap with `using-projects-meta`.** «сколько активных задач по
всем проектам» routed to `using-projects-meta`, not the snapshot. By-design —
the skill defers *precise* per-task work and the `tasks` line is a bonus of the
combined ops view, not its headline — so no fix. Quick «сводка по задачам …»
glances still route here correctly.
2. **Local-container deep diagnosis is unowned.** «локальный контейнер … почему
рестартует» routed to `using-vds-ops` (its incident-phrase triggers grabbed a
*local* container, which its VDS-only tools can't reach). Not this skill's
defect — the snapshot correctly does not claim deep "why". Candidate
`using-vds-ops` scoping follow-up if it recurs.
3. **Deployment scaffold missing.** The skill is committed (`skills/…`, v0.1.0)
but is **not** installed to `~/.claude/skills/`, **not** in
`hermes/mapping.yaml`, and has no `-install` / `-hermes-mapping` /
`-test-trigger` baseline tasks (unlike `meta-host-routing` / `delegate-task`).
Recommended follow-ups before it reaches live sessions; hermes mode could be
`auto` since the skill is read-only (owner's call).

View File

@@ -69,5 +69,6 @@ Parseable: `grep "^## \[" .wiki/log.md | tail -20`.
## [2026-06-09] decision | delegate-task-negative-trigger-fp — `delegate-task` 0.2.0→0.2.1 (PATCH): fixed 5/5-consistent false-positive on «создать задачу себе». Root cause: self-task phrase shares stem «создать задачу» with the «создать задачу на агента» positive trigger; the abstract "Does NOT apply when doing the work yourself" carve-out can't beat a literal stem-match under the 1%-rule. Fix: made the negative literal + routed («создать задачу себе» / «task for myself» / «поставить себе задачу» → using-tasks) in description + body disambiguator («на агента»/«агенту» = delegate; «себе» = own board). Re-verified via fresh-context subagent trigger run: positives 5/5 (no regression), negative 4/5 → using-tasks (was 0/5); the 1 residual miss was an eval-harness artifact (forced skill-name-before-reasoning), not description ambiguity. Concept page written; reusable principle = put the exact colliding negative phrase with an explicit →sibling route, literal beats abstract.
## [2026-06-09] decision | delegate-task-session-break — `delegate-task` 0.2.1→0.2.2 (PATCH): authoring side of the `session_break` marker (consumer = using-tasks v1.2.0). Added pre-flight Q6 (after notify): "Session-break после этой задачи? (domain-switch / milestone / heavy infra)"; if yes → set optional template field `session_break: true | "<hint>"` (trailer, next to weight/notify/allow_upgrade; same lowercase frontmatter key using-tasks reads). Usage guidance lists three set-it cases; What-NOT-to-do bullet warns against setting it routinely (it's a real-boundary marker, not a default). Wiki concept page concepts/delegate-task-session-break.md + index. Pairs with using-tasks-session-break.
## [2026-06-09] decision | using-system-snapshot — new skill v0.1.0: thin read-only wrapper over the single `mcp__projects-meta__meta_system_snapshot` call (poller status + local docker containers + cached cross-project task summary). Replaces the scatter of `tasklist` + `docker ps` + manual `meta_status`. Core rule: no claim about poller / local-docker / task-load state without calling the tool in the current turn (memory + stale earlier snapshot ≠ evidence). Output = three lines, one per section (docker lists only problem containers; tasks gives Σ active/blocked + busiest 23). Liveness split documented: poller+docker live, tasks from cache (defer precise work to using-projects-meta Step 0). Scope boundaries: deep single-container diagnosis → using-vds-ops / `docker logs`; docker section is LOCAL, not the VDS. Read-only, no per-session grant (mirrors using-vds-ops). Output shape verified by a live call 2026-06-09. Concept page concepts/using-system-snapshot-design.md + index. TDD N/A (markdown policy artifact); behavioral smoke-test = paired skill-using-system-snapshot-review task.
## [2026-06-09] review | using-system-snapshot v0.1.0 — VERDICT PASS on all 3 acceptance criteria (skill-using-system-snapshot-review). Tool contract verified by a live `meta_system_snapshot` call (output matches the documented `poller`/`docker`/`tasks` shape exactly). Behavioral trigger smoke = 9 fresh-context subagents over a simulated registry (real descriptions + using-vds-ops/using-projects-meta/using-tasks competitors, no expected-answer hint): 4/4 positives → using-system-snapshot; VDS-logs → using-vds-ops; mutate/full-board → using-projects-meta; `docker-compose.yml` edit → none (no FP on "docker" keyword). No-claim-without-snapshot rule explicit in 4 places; three-line output format confirmed achievable against the live payload. 3 informational findings (none blocking): (1) cross-project task-COUNT phrasings overlap with using-projects-meta — by-design, snapshot defers precise per-task work; (2) LOCAL-container deep diagnosis is unowned — vds-ops incident triggers grab local containers its VDS-only tools can't reach (vds-ops scoping, not this skill); (3) deployment scaffold missing — skill committed but not installed to `~/.claude/skills/`, not in `hermes/mapping.yaml`, no -install/-hermes-mapping/-test-trigger baseline tasks; recommended follow-ups (hermes mode could be `auto`, read-only skill). Review outcome appended to concepts/using-system-snapshot-design.md.
## [2026-06-09] decision | using-tasks-status-archival — `using-tasks` 1.2.0→1.3.0 (MINOR): added done-task archival rule to fix STATUS.md bloat ("huge STATUS.md" complaint). When ≥10 🟢 done blocks pile up — checked at session start (step 7) and after close (Task completion step 7) — move them verbatim to `.tasks/archive/YYYY-MM.md` (append, one file per month, one-time header), leaving only 🔴/🟡/⚪/🔵 on the board; committed on its own. Did NOT follow the task's literal instruction to replace `Read STATUS.md` with `tasks_get_status` for orientation: that tool returns a single task's live status by known slug (`{status, found}`) and cannot enumerate the board, and `tasks_aggregate` is cross-project + cache-based + doesn't index ready/done (its docs say read STATUS.md directly for the current project). So orientation stays a local board-read (kept cheap by archival); skill now warns against both tools for board enumeration and points `tasks_get_status` at its real single-task use. Core goal (kill the bloat) met by archival alone. Concept page concepts/using-tasks-status-archival.md + index. TDD N/A (markdown policy). Deviation flagged for paired review task using-tasks-status-read-perf-review.
## [2026-06-09] decision | using-tasks-session-break — `using-tasks` 1.1.0→1.2.0 (MINOR): added the `session_break` marker. Task author sets `session_break: true | "<hint>"` in task frontmatter (mirrored as `**Session break:**` on the local board); after the task closes 🟢, before `tasks_claim_next`, an autonomous agent prints the verbatim line `🔚 SESSION BOUNDARY — [slug] закрыта. Рекомендую завершить текущую сессию. Следующий трек: [value | "см. STATUS.md"]` and stops instead of chaining the next task. Absent → behaviour unchanged. Enforced in Task completion step 6 + Rules bullet + format docs. Marker not heuristic: the stop-point is an authoring choice, not a runner guess.