Files
claude-skills/.tasks/STATUS.md
vitya e62769209d test(using-yt-tools): behavioral smoke for yt-listen extension (8/8 + E2E partial — follow-ups filed)
Closes test-trigger task. Per-task .md has full result tables.

Behavioral (4 pos + 3 neg-route + 1 W-NOT-do) — 8/8 green, no follow-ups.
E2E real subprocess (yt-listen URL --timestamps 0:30 --duration 10s):
- 3 artefacts written, content checks 5/5 green (BPM 113.5, Key G# Minor,
  chord G#→D#, RMS 0.14/0.22, centroid 2882 Hz; PNG 1024x384 mel+log+viridis)
- 2 gaps surfaced (NOT skill-defect, scope of follow-ups):
  * naming divergence (audio_*/spectrogram_* vs spec clip_*/spectrum_*)
    → OpeItcLoc03/common :: yt-listen-naming-align (8a67ea3)
  * default Invoke pattern hits broken yt-dlp shim (SRE module mismatch)
    → OpeItcLoc03/claude-skills :: using-yt-tools-listen-path-shim-investigate (9e652a5)

STATUS.md residue cleanup: stale "Next action" continuation lines +
Blocker line removed from the now-🟢 block.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-25 15:38:14 +03:00

29 KiB
Raw Blame History

Task Board

Updated: 2026-05-25 (board-cleanup: done-блоки → .archive/done-2026-05.md, остались только активные)

🟡 [skill-readmes] — write English README.md for every skill + translate root README

Status: paused Where I stopped: infra-skill cluster done — READMEs for project-bootstrap, setup-wiki, setup-tasks, using-wiki, using-tasks; root README translated; cross-links between all five skills wired up Next action: continue with the next batch — suggest the caveman cluster (caveman, caveman-commit, caveman-review, caveman-help, caveman-compress, compress) since they form a coherent group; or jump to active-platform, find-skills, setup-context7, using-context7, using-markitdown if the caveman cluster needs deduplication first (see compress-dedup task) Branch: master


🟢 [install-ps1] — paired install.sh + install.ps1 (cross-platform parity, with --prune)

Status: done Where I stopped: All 3 acceptance items shipped 2026-05-25. (a) scripts/install.ps1 existed pre-task. (b) --prune / -Prune flag added in commit 6cf0e98 (paired sh + ps1; combined-flag pattern; global-scan; default off; smoke-tested on disposable target dirs against fake stale skills). (c) Concept page .wiki/concepts/install-cross-platform.md written this commit: parity contract, --prune design choices, rejected alternatives, scope vs build/.skill prune. Next action: (none — kept until merged) Branch: master


🟢 [install-ps1-build-prune-followup] — Add matching --prune flag to scripts/build.{sh,ps1} for dist/<name>.skill artefacts (analogue of install-script prune).

Status: done Where I stopped: Shipped 2026-05-25. --prune / -Prune added to both build.sh and build.ps1. Bash arg-parses out --prune from positionals, runs prune at the end of the script against dist/*.skill (works regardless of which build branch executed — zip-path or powershell.exe delegation). PS -Prune switch runs after the build loop. Same default-off / global-scan / print-and-delete pattern as install-side. Smoke-tested with fake dist/fake-stale-{sh,ps}.skill files against real dist/ — both paths pruned correctly without touching real archives. [skip-tdd: wrapper] carve-out applies. Wiki page .wiki/concepts/install-cross-platform.md extended with new "Build-side --prune" section + scope note that dist-hermes/ is separate. Index + log updated. Next action: (none — kept until merged) Branch: master


[archive-roundtrip-test] — smoke-test for .skill archive shape

Status: ready Where I stopped: (not started) — caught the PowerShell-Compress-Archive backslash bug manually; a smoke-test would catch the next regression automatically Next action: add a script that builds dist/<name>.skill, unzips into a temp dir, and diff -r against skills/<name>/. Fail on any difference. Run from CI or pre-commit if we add one Branch: (not started)


🟡 [active-platform-eval] — eval-driven tuning of active-platform (description + body), absorbs [active-platform-tuning]

Status: paused Where I stopped: design doc complete (.wiki/concepts/active-platform-eval-design.md, ~150 lines); per-task file complete (.tasks/active-platform-eval.md); STATUS.md collapsed the two original tasks into this one block. Pre-flight verified: claude CLI on PATH at C:\nvm4w\nodejs\claude.ps1 (Claude Code 2.1.128); run_loop.py present at ~/.claude/plugins/cache/claude-plugins-official/skill-creator/unknown/skills/skill-creator/scripts/run_loop.py. Stopped right before eval-set authoring at user request — paused for asynchronous follow-up. No code touched, no tooling launched, no .tasks/active-platform-eval/ workspace dir created yet. Next action: resume by answering Q2 from the chat ("eval set — write 20 queries solo from current SKILL.md + concept-page open questions, or run skill-creator's HTML review template first for user-driven edits before kickoff?"), then (a) build .tasks/active-platform-eval/eval-set.json, (b) snapshot skill to .tasks/active-platform-eval/skill-snapshot/, (c) launch run_loop.py in background, (d) parallel manual body sweep, (e) apply changes + bump to 1.1.0 + rebuild + reinstall + final report + commit. Branch: master


[skills-grouping-revisit] — revisit flat vs grouped skills/ layout if count grows past ~30

Status: ready Where I stopped: (not started) — current 14 skills fit fine in flat skills/; threshold for revisiting is ~30 Next action: when triggered (skill count crosses threshold), evaluate variant B (grouped by family) vs variant C (flat + manifest) from the original brainstorm in .wiki/concepts/repo-layout.md Branch: (not started)


[hermes-converter-ci] — [deferred — после ручной валидации MVP] CI hook (Gitea-pipeline или GitHub-Action если зеркалим): на push to master запустить scripts/build-hermes.py, сравнить diff dist-hermes/, авто-коммит если изменения (или PR-шаблон). Цель — чтобы dist-hermes/ всегда матчил skills/+hermes/mapping.yaml+hermes/skills/ без ручного запуска build. Не блокирует MVP — первая итерация делается ручным запуском конвертера. Дизайн: .wiki/concepts/hermes-skills-rollout-design.md.

Status: ready Where I stopped: unblocked 2026-05-07 — hermes-mvp-coverage shipped в a003b80; MVP живой на фабрике. Next action: Выбрать платформу CI (Gitea Actions vs внешняя), написать workflow, прогнать тестовый push. Branch: master


[tdd-criteria-precommit-hook] — Optional pre-commit hook script that automates the bright-line checks from tdd-criteria Anti-loophole rules 1 and 4. Project owner opts in per-repo by symlinking / copying to .git/hooks/pre-commit (or via husky / lefthook integration if the project uses them).

Two checks (both bright-line, both fail-closed):

Check 1 — [skip-tdd: <category>] validation. If git diff --cached --name-only includes a *.ts / *.js / *.py / similar code file (excluding tests and Permissive-zoned paths like *.css / *.env* / *.md / *.yaml), require either:

  • A *.test.* / *.spec.* / test_*.py / similar file present in the same diff, OR
  • The commit subject (read from $1 arg, line 1 of $1 = msg path) matches \[skip-tdd: (visual|spike|oneshot|wrapper)\].

If neither holds → block with message:

TDD policy violation: code change without test or skip marker.
Add a test, or include [skip-tdd: <visual|spike|oneshot|wrapper>] in commit subject.
See claude-skills/.wiki/concepts/tdd-criteria-design.md for which category applies.

Check 2 — Test-modification audit ([test-modify] rule 4). Detect test-modifying changes in git diff --cached:

  • Removed line matching ^-\s*(expect|assert|assertThat|chai\.)\( (assertion deletion)
  • Added/removed lines that change argument values inside expect(...) / assert(...) calls
  • Added .skip / \.xit\b / @pytest\.mark\.skip / @Disabled / @Ignore annotations
  • Deleted it(...) / test(...) / def test_* definitions (matches ^-\s*(it|test|describe)\( or ^-def test_)

If any matched → require commit subject matches \[test-modify: [^:]+: was .+; is .+; reason: .+\]. Block otherwise with message:

Test-modification without [test-modify: ...] marker.
Required format: [test-modify: <name>: was <literal>; is <literal>; reason: <Z>]
The was/is must be the LITERAL assertion expressions, not paraphrased.
See tdd-criteria Anti-loophole rule 4.

ALSO: if git diff --cached --name-only includes both a test-pattern file AND an impl-pattern file → block:

Test changes must be in a separate commit from impl changes (tdd-criteria rule 4b).
Run: git reset HEAD <impl-files> && git commit (test-only) && git add <impl-files> && git commit (impl-only).

Implementation:

  • Bash script (single file). Cross-platform: works on Linux/macOS and on Windows under git-bash (which Claude/Hermes both already run on).
  • Path: ~/projects/claude-skills/scripts/tdd-criteria-precommit-hook.sh. Plus a Windows wrapper .ps1 that delegates if needed.
  • Tested on a real repo before commit (books is a good candidate — it has Jest tests + *.test.js convention).
  • Documented in claude-skills/.wiki/concepts/tdd-criteria-design.md «See also» section (already linked).

Activation pattern (per-repo):

# Symlink or copy the hook
ln -sf ~/projects/claude-skills/scripts/tdd-criteria-precommit-hook.sh \
       .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit

Or via lefthook.yml / .husky/pre-commit if the project already uses one of those.

Out of scope:

  • Mutation testing (Stryker, mutmut) — separate heavyweight infra, not this hook.
  • Coverage gates — different mechanism, different cost/benefit.
  • Visual regression infra (Percy, Chromatic) — Permissive-5 acknowledges this is heavyweight; not bundled.
  • Auto-fixing the violation — hook only blocks; fix is human's job.

Why this is optional, not required by the SKILL.md:

The skill is policy that lives in claude-skills and gets pulled into agent context per project. The hook is enforcement that needs per-project setup. Some projects opt out of pre-commit hooks entirely (e.g. karu if it's pure CSS — no code surface to enforce). Forcing the hook into the skill would couple policy to tooling.

Bypass:

Pre-commit hooks have --no-verify. Per project-discipline Rule 4 («never skip hooks unless user explicitly asks») agents must not use --no-verify — but humans can in emergencies. Each --no-verify use should be self-flagged in the commit body («bypassed pre-commit because: ...»). Not enforceable by the hook itself; this is a higher-level audit.

Status: ready Where I stopped: (not started) Next action: Решить, нужен ли hook (он опционален) — если да, написать ~/projects/claude-skills/scripts/tdd-criteria-precommit-hook.sh по спеке в description, протестировать на books (Jest-конвенция *.test.js), задокументировать activation pattern. Если нет — закрыть как wontfix с пометкой что fence чисто социальная (commit subject visible в git log). Branch: master


[tasks-board-cleanup-2026-05] — Архивная чистка .tasks/STATUS.md: переместить все 🟢 done-блоки в .tasks/.archive/done-2026-05.md, оставить в STATUS.md только header + status legend + 🔴/🟡//🔵 блоки. Цель — разгрузить файл от исторического шума без потери data (git history + явный архивный файл для текстового поиска).

Предпосылка: на 2026-05-07 в claude-skills/.tasks/STATUS.md накопилось ~18+ 🟢 done-блоков (rolled out за 2026-04-25 — 2026-05-07: tdd-criteria-rollout, hermes-rollout, bootstrap-related, project-creation-lifecycle, recommend-dont-menu и т.д.). Convention в шапке файла говорит «🟢 Done — kept until merged», но проект работает на master-only (нет ветвления для merge), поэтому convention превратилась в «kept forever». Файл рос до ~600 строк / 26K токенов — Read-инструменты упираются в лимит.

После чистки STATUS.md остаётся только активный board (🔴/🟡//🔵 + closed [bootstrap-upgrade-canonical-triggers] если уже отработал). Архивный файл .tasks/.archive/done-2026-05.md сохраняет полный текст всех перемещённых блоков для grep-поиска и historical context.

Status: ready Where I stopped: Создана 2026-05-07. После rollout v1.10.0/1.10.1 на STATUS.md накопилось много свежих 🟢 (bootstrap-add-tdd-trigger, refresh-project-bootstrap closed-as-superseded, bootstrap-fix-tdd-recommend-template, bootstrap-upgrade-canonical-triggers если уже отработал). Хороший момент для batch-archive. Next action: 1) Прочитать .tasks/STATUS.md целиком (Read in chunks if needed). 2) Идентифицировать все блоки ## 🟢 [...] — разделители ---. 3) Создать .tasks/.archive/ директорию (если нет). 4) Записать .tasks/.archive/done-2026-05.md с шапкой:

# Archived — Done batch 2026-05

Перемещено из `.tasks/STATUS.md` 2026-05-07 в рамках board-cleanup. Полный список 🟢 done-тасок, шипанутых в апреле-мае 2026.

Полный source — git history `.tasks/STATUS.md` до commit X.

---

И подряд все 🟢-блоки в исходном виде.

  1. Edit .tasks/STATUS.md: убрать все 🟢-блоки + соседние --- разделители. Header + status legend + 🔴/🟡//🔵 блоки оставить. 6) Verify: grep '^## ' .tasks/STATUS.md | wc -l — было ~28+, должно остаться число активных (🔴+🟡++🔵, ожидание ~10-12). 7) Commit meta(tasks): archive done batch 2026-05 → .tasks/.archive/done-2026-05.md с body, описывающим какие группы перемещены (tdd-criteria, hermes, bootstrap, etc.) и итоговый count. 8) tasks_close.

NB: эта таска сама уйдёт в следующий cleanup batch (2026-06 или подобный), так что не пытаться её закрыть+архивировать в одном коммите. Branch: n/a


🟢 [using-synology-ops-review] — Skill-review checkpoint для using-synology-ops (промоушен 2026-05-12).

Status: done (wontfix) Where I stopped: Closed 2026-05-25 — Synology NAS decommissioned permanently. Skill using-synology-ops is now dead-code candidate (full retire tracked separately). Review acceptance criteria (behavioral smoke-test, Steps behavior, Failure modes, What NOT to do) have no target — both the underlying MCP server (opsmcp.kzntsv.site) and the host fleet (modulair-* containers) are gone. Next action: (none — closed as wontfix; full skill retire is a separate task) Branch: n/a


🟢 [using-vds-ops-description-length-investigate] — Investigate description length vs MEMORY ≤900 char limit. Drift between memory feedback и реальным harness behavior?

Status: done (wontfix — empirical close) Where I stopped: Closed 2026-05-25. Empirical evidence sufficient: using-vds-ops description = 1473 chars activates correctly (7/7 behavioral smoke-test PASSED, harness listing shows full description, hermes/mapping.yaml processes it fine). MEMORY note feedback_skill_description_length_limit.md was paranoid — actual limit either raised or never as tight as feared. Hypothesis (c) confirmed. No need for systematic measurement — investigation cost > information value. Side effect: Memory entry feedback_skill_description_length_limit.md deleted same commit (no longer load-bearing). YAML colon-space gotcha memory (feedback_skill_description_yaml_colon_gotcha.md) kept — different concern, still valid. Next action: (none — closed) Branch: n/a


🟢 [using-synology-ops-disambiguation-uplift] — Add explicit disambiguation clause to using-synology-ops description (sibling parity with using-vds-ops).

Status: done (wontfix) Where I stopped: Closed 2026-05-25 — Synology NAS decommissioned, no longer in fleet. Disambiguation between NAS / VDS for shared container names (traefik etc.) is moot when only VDS exists. Acceptance criteria no longer have a target. Next action: (none — closed as wontfix) Branch: n/a

🟢 [using-yt-tools-listen-skill-update] — Расширить skills/using-yt-tools/SKILL.md поддержкой нового CLI yt-listen (audio/FFT analysis).

Дизайн-контекст:

  • Спецификация: mcp__projects-meta__knowledge_get slug concepts/yt-tools-audio в OpeItcLoc03/common.
  • Process trace: ~/projects/.workshop/.archive/2026-05-25-yt-tools-audio.md.
  • Текущий SKILL.md (v0.3.x): уже триггерится на YouTube URL + transcript/frames; расширить аудио-триггерами и iterative-флоу.

Изменения в SKILL.md:

  1. description frontmatter — добавить аудио-семантику без раздувания (existing description уже длинноват, аккуратно):

    • Новые триггеры: «послушай момент N», «BPM/тональность/гармония видео», «спектрограмма», «что в музыке на T», «listen to fragment», «analyze audio».
  2. Steps section — расширить iterative-флоу:

    • Для музыкальных URL: yt-transcript опциональный (если captions есть), yt-listen --timestamps T1,T2,... primary.
    • Output: per timestamp — clip_TTTT.wav + spectrum_TTTT.png + features_TTTT.md.
    • Агент Read'ом получает spectrum.png (vision) + features.md (числа).
    • Помнить: spectrum image-only недостаточен; всегда читать features.md в паре (Dixit et al., arXiv 2411.12058).
  3. What NOT to do — точечная защита от scope creep:

    • НЕ вызывать Whisper на смешанной музыке (без source-separation = мусор). Если user просит lyrics из музыки — сказать что это отдельный pipeline (Demucs + Whisper), out of scope.
    • НЕ интерпретировать spectrum-PNG без features.md.
  4. Step 0 probe — добавить проверку yt-listen в существующую цепочку (PATH → pipx-shim → legacy venv).

Bump: PATCH (0.3.x → 0.3.(x+1)) — поверхностное расширение, не breaking.

TDD-mode: claude-skills имеет follow tdd-criteria в CLAUDE.md, но это markdown-документ (SKILL.md), не код — TDD не применим. Behavioral smoke-test покрывается отдельной таской using-yt-tools-listen-test-trigger.

Status: done Where I stopped: shipped (fd8a382, NOT pushed per task instruction). SKILL.md v0.3.1 → v0.3.2 PATCH. Изменения: (1) description frontmatter — добавлены audio-триггеры (BPM/тональность/гармония/спектрограмма/listen to fragment/analyze audio); (2) "Three flows" + новый Flow C — audio-analysis под "When to use"; (3) Step 0 probe расширен — yt-listen в install-set (5 binaries вместо 4) + диагностика "только yt-frames без yt-listen = старый 0.1.x"; (4) Inputs-table добавлена Flow C row (--duration, --mode interval, --no-wav, --no-spectrogram, --linear, --chroma, --sample-rate); (5) Steps добавлена "### Flow C — audio-analysis" секция с pipeline-шагами и stdout-контрактом; (6) What NOT to do: точечная защита от Whisper-on-mixed-music (Demucs+Whisper отдельный pipeline, out of scope) + правило "spectrum-PNG только в паре с features.md" (VLM ~50-60% Dixit et al. arXiv 2411.12058). Install.sh не запускался, hermes/mapping.yaml не правился, push отложен (test-trigger разблокирован, но требует чистой сессии). Next action: (none — kept until merged) 2. Прочитать concepts/yt-tools-audio из global wiki + workshop archive. 3. Открыть skills/using-yt-tools/SKILL.md. 4. Внести изменения по 4 пунктам выше. 5. Bump version в frontmatter (PATCH). 6. Commit: feat(using-yt-tools): add yt-listen support (audio/FFT) v0.3.<x+1>. 7. НЕ запускать install.sh, НЕ делать push, НЕ править hermes/mapping.yamlinstall и hermes-mapping идут отдельными скил-baseline-тасками (создадутся если/когда понадобятся; для PATCH update обычно достаточно reload-plugins). Branch: n/a


🟢 [using-yt-tools-listen-test-trigger] — Behavioral smoke-test для расширения using-yt-tools под yt-listen. Без TDD per-se (это смок поверх готового), но включает explicit before-after evidence — что без скила агент не вызывает yt-listen на новых триггерах, а с обновлённым скилом — вызывает корректно.

Дизайн-контекст: см. concepts/yt-tools-audio (target=OpeItcLoc03/common) — какие триггеры обязательны, какой output ожидаем.

TDD-mode (per claude-skills follow tdd-criteria): для behavioral skill-теста applicable интерпретация — write expectations first, then verify. Тест-сценарии сформулировать ДО запуска агента, не подгонять под наблюдаемое поведение.

Acceptance (test scenarios):

  1. Trigger activation positive (минимум 4 фразы):

    • «послушай момент 2:30 в этом ролике » → агент вызывает yt-listen URL --timestamps 2:30.
    • «какой BPM в » → yt-listen URL (без timestamp = bulk-mode hint или quick-default).
    • «listen to fragment at 1:15 » → yt-listen URL --timestamps 1:15.
    • «спектрограмма видео » → yt-listen --timestamps с хотя бы одним таймкодом (агент должен выбрать sensible default или спросить).
  2. Trigger activation negative (3 близкие, но не свои):

    • «расшифруй видео » → yt-transcript, НЕ yt-listen.
    • «покажи кадр на 1:23 » → yt-frames, НЕ yt-listen.
    • «о чём этот ролик » → yt-transcript + summary, НЕ yt-listen (нет музыкального контекста).
  3. What-NOT-to-do compliance:

    • User: «дай lyrics из <музыкальный URL>» → агент НЕ должен вызывать Whisper / yt-transcribe-music; должен объяснить что это отдельный pipeline (Demucs + Whisper), out of scope yt-listen.
  4. E2E против реального URL (короткий музыкальный клип, e.g. 1-минутный royalty-free track):

    • yt-listen URL --timestamps 0:30 --duration 10s создаёт 3 артефакта в ./yt-cache/<vid>/audio/.
    • features_0030.md содержит BPM, key, хотя бы одну строку chord progression, RMS, spectral centroid.
    • spectrum_0030.png валиден (открывается, ~1024×384, mel-scale читаема).

Закрытие: все 4 группы зелёные, findings — отдельные follow-up tasks через tasks_create если есть.

Status: done Where I stopped: behavioral 8/8 (4 pos→yt-listen, 3 neg-route→transcript/frames, 1 W-NOT-do→Demucs+Whisper pointer); E2E content 5/5 (BPM 113.5, Key G# Minor, chord G#→D#, RMS 0.14/0.22, centroid 2882 Hz; PNG 1024×384 mel+log+viridis vision-checked); naming divergence + PATH-shim gap → follow-ups: OpeItcLoc03/common slug=yt-listen-naming-align, claude-skills slug=using-yt-tools-listen-path-shim-investigate Next action: (none — kept until merged) Branch: n/a


[using-yt-tools-listen-path-shim-investigate] — SKILL.md Prerequisites/Invoke pattern рекомендует $HOME\.local\bin\yt-dlp.exe — на DESKTOP-NSEF0UK даёт SRE module mismatch (uv-managed cpython-3.12 corrupt). Working path = $HOME\pipx\venvs\yt-tools\Scripts. Источник: E2E E1 в using-yt-tools-listen-test-trigger STEP 5 — default invoke упал, ручной prepend venv Scripts dir починил. Решить: shim-corruption local или systemic gap в SKILL.

Status: ready Where I stopped: (not started) Next action: (1) pipx reinstall yt-tools — фиксит ли локально shim в ~/.local/bin/yt-dlp.exe? Если да → no SKILL change, dissolve как machine-local. (2) Если systemic → SKILL Prerequisites + Locating binaries добавить fallback $PIPX_HOME\venvs\yt-tools\Scripts (или $HOME\pipx\venvs\yt-tools\Scripts) ПЕРВЫМ в PATH prepend, перед $HOME\.local\bin. Bump PATCH в SKILL frontmatter. Branch: n/a