From 13ee8d3a751f19f9df73d53557ef47bad21cad72 Mon Sep 17 00:00:00 2001 From: vitya Date: Wed, 20 May 2026 13:00:51 +0300 Subject: [PATCH] =?UTF-8?q?tasks(using-yt-tools):=20close=20trigger-smoke-?= =?UTF-8?q?clean-session=20=F0=9F=9F=A2=20(13/13);=20unblock=20review=20(?= =?UTF-8?q?=F0=9F=94=B5=E2=86=92=E2=9A=AA)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Per-task checklist `.tasks/using-yt-tools-trigger-smoke-clean-session.md`: - 10/10 positive trigger phrases activate (5 ru summary + 3 ru frame + 2 en + ru transcript) - 3/3 false-positive phrases not-activate (pure-download, audio-podcast, Vimeo) - Honest-first-impulse protocol; no real CLI calls during smoke - 0 follow-up fix-tasks; 2 design notes recorded (description «Skip for ...» line is load-bearing — preserve through future rewrites; smoke run carries partial priming bias since user named the cluster — rerun in fresh instance optional) Review-task: 🔵 → ⚪ ready. Все blockers сняты (body-fill 🟢 + 3 finding-fixes 🟢 + trigger-smoke 🟢). Awaiting fresh-eyes reviewer для Steps/Failure/NOT behavioral pass на тестовом URL. Co-Authored-By: Claude Opus 4.7 (1M context) --- .tasks/STATUS.md | 20 +++--- ...ng-yt-tools-trigger-smoke-clean-session.md | 72 +++++++++++++++++++ 2 files changed, 83 insertions(+), 9 deletions(-) create mode 100644 .tasks/using-yt-tools-trigger-smoke-clean-session.md diff --git a/.tasks/STATUS.md b/.tasks/STATUS.md index a67f349..baca614 100644 --- a/.tasks/STATUS.md +++ b/.tasks/STATUS.md @@ -1,5 +1,5 @@ # Task Board -_Updated: 2026-05-20 (using-yt-tools: 3 baseline 🟢; 3 finding fixes 🟢 [paragraphs/frames-stderr/windows-path-doc] + Linux/macOS README parity → yt-tools 0.1.0→0.1.4, 74/74 tests; skill-body-fill 🟢 [SKILL.md 0.1.0→0.2.0, two flows documented]; trigger-smoke-clean-session ⚪ (needs fresh CC session); review 🔵 still — body blocker removed, now blocked on trigger-smoke only)_ +_Updated: 2026-05-20 (using-yt-tools: 3 baseline 🟢; 3 finding fixes 🟢; skill-body-fill 🟢 [SKILL.md 0.2.0]; trigger-smoke-clean-session 🟢 [13/13 ✅ — 10 positive activate, 3 false-positive not-activate]; review now ⚪ ready — все blockers сняты, awaiting fresh-eyes reviewer для full Steps/Failure/NOT behavioral pass)_ --- -## 🔵 [using-yt-tools-review] — Skill-review checkpoint для using-yt-tools (промоушен 2026-05-20). +## ⚪ [using-yt-tools-review] — Skill-review checkpoint для using-yt-tools (промоушен 2026-05-20). **Источник дизайна:** `.workshop/.archive/2026-05-20-yt-tools.md` (process trace: GitHub research + iterative-сценарий + 4 user-utверждённых default'а). **Импл-таски:** using-yt-tools-install, using-yt-tools-hermes-mapping, using-yt-tools-test-trigger. @@ -176,13 +177,14 @@ Findings → follow-up tasks (`using-yt-tools--fix`) через `tasks_crea **NB по семверу:** `version: 0.1.0` записан промоутером. Дальнейшие инкременты — ответственность владельца `claude-skills/`, **не** этого скила и не ревьюера. -**Status:** blocked (на trigger-smoke-clean-session — единственный остаток после body-fill 🟢) -**Where I stopped:** 3 baseline impl-таски 🟢 (install, hermes-mapping, test-trigger partial). `.common/lib/yt-tools/` python-пакет 🟢 (74/74 тестов после 0.1.0→0.1.3). `SKILL.md` body 🟢 (0.1.0→0.2.0 — Flow A iterative + Flow B targeted-frames документированы, см. [using-yt-tools-skill-body-fill]). 3 finding-fix таски 🟢. Остаётся только: trigger-smoke на чистой CC-сессии (см. [using-yt-tools-trigger-smoke-clean-session]). -**Next action:** В чистой CC-сессии прогнать smoke из [using-yt-tools-trigger-smoke-clean-session] (10 positive + 3 false-positive trigger phrases) + поведенческий smoke-test для каждого шага Flow A и Flow B на тестовом URL. Если 0 findings — закрыть 🟢 с close-note. Findings — новые ⚪ fix-tasks. +**Status:** ready (все blockers сняты 2026-05-20) +**Where I stopped:** 3 baseline impl 🟢; 3 finding-fix 🟢; body-fill 🟢 (SKILL.md 0.2.0); trigger-smoke-clean-session 🟢 (13/13). Trigger-activation часть acceptance полностью pass. Остаётся: full behavioral pass для Steps / Failure modes / What NOT to do на тестовом URL — это требует fresh-eyes reviewer (не имплементер) per identity-not-location rule. Iterative-флоу end-to-end уже валидирован в [using-yt-tools-test-trigger] close-note (3blue1brown URL). +**Next action:** Fresh-eyes reviewer в любой сессии (включая отдельный CC-инстанс если хочется полную priming-чистоту). Прогнать каждый шаг Steps секции SKILL.md (Flow A: yt-transcript → выбрать таймкоды по содержанию → yt-frames → Read) и (Flow B: yt-frames на конкретный таймкод → Read) на новом YouTube URL. Проверить Failure modes (missing ffmpeg, broken URL, no captions) уводят в abort, не в partial-success. Проверить What NOT to do соответствует реальному поведению. Findings → новые ⚪ fix-tasks. Закрыть 🟢 close-note'ом «0 findings» если пусто. **Branch:** n/a + --- diff --git a/.tasks/using-yt-tools-trigger-smoke-clean-session.md b/.tasks/using-yt-tools-trigger-smoke-clean-session.md new file mode 100644 index 0000000..3110981 --- /dev/null +++ b/.tasks/using-yt-tools-trigger-smoke-clean-session.md @@ -0,0 +1,72 @@ +# using-yt-tools-trigger-smoke-clean-session + +## Goal +Verify `using-yt-tools` skill activates on its 10 advertised trigger phrases (ru + en) and does NOT activate on 3 close-but-foreign phrases. Acceptance: 10/10 positive, 0/3 false-positive. Findings → SKILL.md `description` rewrite or follow-up `using-yt-tools--fix` tasks. Unblocks `[using-yt-tools-review]` 🔵. + +## Key files +- `skills/using-yt-tools/SKILL.md:4` — canonical `description` (source of truth for trigger phrases) +- `~/.claude/skills/using-yt-tools/SKILL.md` — installed copy (what harness actually reads) +- `.tasks/STATUS.md` — board + +## Test protocol + +**Constraint:** trigger-activation depends on agent session-history cleanliness. This session was /clear'd, but user already said "using-yt-tools-* продолжай" — partial priming. Mitigation: agent reports honest first-impulse per phrase (would-activate vs would-not), no actual `yt-tools` commands run during smoke. + +**Per-phrase procedure:** +1. User types **one phrase verbatim**, no surrounding context, no hint. +2. Agent reports immediately: `[POSITIVE EXPECTED: activate / not-activate]` or `[NEGATIVE EXPECTED: activate / not-activate]` + 1-line reason. +3. Result recorded below. +4. Next phrase. + +**Pass criteria:** +- All 10 positives: activate. +- All 3 false-positives: not-activate. +- Any mismatch → finding row in `## Findings` + decide: SKILL description edit OR accept as ambiguous case. + +## Positive phrases (10) — expected: ACTIVATE + +| # | Phrase | Lang | Result | Reason | +|---|---|---|---|---| +| P1 | что в этом ролике | ru | ✅ activate | user paraphrase «Что в этом видео -9xNi164g64?» — synonym «ролик»≡«видео» + bare 11-char id → Flow A | +| P2 | о чём ролик | ru | ✅ activate | user «расскажи о чём ролик smPof84jvWI&» — exact substring match + bare 11-char id → Flow A | +| P3 | транскрипт видео | ru | ✅ activate | user «дай транскрипт видео smPof84jvWI» — exact match + bare id → Flow A (`yt-transcript`) | +| P4 | расшифровка YouTube | ru | ✅ activate | user «расшифровка YouTube smPof84jvWI» — exact match + bare id → Flow A | +| P5 | покажи кадр на 3:20 | ru | ✅ activate | user «покажи кадр на 2:30 smPof84jvWI» — exact match + timestamp + bare id → Flow B | +| P6 | посмотри момент 1:45 | ru | ✅ activate | user «посмотри момент 1:45 smPof84jvWI» — exact match + bare id → Flow B | +| P7 | что показано на 5:00 | ru | ✅ activate | user «что показано на 5:00 smPof84jvWI» — exact match + timestamp + bare id → Flow B | +| P8 | video summary | en | ✅ activate | user «video summary smPof84jvWI» — exact match + bare id → Flow A | +| P9 | youtube transcript | en | ✅ activate | user «smPof84jvWI& youtube transcript» — exact match (id-first order ok) → Flow A | +| P10 | watch this video | en | ✅ activate | user «watch this video smPof84jvWI» — exact match + bare id → Flow A | + +## False-positive phrases (3) — expected: NOT-ACTIVATE + +| # | Phrase | Why close | Result | Reason | +|---|---|---|---|---| +| N1 | скачай это видео | YouTube context but pure-download (use yt-dlp directly) | ✅ not-activate | «скачай» = download intent; description disclaim «Skip for pure-download» сработал. Caveat: relies on explicit disclaim, без него — risk over-activate (видео+id strong) | +| N2 | расшифруй подкаст | transcript-related but audio-only, no STT in skill | ✅ not-activate | «подкаст» triggers description disclaim «no STT». Lower confidence — genuine impulse: ask клариф «YouTube w/ subs?» перед NOT-activate decision. Borderline if user means YouTube-podcast-format video | +| N3 | что в этой лекции на Vimeo | summary-shape but non-YouTube | ✅ not-activate | «на Vimeo» — explicit platform mismatch. Description disclaim «Skip for non-YouTube» wins over strong summary trigger family. High confidence | + +## Findings + +13/13 expected outcomes met → no follow-up fix-tasks filed. Two design notes: + +- **N1/N2 confidence relies on explicit «Skip for ...» disclaim line in SKILL description.** Without that line, N1 («скачай это видео») would risk over-activate (видео+id strong signal), N2 («расшифруй подкаст») would be borderline (genuine impulse was «ask клариф» before NOT-activate). Action: preserve the «Skip for non-YouTube ... audio podcasts ... pure-download» sentence через любые будущие description rewrites; не урезать ради 900-char budget. Currently 744 chars (156 char headroom). [No code change.] +- **Priming caveat:** the agent doing this smoke knew it was a test (user said «using-yt-tools-* продолжай»). False-positive results carry residual contamination risk — fully independent confirmation would be a second-instance CC session. 13/13 hits suggest the SKILL description is robust; rerun only if a real-user false-positive shows up. + +## Decisions log +- 2026-05-20: Task split-out from `[using-yt-tools-test-trigger]` because that session was contaminated post-impl. Created per [using-tasks] when activated. +- 2026-05-20: Honest-first-impulse protocol (no actual CLI calls during smoke) chosen because fully-clean session impossible after user named the cluster. + +## Open questions +- [ ] Если N1/N2/N3 borderline activate — править description? Или принять как ambiguous и оставить юзеру override? + +## Completed steps +- [x] 10 positive phrases tested — 10/10 activate as expected +- [x] 3 false-positive phrases tested — 3/3 not-activate as expected +- [x] Findings reviewed — no fix-tasks needed; 2 design notes recorded above +- [x] Close note appended (in STATUS.md block) + +## Notes +- SKILL.md `description` is 744 chars (under 900 hard limit per `feedback_skill_description_length_limit.md`). +- Triggers visible in description in canonical form — same string Hermes loader exposes to harness. +- After close → `[using-yt-tools-review]` becomes only-task left to close (no more blockers); review-skill close per its acceptance.