diff --git a/README.md b/README.md index 18686eb..7a68164 100644 --- a/README.md +++ b/README.md @@ -115,6 +115,8 @@ an explicit `adapted-from` marker in its frontmatter. | `find-skills` | `adapted-from: vercel-labs/skills` (MIT) — vendored copy, upstream pin TBD | | `grilling` | `adapted-from: mattpocock/skills @ 84fdeffd` (MIT) — family collapsed to one skill (pi hides `disable-model-invocation` wrappers) | | `brainstorming` | `adapted-from: obra/superpowers @ 6.2.0` (MIT) — divergent phase, visual-companion dropped | +| `diagnosing-bugs` | `adapted-from: mattpocock/skills @ 84fdeffd` (MIT) + superpowers 6.2.0 concepts (Iron Law, red flags) | +| `writing-skills` | `adapted-from: obra/superpowers @ 6.2.0` (MIT) — TDD-for-skills core + ideya 8 self-skill-authoring | | all other `skills/*` | `author: ours` | Adaptation policy: a clone is rewritten to our conventions (`.tasks/` boards, diff --git a/README.ru.md b/README.ru.md index b42558f..943bf6d 100644 --- a/README.ru.md +++ b/README.ru.md @@ -85,6 +85,8 @@ bash scripts/build.sh caveman # один | `find-skills` | `adapted-from: vercel-labs/skills` (MIT) — вендорная копия, пин апстрима TBD | | `grilling` | `adapted-from: mattpocock/skills @ 84fdeffd` (MIT) — семейство схлопнуто в один скил (pi прячет `disable-model-invocation` обёртки) | | `brainstorming` | `adapted-from: obra/superpowers @ 6.2.0` (MIT) — расходящаяся фаза, visual-companion выброшен | +| `diagnosing-bugs` | `adapted-from: mattpocock/skills @ 84fdeffd` (MIT) + superpowers 6.2.0 (Iron Law, red flags) | +| `writing-skills` | `adapted-from: obra/superpowers @ 6.2.0` (MIT) — TDD-for-skills ядро + идея 8 self-skill-authoring | | остальные `skills/*` | `author: ours` | Политика адаптации: клон переписывается под наши конвенции (доски `.tasks/`, diff --git a/dist/diagnosing-bugs.skill b/dist/diagnosing-bugs.skill new file mode 100644 index 0000000..48412d7 Binary files /dev/null and b/dist/diagnosing-bugs.skill differ diff --git a/dist/writing-skills.skill b/dist/writing-skills.skill new file mode 100644 index 0000000..10f88b7 Binary files /dev/null and b/dist/writing-skills.skill differ diff --git a/skills/diagnosing-bugs/SKILL.md b/skills/diagnosing-bugs/SKILL.md new file mode 100644 index 0000000..9759352 --- /dev/null +++ b/skills/diagnosing-bugs/SKILL.md @@ -0,0 +1,298 @@ +--- +name: diagnosing-bugs +adapted-from: mattpocock/skills @ 84fdeffd12f2ee307994d1eb6feb48173b6e0502 (MIT); concepts from obra/superpowers @ 6.2.0 (MIT) +version: 0.1.0 +description: > + Diagnosis loop for hard bugs and performance regressions. Use when the user + says "diagnose"/"debug this", or reports something broken/throwing/failing/ + slow, or any test failure / unexpected behavior / build failure / integration + issue — before proposing fixes. Triggers: «диагностируй», «почему падает», + «разберись с багом», "debug this", "diagnose", "it's broken", "why is it + failing". Cross-agent — no tool refs beyond generic harness commands. +--- + +# Diagnosing Bugs + +A discipline for hard bugs. Skip phases only when explicitly justified. + + +NO FIXES WITHOUT ROOT CAUSE INVESTIGATION FIRST. If you haven't completed +Phase 1 (a tight red-capable feedback loop), you cannot propose fixes. + + +When exploring the codebase, read `CONTEXT.md` (if it exists) to get a clear +mental model of the relevant modules, and check ADRs in the area you're +touching. + +## Redact + +This skill has you show commands, outputs and captured artifacts. **Redact +every secret first** — write `` in its place. Build loops against env +vars, so the credential stays in the environment rather than in what you show. +Captured artifacts carry auth headers: quote only the lines that carry the +signal. + +If the redacted output is not enough to diagnose the bug, say so and ask the +user. + +## Phase 1 — Build a feedback loop + +**This is the skill.** Everything else is mechanical. If you have a **tight** +pass/fail signal for the bug — one that goes red on _this_ bug — you will find +the cause; bisection, hypothesis-testing, and instrumentation all just consume +it. If you don't have one, no amount of staring at code will save you. + +Spend disproportionate effort here. **Be aggressive. Be creative. Refuse to +give up.** + +### Ways to construct one — try them in roughly this order + +1. **Failing test** at whatever seam reaches the bug — unit, integration, e2e. +2. **Curl / HTTP script** against a running dev server. +3. **CLI invocation** with a fixture input, diffing stdout against a known-good + snapshot. +4. **Headless browser script** (Playwright / Puppeteer) — drives the UI, + asserts on DOM/console/network. +5. **Replay a captured trace.** Save a real network request / payload / event + log to disk; replay it through the code path in isolation. +6. **Throwaway harness.** Spin up a minimal subset of the system (one service, + mocked deps) that exercises the bug code path with a single function call. +7. **Property / fuzz loop.** If the bug is "sometimes wrong output", run 1000 + random inputs and look for the failure mode. +8. **Bisection harness.** If the bug appeared between two known states + (commit, dataset, version), automate "boot at state X, check, repeat" so you + can `git bisect run` it. +9. **Differential loop.** Run the same input through old-version vs new-version + (or two configs) and diff outputs. +10. **HITL bash script.** Last resort. If a human must click, drive _them_ + with a structured loop so the captured output feeds back to you. + +Build the right feedback loop, and the bug is 90% fixed. + +### Tighten the loop + +Treat the loop as a product. Once you have _a_ loop, **tighten** it: + +- Can I make it faster? (Cache setup, skip unrelated init, narrow the test + scope.) +- Can I make the signal sharper? (Assert on the specific symptom, not "didn't + crash".) +- Can I make it more deterministic? (Pin time, seed RNG, isolate filesystem, + freeze network.) + +A 30-second flaky loop is barely better than no loop; a 2-second deterministic +one is tight — a debugging superpower. + +### Non-deterministic bugs + +The goal is not a clean repro but a **higher reproduction rate**. Loop the +trigger 100×, parallelise, add stress, narrow timing windows, inject sleeps. A +50%-flake bug is debuggable; 1% is not — keep raising the rate until it's +debuggable. + +### When you genuinely cannot build a loop + +Stop and say so explicitly. List what you tried. Ask the user for: (a) access +to whatever environment reproduces it, (b) a redacted captured artifact (HAR +file, log dump, core dump, screen recording with timestamps), or (c) permission +to add temporary production instrumentation. Do **not** proceed to hypothesise +without a loop. + +### Completion criterion — a tight loop that goes red + +Phase 1 is done when the loop is **tight** and **red-capable**: you can name +**one command** — a script path, a test invocation, a curl — that you have +**already run at least once** (show the invocation and its output, redacted), +and that is: + +- [ ] **Red-capable** — it drives the actual bug code path and asserts the + **user's exact symptom**, so it can go red on this bug and green once + fixed. Not "runs without erroring" — it must be able to _catch this + specific bug_. +- [ ] **Deterministic** — same verdict every run (flaky bugs: a pinned, high + reproduction rate, per above). +- [ ] **Fast** — seconds, not minutes. +- [ ] **Agent-runnable** — you can run it unattended. + +If you catch yourself reading code to build a theory before this command +exists, **stop — jumping straight to a hypothesis is the exact failure this +skill prevents.** No red-capable command, no Phase 2. + +## Phase 2 — Reproduce + minimise + +Run the loop. Watch it go red — the bug appears. + +Confirm: + +- [ ] The loop produces the failure mode the **user** described — not a + different failure that happens to be nearby. Wrong bug = wrong fix. +- [ ] The failure is reproducible across multiple runs (or, for + non-deterministic bugs, reproducible at a high enough rate to debug + against). +- [ ] You have captured the exact symptom (error message, wrong output, slow + timing) so later phases can verify the fix actually addresses it. + +### Minimise + +Once it's red, shrink the repro to the **smallest scenario that still goes +red**. Cut inputs, callers, config, data, and steps **one at a time**, re-running +the loop after each cut — keep only what's load-bearing for the failure. + +Why bother: a minimal repro shrinks the hypothesis space in Phase 3 (fewer +moving parts left to suspect) and becomes the clean regression test in Phase 5. + +Done when **every remaining element is load-bearing** — removing any one of +them makes the loop go green. + +Do not proceed until you have reproduced **and** minimised. + +## Phase 3 — Hypothesise + +Generate **3–5 ranked hypotheses** before testing any of them. +Single-hypothesis generation anchors on the first plausible idea. + +Each hypothesis must be **falsifiable**: state the prediction it makes. + +> Format: "If is the cause, then will make the bug disappear +> / will make it worse." + +If you cannot state the prediction, the hypothesis is a vibe — discard or +sharpen it. + +**Show the ranked list to the user before testing.** They often have domain +knowledge that re-ranks instantly ("we just deployed a change to #3"), or know +hypotheses they've already ruled out. Cheap checkpoint, big time saver. Don't +block on it — proceed with your ranking if the user is AFK. + +## Phase 4 — Instrument + +Each probe must map to a specific prediction from Phase 3. **Change one +variable at a time.** + +Tool preference: + +1. **Debugger / REPL inspection** if the env supports it. One breakpoint beats + ten logs. +2. **Targeted logs** at the boundaries that distinguish hypotheses. +3. Never "log everything and grep". + +**Tag every debug log** with a unique prefix, e.g. `[DEBUG-a4f2]`. Cleanup at +the end becomes a single grep. Untagged logs survive; tagged logs die. + +**Perf branch.** For performance regressions, logs are usually wrong. Instead: +establish a baseline measurement (timing harness, `performance.now()`, +profiler, query plan), then bisect. Measure first, fix second. + +**Multi-component systems:** when the failure path crosses components +(CI → build → signing, API → service → database), before proposing fixes add +diagnostic instrumentation at each component boundary — log what enters, what +exits, and verify environment/config propagation at each layer. Run once to +gather evidence showing WHERE it breaks, then investigate that component. + +## Phase 5 — Fix + regression test + +Write the regression test **before the fix** — but only if there is a +**correct seam** for it. + +A correct seam is one where the test exercises the **real bug pattern** as it +occurs at the call site. If the only available seam is too shallow +(single-caller test when the bug needs multiple callers, unit test that can't +replicate the chain that triggered the bug), a regression test there gives +false confidence. + +**If no correct seam exists, that itself is the finding.** Note it. The +codebase architecture is preventing the bug from being locked down. Flag this +in the post-mortem. + +If a correct seam exists: + +1. Turn the minimised repro into a failing test at that seam. +2. Watch it fail. +3. Apply the fix. +4. Watch it pass. +5. Re-run the Phase 1 feedback loop against the original (un-minimised) + scenario. + +**One change at a time.** No "while I'm here" improvements, no bundled +refactoring. + +### If the fix doesn't work + +- Count how many fixes you've tried. +- If < 3: return to Phase 1, re-analyze with new information. +- **If ≥ 3: STOP and question the architecture.** Each fix revealing new + shared state / coupling / problems in different places is the pattern of an + architectural problem, not a failed hypothesis. Discuss with the user before + attempting more fixes. This is NOT a failed hypothesis — this is a wrong + architecture. + +## Phase 6 — Cleanup + post-mortem + +Required before declaring done: + +- [ ] Original repro no longer reproduces (re-run the Phase 1 loop) +- [ ] Regression test passes (or absence of seam is documented) +- [ ] All `[DEBUG-...]` instrumentation removed (`grep` the prefix) +- [ ] Throwaway prototypes deleted (or moved to a clearly-marked debug + location) +- [ ] The hypothesis that turned out correct is stated in the commit / PR + message — so the next debugger learns + +**Then ask: what would have prevented this bug?** If the answer involves +architectural change (no good test seam, tangled callers, hidden coupling) +hand the specifics off to the project owner / architecture skill. Make the +recommendation **after** the fix is in, not before — you have more information +now than when you started. + +## Red Flags — STOP and return to Phase 1 + +If you catch yourself thinking any of these, stop and go back: + +- "Quick fix for now, investigate later" +- "Just try changing X and see if it works" +- "Add multiple changes, run tests" +- "Skip the test, I'll manually verify" +- "It's probably X, let me fix that" +- "I don't fully understand but this might work" +- "Pattern says X but I'll adapt it differently" +- Proposing solutions before tracing data flow +- "One more fix attempt" (when already tried 2+) +- Each fix reveals a new problem in a different place + +**All of these mean: STOP. Return to Phase 1.** + +## Common Rationalizations + +| Excuse | Reality | +|--------|---------| +| "Issue is simple, don't need process" | Simple issues have root causes too. Process is fast for simple bugs. | +| "Emergency, no time for process" | Systematic debugging is FASTER than guess-and-check thrashing. | +| "Just try this first, then investigate" | First fix sets the pattern. Do it right from the start. | +| "I'll write test after confirming fix works" | Untested fixes don't stick. Test first proves it. | +| "Multiple fixes at once saves time" | Can't isolate what worked. Causes new bugs. | +| "I see the problem, let me fix it" | Seeing symptoms ≠ understanding root cause. | +| "One more fix attempt" (after 2+ failures) | 3+ failures = architectural problem. Question the architecture, don't fix again. | + +## When Process Reveals "No Root Cause" + +If systematic investigation reveals the issue is truly environmental, +timing-dependent, or external: + +1. You've completed the process. +2. Document what you investigated. +3. Implement appropriate handling (retry, timeout, error message). +4. Add monitoring/logging for future investigation. + +**But:** 95% of "no root cause" cases are incomplete investigation. + +## Cross-agent applicability + +Pure methodology — no harness-specific tool references. Works on pi, Claude, +or any agent. The sub-agent mention is a generic capability note; without +sub-agent support the agent looks facts up directly. + +## Out of scope + +- Does NOT cover code review (that's a separate review process). +- Does NOT write the regression-test policy (see `tdd-criteria` for the + bright-line rules on when tests are required). diff --git a/skills/writing-skills/SKILL.md b/skills/writing-skills/SKILL.md new file mode 100644 index 0000000..463ad19 --- /dev/null +++ b/skills/writing-skills/SKILL.md @@ -0,0 +1,208 @@ +--- +name: writing-skills +adapted-from: obra/superpowers @ 6.2.0 (MIT) — TDD-for-skills core; ideya 8 self-skill-authoring (workshop record) +version: 0.1.0 +description: > + Authoring agent skills TDD-style — RED-GREEN-REFACTOR applied to SKILL.md + documents. Use when creating a new skill, editing an existing one, or + verifying a skill works before deploying. Trigger (user): «напиши скил», + «создай скил», "write a skill", "create a skill"; or self-authoring trigger + (Hermes-mode): the third time you do the same thing without instruction, or + the user says «запомни»/«зафиксируй»/«в следующий раз». Anti-sproul: BEFORE + writing a new skill, check the catalog for an existing cover. +--- + +# Writing Skills + +**Writing skills IS Test-Driven Development applied to process documentation.** + +You write test cases (pressure scenarios with subagents), watch them fail +(baseline behavior without the skill), write the skill (documentation), watch +tests pass (agents comply), and refactor (close loopholes). + +**Core principle:** If you didn't watch an agent fail without the skill, you +don't know if the skill teaches the right thing. + + +NO SKILL WITHOUT A FAILING TEST FIRST. This applies to NEW skills AND edits to +existing skills. Write the skill before testing? Delete it. Start over. No +exceptions — not for "simple additions", not for "documentation updates". + + +## Anti-sproul guard (check before writing) + +At 45+ skills in the catalog, "just write another one" is harm, not help. On +any self-authoring trigger, FIRST check for a double/coverage: + +1. Does a skill already exist that covers this? (catalog + adapted-from + sources: mattpocock/skills, obra/superpowers, vendor skills) +2. Should this be a **new skill**, an **extension of an existing one**, a + **rule in CLAUDE.md**, or **nothing** (one-off coincidence)? + +Only proceed to the TDD cycle if the answer is genuinely "new skill". If the +pattern is project-specific, it belongs in the project's `.agents/skills/` +(copy, project owns it) — NOT the catalog. Cross-project value → promote to +the sovereign catalog. + +## What is a Skill? + +A **skill** is a reference guide for proven techniques, patterns, or tools. +Skills help future agents find and apply effective approaches. + +**Skills are:** reusable techniques, patterns, tools, reference guides. +**Skills are NOT:** narratives about how you solved a problem once. + +## TDD Mapping for Skills + +| TDD Concept | Skill Creation | +|---|---| +| **Test case** | Pressure scenario with subagent | +| **Production code** | Skill document (SKILL.md) | +| **Test fails (RED)** | Agent violates rule without skill (baseline) | +| **Test passes (GREEN)** | Agent complies with skill present | +| **Refactor** | Close loopholes while maintaining compliance | +| **Write test first** | Run baseline scenario BEFORE writing skill | +| **Watch it fail** | Document exact rationalizations agent uses | +| **Minimal code** | Write skill addressing those specific violations | +| **Watch it pass** | Verify agent now complies | + +## RED — Write the failing test (baseline) + +Run a pressure scenario with a fresh-context subagent **WITHOUT the skill**. +Document exact behavior: + +- What choices did they make? +- What rationalizations did they use (verbatim)? +- Which pressures triggered violations? + +This is "watch the test fail" — you must see what agents naturally do before +writing the skill. + +**Pressure types for discipline skills** (combine 3+): time pressure, sunk +cost, authority ("the user asked for it"), exhaustion/length, "it's simple". + +## GREEN — Write the minimal skill + +Write the skill that addresses those **specific** rationalizations. Don't add +content for hypothetical cases. + +### Skill structure (our catalog conventions) + +``` +skills//SKILL.md +``` + +Frontmatter (YAML): + +- `name` — letters, numbers, hyphens only. Verb-first, active voice: + `pulling-before-work`, `diagnosing-bugs`, `writing-skills` (gerund works for + processes). +- `description` — **when to use, NOT what it does.** Start with "Use when..." / + trigger phrases. Third person (injected into system prompt). NEVER summarize + the skill's process or workflow — agents follow the description instead of + reading the body. +- `version` — semver; bump on every edit (project-discipline Rule 3). +- `author: ours` or `adapted-from: / @ (license)` for + vendored/adapted copies — with the real upstream pin, not "TBD". + +Body: + +- Overview: core principle in 1-2 sentences. +- When to use: bullet list with symptoms and triggers; when NOT to use. +- Core pattern / Quick reference: table or bullets for scanning. +- Common mistakes / rationalizations: table (excuse → reality). +- Red flags: self-check list ("all of these mean STOP and start over"). +- Cross-agent applicability note (no harness-specific tool refs). +- Out of scope: what this skill explicitly does NOT do. + +**Guidance form must match the failure type:** + +| Baseline failure | Right form | Wrong form | +|---|---|---| +| Skips/violates a rule under pressure | Prohibition + rationalization table + red flags | Soft guidance ("prefer...") | +| Complies, but output has wrong shape | Positive recipe/contract: state what the output IS | Prohibition list | +| Omits a required element | Structural: REQUIRED field in the template | Prose reminders | + +**No nuance clauses.** "Don't X unless it matters" reopens the negotiation — +express a real exception as its own conditional on an observable predicate. + +### Micro-test wording before full scenarios + +Full pressure-scenario runs are the final gate but slow. Verify the wording +first: + +1. One fresh-context sample per call; system prompt = the realistic context. +2. Always include a **no-guidance control**. If the control doesn't exhibit + the failure, there is nothing to fix — stop. +3. 5+ reps per variant. Single samples lie. +4. Manually read every flagged match (template echoes masquerade as hits). +5. Variance is a metric: five different interpretations across five reps means + the wording isn't binding — tighten the form. + +Micro-tests verify wording; they do not replace pressure scenarios for +discipline skills. + +## REFACTOR — close loopholes + +- Agent found a new rationalization? Add an explicit counter. +- Build the rationalization table from all test iterations. +- Create the red-flags list. +- Re-test until bulletproof. + +## Verification checklist (before declaring done) + +- [ ] Baseline failure documented (RED evidence: what the agent did without + the skill) +- [ ] Skill addresses those specific failures (not hypotheticals) +- [ ] Scenario passes WITH the skill (GREEN evidence) +- [ ] Description = when to use only, no workflow summary +- [ ] Frontmatter: name (verb-first, hyphens), description (triggers), version + (bumped), provenance (author/adapted-from with real pin) +- [ ] Lint passes (catalog: `scripts/lint-skills.py`), dist rebuilt + (`scripts/build.sh`), installed (`scripts/install.sh `) +- [ ] README provenance table updated (catalog) +- [ ] Cross-agent — no harness-specific tool references +- [ ] One excellent example, not multi-language dilution + +## Deploying (catalog) + +1. Edit `skills//SKILL.md`. +2. Lint: `python scripts/lint-skills.py`. +3. Build: `bash scripts/build.sh` (→ `dist/.skill`). +4. Install: `bash scripts/install.sh ` (→ `~/.claude/skills/`). +5. Update README provenance table (en+ru). +6. Commit + push; bump version in the commit (project-discipline Rule 3). + +**STOP after each skill.** Don't batch-create skills without testing each one. +Deploying untested skills = deploying untested code. + +## Common Rationalizations for Skipping Testing + +| Excuse | Reality | +|--------|---------| +| "Skill is obviously clear" | Clear to you ≠ clear to other agents. Test it. | +| "It's just a reference" | References can have gaps, unclear sections. | +| "Testing is overkill" | Untested skills have issues. Always. | +| "I'll test if problems emerge" | Problems = agents can't use skill. Test BEFORE deploying. | +| "Too tedious to test" | Testing is less tedious than debugging a bad skill in production. | +| "I'm confident it's good" | Overconfidence guarantees issues. Test anyway. | +| "Academic review is enough" | Reading ≠ using. Test application scenarios. | +| "There's already a similar skill, close enough" | Similar ≠ cover. Verify the double actually covers the failure. | + +**All of these mean: test before deploying. No exceptions.** + +## Cross-agent applicability + +Pure authoring methodology — no harness-specific tool references. The +"subagent" for pressure scenarios is any fresh-context agent run (pi -p, +claude -p, codex exec, hermes headless). Local skills live in the agent's +skills directory (`~/.claude/skills/`, `~/.hermes/skills/`, project +`.agents/skills/`); the catalog deploy steps are our repo's convention. + +## Out of scope + +- Does NOT tell you WHAT skill to write — that's the anti-sproul guard + the + user's call. +- Does NOT cover how to run the catalog pipeline beyond the deploy steps above + (see the repo's `scripts/` and README). +- Does NOT enforce TDD for code — that's `tdd-criteria`.