meta(tasks): pause [active-platform-eval] after design + pre-flight
Combined the two backlog tasks [active-platform-tuning] + [active-platform-eval] into a single workstream. Eval IS the tuning mechanism; "wait for 5 real signals" was a placeholder replaced by a 20-query synthetic eval set balanced across Win/Lin/Mac. Spec written at .wiki/concepts/active-platform-eval-design.md (~150 lines): 20 queries (>=3 should-trigger per OS + near-miss negatives), run_loop.py 5-iter autoloop in parallel with manual body sweep (WSL clarity, BSD/macOS expansion, ambiguity policy). Workspace at .tasks/active-platform-eval/ (eval-set.json committed, iterations gitignored). Version bump 1.0.0 -> 1.1.0 planned (MINOR). Pre-flight verified: claude CLI on PATH at C:\nvm4w\nodejs\claude.ps1 (Claude Code 2.1.128); run_loop.py present in skill-creator install. Both autoloop deps satisfied -- no fallback to manual single-pass. Per-task file at .tasks/active-platform-eval.md (Goal, Key files, Decisions log, Open questions, Notes). STATUS.md collapsed the two original blocks into one paused block; resume point is Q2 (write 20 queries solo vs run skill-creator HTML-review template). Also fixed in same pause: [install-ps1] STATUS scope expanded to paired install.sh + install.ps1, cross-platform parity, prune flag (lesson from [compress-dedup]). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -29,6 +29,7 @@ Catalog of all wiki pages. One line per page, organized by type. Updated on ever
|
||||
- [wiki-realignment.md](concepts/wiki-realignment.md) — fixing `project-bootstrap` to create the Karpathy-canonical wiki layout
|
||||
- [interns-design](concepts/interns-design.md) — interns-design
|
||||
- [compress-dedup.md](concepts/compress-dedup.md) — `skills/compress/` deleted as a byte-identical dupe of `skills/caveman-compress/`; canonical kept for README + SECURITY + caveman-toolkit branding; better Process-step wording ported across; `version: 1.0.0` added to caveman-compress frontmatter
|
||||
- [active-platform-eval-design.md](concepts/active-platform-eval-design.md) — spec for eval-driven tuning of `active-platform`: combine the two ⚪ tasks into one workstream, 20-query cross-platform eval set (≥3 per OS + near-miss negatives), `run_loop.py` autoloop **in parallel** with manual body sweep (WSL / BSD / ambiguity), version 1.0.0 → 1.1.0 (MINOR). Status: paused after design + pre-flight check, before eval-set authorship
|
||||
|
||||
## Packages
|
||||
|
||||
|
||||
Reference in New Issue
Block a user