chore: add agentic development toolkit (#33)

This commit is contained in:
Aman Tahiliani
2026-07-03 23:18:31 -04:00
committed by GitHub
parent 7263949260
commit 094620fa08
18 changed files with 875 additions and 1 deletions

49
.agents/README.md Normal file
View File

@@ -0,0 +1,49 @@
# `.agents/` — harness-agnostic agentic dev toolkit
Portable skills, prompts, and scripts that drive the box-box development lifecycle.
Any harness (Claude, Codex, opencode, …) can read these — the canonical workflows
live here, not in a tool-specific folder. The shared project context every harness
reads is `AGENTS.md` (→ `CLAUDE.md`).
## Layout
```
skills/ Codex/open agent skills: groom, write-spec, implement, review, lenses
personas/ grill.md (base) + lens overlays (architect, …) — the interrogation voices
prompts/ ready-spec.md (groomed spec), implement.md/review.md dispatch prompts
lib/gh.sh GitHub issue + Project (#2) state helpers: issue_*, set_stage/effort/priority
lib/dispatch.sh dispatch(): Ready issue → worktree → harness → gate → PR
harnesses.sh headless adapters (one fn per harness) + run_gate — the ONLY tool-specific code
bin/dev CLI: `dev implement <issue#> --harness <name> [--dry-run]`
```
## The lifecycle
`Icebox → Research → Ready → In Progress → In Review → Done` (the Project `Stage` field).
- **Groom** (interactive, Claude): `/groom <issue#>` runs a seeded grill-me → writes a
Ready spec into the issue body → sets Effort/Priority → leaves Stage at `Research`.
You review and flip to `Ready`.
- **Implement** (any harness): `.agents/bin/dev implement <issue#> --harness <name>`
(or `/implement …` in Claude to supervise) → isolated worktree → runs the harness
headless on the spec → build gate → opens a PR → sets Stage `In Review`.
- **Review + merge**: use the `review` skill from a harness different from the
implementer to create a local review packet and PR comment, then you merge.
## Skills and harnesses
`.agents/skills` is the canonical home for reusable workflows. Codex discovers
repo skills from that path directly, and Claude can use the same files through
`.claude/skills -> ../.agents/skills`. Other harnesses can read the same
`SKILL.md` files explicitly or enter the workflow through `.agents/bin/dev`.
Do not put canonical workflow instructions under `.claude/`; that directory is
local adapter state.
## Adding / fixing a harness
Edit one function in `harnesses.sh`: `harness_<name> <workdir> <promptfile>`, running the
tool non-interactively in `<workdir>` on the prompt. `claude`/`codex`/`opencode` are
wired; `pi`/`cursor` are stubs — confirm their headless flags before trusting.
Always `--dry-run` a new harness first: it renders the exact prompt and plan, touching
nothing (no worktree, PR, or state change).

45
.agents/bin/dev Executable file
View File

@@ -0,0 +1,45 @@
#!/usr/bin/env bash
# .agents/bin/dev — box-box agentic dev CLI. Works from any harness's shell.
#
# dev implement <issue#> --harness <claude|codex|opencode|pi|cursor> [--dry-run] [--base <branch>]
#
# Grooming is driven interactively via the Claude Code /groom skill; this CLI covers
# the implement lane (dispatch a Ready issue to a harness → worktree → gate → PR).
set -o pipefail
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]:-$0}")/../.." && pwd)"
usage() {
cat >&2 <<EOF
box-box dev CLI
dev implement <issue#> --harness <name> [--dry-run] [--base <branch>]
harnesses: claude, codex, opencode (supported) · pi, cursor (verify flags in .agents/harnesses.sh)
--dry-run render the prompt + plan, touch nothing (no worktree/PR/state change)
--base base branch for the worktree/PR (default: main)
EOF
exit 2
}
cmd="${1:-}"; shift 2>/dev/null || true
case "$cmd" in
implement)
issue="${1:-}"; shift 2>/dev/null || true
harness=""; passthru=()
while [ $# -gt 0 ]; do
case "$1" in
--harness) harness="${2:-}"; shift 2 ;;
--dry-run) passthru+=(--dry-run); shift ;;
--base) passthru+=(--base "${2:-}"); shift 2 ;;
*) echo "unknown arg: $1" >&2; usage ;;
esac
done
[ -n "$issue" ] && [ -n "$harness" ] || usage
# shellcheck source=/dev/null
source "$ROOT/.agents/lib/dispatch.sh"
dispatch "$issue" "$harness" ${passthru[@]+"${passthru[@]}"}
;;
""|-h|--help|help) usage ;;
*) echo "unknown command: $cmd" >&2; usage ;;
esac

52
.agents/harnesses.sh Normal file
View File

@@ -0,0 +1,52 @@
#!/usr/bin/env bash
# .agents/harnesses.sh — headless harness adapters.
#
# ONE function per harness: harness_<name> <workdir> <promptfile>
# Each runs the harness NON-INTERACTIVELY inside <workdir>, feeding it the rendered
# prompt, expected to edit files (and ideally commit). This is the ONLY harness-specific
# code in the toolkit — tune the exact flags here per tool/version.
#
# Autonomy note: these run unattended, so each uses the tool's "just do it" mode
# (accept edits / full-auto). Only run harnesses you trust on code you can review via PR.
# ---- MUST-HAVE ----
harness_claude() { # Claude Code — print mode, auto-accept edits
local dir="$1" prompt="$2"
( cd "$dir" && claude -p "$(cat "$prompt")" --permission-mode acceptEdits )
}
harness_codex() { # OpenAI Codex CLI — non-interactive exec, full auto
local dir="$1" prompt="$2"
( cd "$dir" && codex exec --full-auto "$(cat "$prompt")" )
}
harness_opencode() { # opencode — non-interactive run
local dir="$1" prompt="$2"
( cd "$dir" && opencode run "$(cat "$prompt")" )
}
# ---- NICE-TO-HAVE (verify the exact invocation for your version before trusting) ----
harness_pi() { # pi — CONFIRM headless CLI + flags
local dir="$1" prompt="$2"
( cd "$dir" && pi run "$(cat "$prompt")" ) # placeholder — verify
}
harness_cursor() { # Cursor CLI agent — CONFIRM flags
local dir="$1" prompt="$2"
( cd "$dir" && cursor-agent -p "$(cat "$prompt")" --force ) # placeholder — verify
}
# ---- build/typecheck gate (fast, local) ----
# Returns non-zero on failure. This is a smoke gate — CI runs the full suite. Tune freely.
run_gate() {
local dir="$1"
( cd "$dir" && go build ./... ) || return 1
if [ -d "$dir/frontend/node_modules" ]; then
( cd "$dir/frontend" && npx tsc --noEmit ) || return 1
else
echo " (gate: frontend deps absent in worktree — tsc/vitest deferred to CI)" >&2
fi
return 0
}

121
.agents/lib/dispatch.sh Normal file
View File

@@ -0,0 +1,121 @@
#!/usr/bin/env bash
# .agents/lib/dispatch.sh — implement a Ready issue with a chosen harness in an
# isolated git worktree, run the build gate, and open a PR.
#
# Source it, then: dispatch <issue#> <harness> [--dry-run] [--base <branch>]
# (or use the CLI: .agents/bin/dev implement <issue#> --harness <name> [--dry-run])
_AGENTS_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
# shellcheck source=/dev/null
source "$_AGENTS_DIR/lib/gh.sh"
# shellcheck source=/dev/null
source "$_AGENTS_DIR/harnesses.sh"
_slug() { echo "$1" | tr '[:upper:]' '[:lower:]' | sed -E 's/[^a-z0-9]+/-/g; s/^-+//; s/-+$//' | cut -c1-40; }
_render_prompt() { # <issue#> <title> <body>
echo "# Implement: $2"
echo
cat "$_AGENTS_DIR/prompts/implement.md"
echo; echo "---"; echo
echo "## Spec — issue #$1"
echo
echo "$3"
}
_pr_body() { # <issue#> <harness> <gate>
cat <<EOF
Implements #$1.
- **Harness:** $2 (dispatched via \`.agents/bin/dev\`)
- **Local gate** (\`go build\` + \`tsc --noEmit\`): **$3**
- Full test suite + independent review run in CI / by a reviewer harness.
See #$1 for the groomed spec, Test Plan, and Definition of Done.
Closes #$1
EOF
}
dispatch() { # <issue#> <harness> [--dry-run] [--base <branch>]
local issue="$1" harness="$2"; shift 2 || { echo "usage: dispatch <issue#> <harness> [--dry-run] [--base <branch>]"; return 2; }
local dry=0 base="main"
while [ $# -gt 0 ]; do
case "$1" in
--dry-run) dry=1 ;;
--base) base="$2"; shift ;;
*) echo "dispatch: unknown flag '$1'" >&2; return 2 ;;
esac; shift
done
# adapter must exist
if ! declare -f "harness_$harness" >/dev/null 2>&1; then
echo "no adapter for harness '$harness' — add harness_$harness() to .agents/harnesses.sh" >&2; return 2
fi
local repo_root title body slug branch wt prompt
repo_root="$(git rev-parse --show-toplevel)" || return 1
title="$(issue_title "$issue")" || { echo "issue #$issue not found on $REPO" >&2; return 1; }
body="$(issue_body "$issue")"
slug="$(_slug "$title")"
branch="feat/issue-${issue}-${slug}"
wt="$repo_root/.worktrees/issue-${issue}"
prompt="$(mktemp "${TMPDIR:-/tmp}/boxbox-prompt-${issue}.XXXX")"
_render_prompt "$issue" "$title" "$body" > "$prompt"
echo "── dispatch #$issue$harness ──"
echo " title : $title"
echo " branch : $branch"
echo " worktree : $wt"
echo " base : $base"
echo " prompt : $prompt"
if [ "$dry" = 1 ]; then
echo " [dry-run] no worktree / harness / PR / state change. Prompt preview:"
sed 's/^/ | /' "$prompt"
return 0
fi
# preflight: warn (don't block) if not Ready
local stage; stage="$(get_field "$issue" Stage)"
[ "$stage" = "Ready" ] || echo " ⚠ Stage is '$stage' (not Ready) — dispatching anyway"
# isolated worktree
if [ -d "$wt" ]; then
echo " worktree exists — reusing"
else
git -C "$repo_root" worktree add -b "$branch" "$wt" "$base" || return 1
fi
set_stage "$issue" "In Progress"
echo " running $harness (headless)…"
( harness_"$harness" "$wt" "$prompt" ); local hrc=$?
echo " $harness exited ($hrc)"
# fallback commit: guarantee a PR-able branch even if the harness didn't commit
if [ -n "$(git -C "$wt" status --porcelain)" ]; then
git -C "$wt" add -A
git -C "$wt" commit -q -m "feat(#$issue): $title
Implemented by $harness via .agents/dev dispatch." && echo " committed leftover changes"
fi
# gate
local gate="passed"
run_gate "$wt" || gate="FAILED"
echo " gate: $gate"
# PR (only if there are commits ahead of base)
if [ -n "$(git -C "$wt" log "$base..$branch" --oneline 2>/dev/null)" ]; then
git -C "$wt" push -u origin "$branch" || { echo " push failed — inspect $wt" >&2; return 1; }
local draft=""; [ "$gate" = "FAILED" ] && draft="--draft"
local pr
pr="$(gh pr create -R "$REPO" --head "$branch" --base "$base" $draft \
--title "$title (#$issue)" --body "$(_pr_body "$issue" "$harness" "$gate")")" || { echo " gh pr create failed" >&2; return 1; }
echo " PR: $pr${draft:+ (draft — gate failed)}"
set_stage "$issue" "In Review"
else
echo " no commits on $branch — leaving Stage 'In Progress'. Inspect the worktree: $wt" >&2
fi
}

68
.agents/lib/gh.sh Normal file
View File

@@ -0,0 +1,68 @@
#!/usr/bin/env bash
# .agents/lib/gh.sh — harness-agnostic helpers for the box-box agentic dev lifecycle.
#
# Source it, then call the functions:
# source .agents/lib/gh.sh
# issue_json 9 ; set_stage 9 Research ; set_effort 9 M
#
# Requires: gh (authed, with `project` scope), jq.
# Config is overridable via env vars.
REPO="${BOXBOX_REPO:-AmanTahiliani/box-box}"
PROJECT_OWNER="${BOXBOX_PROJECT_OWNER:-AmanTahiliani}"
PROJECT_NUMBER="${BOXBOX_PROJECT_NUMBER:-2}"
_BOXBOX_CACHE="${TMPDIR:-/tmp}/boxbox-agent"
mkdir -p "$_BOXBOX_CACHE" 2>/dev/null
# ---------- issues ----------
issue_json() { gh issue view "$1" -R "$REPO" --json number,title,body,labels,url,state; }
issue_body() { gh issue view "$1" -R "$REPO" --json body -q .body; }
issue_title() { gh issue view "$1" -R "$REPO" --json title -q .title; }
issue_url() { gh issue view "$1" -R "$REPO" --json url -q .url; }
set_issue_body() { gh issue edit "$1" -R "$REPO" --body-file "$2"; } # <issue#> <file>
add_comment() { gh issue comment "$1" -R "$REPO" --body-file "$2"; } # <issue#> <file>
# Native sub-issue children of an epic (issue numbers, one per line).
sub_issues() {
gh api graphql -H "GraphQL-Features: sub_issues" -f query='
query($owner:String!,$repo:String!,$num:Int!){
repository(owner:$owner,name:$repo){
issue(number:$num){ subIssues(first:50){ nodes{ number } } } } }' \
-F owner="${REPO%/*}" -F repo="${REPO#*/}" -F num="$1" \
-q '.data.repository.issue.subIssues.nodes[].number'
}
# ---------- project fields (cached per shell invocation) ----------
_fields_file="$_BOXBOX_CACHE/fields.json"
_items_file="$_BOXBOX_CACHE/items.json"
_pid_file="$_BOXBOX_CACHE/project_id"
_refresh_fields() { gh project field-list "$PROJECT_NUMBER" --owner "$PROJECT_OWNER" --format json > "$_fields_file"; }
_refresh_items() { gh project item-list "$PROJECT_NUMBER" --owner "$PROJECT_OWNER" --format json --limit 200 > "$_items_file"; }
project_refresh() { _refresh_fields; _refresh_items; } # call once at the start of a session to get fresh state
_project_id() { [ -s "$_pid_file" ] || gh project view "$PROJECT_NUMBER" --owner "$PROJECT_OWNER" --format json | jq -r .id > "$_pid_file"; cat "$_pid_file"; }
_field_id() { [ -s "$_fields_file" ] || _refresh_fields; jq -r --arg n "$1" '.fields[]|select(.name==$n)|.id' "$_fields_file"; }
_option_id() { [ -s "$_fields_file" ] || _refresh_fields; jq -r --arg f "$1" --arg o "$2" '.fields[]|select(.name==$f)|.options[]?|select(.name==$o)|.id' "$_fields_file"; }
_item_id() { [ -s "$_items_file" ] || _refresh_items; jq -r --arg n "$1" '.items[]|select(.content.number==($n|tonumber))|.id' "$_items_file"; }
# set_field <issue#> <FieldName> <OptionName> (single-select fields: Stage/Priority/Effort/Phase)
set_field() {
local item opt fld pid
item="$(_item_id "$1")"; fld="$(_field_id "$2")"; opt="$(_option_id "$2" "$3")"; pid="$(_project_id)"
if [ -z "$item" ] || [ -z "$fld" ] || [ -z "$opt" ]; then
echo "set_field: could not resolve issue=$1 field=$2 option=$3 (item=$item field=$fld opt=$opt)" >&2; return 1
fi
gh project item-edit --id "$item" --project-id "$pid" --field-id "$fld" --single-select-option-id "$opt" >/dev/null \
&& echo "set #$1 $2=$3"
}
set_stage() { set_field "$1" Stage "$2"; }
set_priority() { set_field "$1" Priority "$2"; }
set_effort() { set_field "$1" Effort "$2"; }
# get_field <issue#> <FieldName> -> current value (single-word field names only)
get_field() {
[ -s "$_items_file" ] || _refresh_items
jq -r --arg n "$1" --arg f "$2" '.items[]|select(.content.number==($n|tonumber))|.[($f|ascii_downcase)] // "-"' "$_items_file"
}

View File

@@ -0,0 +1,35 @@
# Lens overlay: Architect
Apply this **on top of** the base grill (`grill.md`). Bias every question toward
technical soundness and fit with the existing box-box architecture.
## box-box context to hold
- **Backend (Go):** `internal/api` (OpenF1 client, cache, 37 endpoints), `internal/web`
(REST handlers + SSE hub, route table in `server.go`), `internal/store` (domain
SQLite), `internal/query` (read models over the domain DB), `internal/ingest`,
`internal/live` (SignalR).
- **Frontend (React+Vite+TS):** TanStack Router/Query, `src/api.ts` (typed fetchers),
`src/types.ts` (payload mirrors), `src/lib` (client helpers), `src/pages`,
`src/components` (incl. `components/live`).
- **Patterns to respect** (see CLAUDE.md "How To Extend"): ServeMux longest-prefix
route ordering in `server.go`, cache TTL tiers, `?source=openf1|local|auto`
resolution, two-phase standings load, lazy tab loads, stale-data fallback banner.
## Grill especially on
- **Reuse vs new:** does existing code already do this (a TUI equivalent in
`internal/ui/*.go`, a query model, an `api.ts` fetcher)? Port vs rebuild.
- **Data flow & source:** OpenF1 live vs domain DB vs cache; payload size; rate
limits; how `?source` is handled.
- **Seams:** which files/modules change; new endpoint (mind registration order!) vs
extend an existing one; new component vs extend; where shared logic lives
(`frontend/src/lib/*`).
- **Testability:** how does this land in `go test` / `vitest` / hermetic Playwright?
What seam makes it testable without live OpenF1?
- **Risk:** domain-DB migrations, perf on large sessions, backward compat, and
failure / stale-data behavior.
Keep questions concrete and decision-shaped — e.g. *"port the GPS normalization from
`internal/ui/trackmap.go`, or recompute in a shared `frontend/src/lib/trackmap.ts` so
it's unit-testable?"* — each with your recommendation.

51
.agents/personas/grill.md Normal file
View File

@@ -0,0 +1,51 @@
# Persona: The Grill
You run a **grill-me** session. Instead of the human prompting you, **you interrogate
the human** until you share a design concept for one specific piece of work (a GitHub
issue or epic). The shared understanding — not the document — is the real output.
## Before you ask anything
Load the full context of the target:
- Read the issue title + body (and any notes already on it).
- Read `CLAUDE.md`.
- Explore the code paths the work implicates.
## Rules
1. **One question at a time.** Walk each branch of the design tree and resolve
dependencies in order — a later question often depends on an earlier answer.
2. **Recommend, consequence-first.** Every question carries your recommended answer
and a short "why". **Calibrate to a technically fluent reader who does not know
*this project's* internals.** Assume general engineering literacy (APIs, streaming,
latency, front/back-end, caching, etc.) and don't explain those. **Do** unpack
anything project-specific: internal file/type/endpoint names, bespoke architecture
choices, and why they matter *here* — a few words is enough, no lectures. Above all,
lead each option with the **practical consequence** a decision can be made on
(effort, risk, what ships sooner, how it feels to use), so the reader can choose
without needing the implementation detail. Recommendation first, with why it's the
better call **for them**.
- *Example — keep the mechanism, but lead with the tradeoff:* "**A (recommended):**
reuse the existing SSE snapshot — cheapest to build, but cars jump a little
between updates. **B:** a dedicated ~4Hz position stream — more work now, but
motion is smooth and it sets up interpolation later." (Names the real mechanism;
the choice is still obvious from the consequences.)
3. **Hybrid asking.**
- Decision with clear discrete options → present a **structured choice**, the
recommendation first. *(In Claude Code: use the AskUserQuestion tool; put the
recommended option first and end its label with "(Recommended)".)* Write each
option's description in the plain-language, consequence-first style from rule 2 —
the label can be terse, but the description must be understandable on its own.
- Genuinely open-ended → ask in **prose**.
4. **Explore before you ask.** If the codebase or the issue already answers a
question, do **not** ask — state what you found and the assumption you're
proceeding with, then move on. Only ask about real forks the human must decide.
5. **Stay in scope.** Grill the design of *this* work, not the whole app. Note
out-of-scope temptations instead of chasing them.
## Termination
Stop when no unresolved branches remain and you could write the spec yourself with no
open questions. Summarize the shared design concept in 36 bullets, confirm it with
the human, then hand off to `write-spec`.

View File

@@ -0,0 +1,21 @@
You are an autonomous coding agent working in an **isolated git worktree** on the
**box-box** repo. Implement the groomed spec below as a single, focused, story-sized
change — then commit it.
## Ground rules
- Read `AGENTS.md` / `CLAUDE.md` first and follow the project's conventions exactly
(architecture, route-registration order in `server.go`, cache TTL tiers,
`api.ts`/`types.ts` mirrors, test layout).
- Implement ONLY this story's scope. Honor the spec's **Out of Scope** — do not build
deferred items, even if tempting.
- If the spec has an **early spike / risk** step, do that FIRST and note the result in
your commit message (and adjust the approach if the spike says to).
- Add or extend tests per the **Test Plan**. Make the relevant suites pass:
`go build ./...`, `go test ./...`, and in `frontend/`: `npm run test`, `tsc --noEmit`.
- Keep the change reviewable and story-sized. **Commit your work** with a clear,
conventional message when done (the dispatcher opens the PR).
- Satisfy every item in the spec's **Definition of Done**.
If something in the spec is ambiguous or turns out to be wrong once you're in the code,
make the smallest reasonable decision, implement it, and call it out clearly in the
commit message / PR so the reviewer can catch it — do not silently expand scope.

View File

@@ -0,0 +1,27 @@
## Context
{one paragraph: the problem and why it matters, grounded in the grooming}
## Acceptance Criteria
- [ ] {observable behaviour 1}
- [ ] {observable behaviour 2}
## Technical Approach
- **Files to touch:** {real paths}
- **Data source:** {OpenF1 live | domain DB | cache | ...}
- **Endpoints/components:** {new or extended — note route-registration order if new}
- **Key decisions:** {the forks resolved during the grill}
## Test Plan
- {cases mapped to the real suites: go test · vitest · hermetic Playwright}
## Out of Scope
- {explicitly deferred}
## Definition of Done
- [ ] Tests added and green (`go test` · `vitest` · `tsc --noEmit` · hermetic Playwright as applicable)
- [ ] Matches CLAUDE.md conventions (route order, cache TTLs, `api.ts`/`types.ts` mirrors)
- [ ] No console / preview errors (UI verified in preview)
- [ ] Story-sized PR, linked to this issue
---
_Groomed {date} · Effort {S|M|L} · Priority {P0|P1|P2} · via /groom_

20
.agents/prompts/review.md Normal file
View File

@@ -0,0 +1,20 @@
You are an independent reviewer for a box-box pull request. Review the PR against
the linked GitHub issue spec, not against your own preferred scope.
## Ground rules
- Use a harness different from the implementer when possible.
- Read `AGENTS.md` / `CLAUDE.md`, the PR body/diff, and the linked issue body.
- Run the relevant local gates and capture logs under `.review/issue-<n>-pr-<pr>/logs`.
- For UI changes, create a local visual packet under `.review/issue-<n>-pr-<pr>/`
with desktop and mobile screenshots for the affected routes.
- If screenshots should appear inline on GitHub, publish only review artifacts to
a separate artifact branch, never to `main` or the product PR branch.
- Post a PR comment with pass/fail status, acceptance-criteria alignment, local
artifact paths, screenshot links when available, and caveats.
- Do not merge. The human owns the merge gate.
## Output
- A local packet with `summary.md`, optional `index.html`, screenshots, logs, and
any visual diffs.
- A GitHub PR comment that makes the review visually scannable.
- A clear recommendation: pass, pass with caveats, or needs changes.

View File

@@ -0,0 +1,42 @@
---
name: groom
description: Groom a box-box GitHub issue into a Ready spec via a seeded grill-me interrogation. Use when the user asks to groom an issue, for example "/groom <issue-number>" or "$groom <issue-number>" with optional "--lens architect". Single-issue path: asks targeted questions, writes a structured spec into the issue body, sets Effort/Priority, and leaves Stage at Research for approval.
argument-hint: <issue-number> [--lens architect]
---
# /groom — refine one issue into a Ready spec
Groom issue **$ARGUMENTS** on the box-box roadmap (Project #2): drive it from the
backlog into a fully-specified Ready ticket through a grill-me session.
## 0. Setup
- Parse the first token of the arguments as the **issue number**. An optional
`--lens <name>` pulls in a lens overlay (currently: `architect`).
- Run: `source .agents/lib/gh.sh && project_refresh` (fresh field/item state).
- Load context **before asking anything**: `issue_json <n>`, read `CLAUDE.md`, and
explore the code paths the issue implicates.
- Move it into grooming if it isn't already there: `set_stage <n> Research`.
## 1. Grill
- Read `.agents/personas/grill.md`. If `--lens <name>` was given, also read
`.agents/personas/<name>.md` and apply it on top.
- Run the grill exactly per those rules: **one question at a time**, recommendation
first, **hybrid** asking (AskUserQuestion for discrete decisions with the
recommended option first and labelled "(Recommended)"; prose for open-ended), and
**explore the code to self-answer** wherever possible — only ask about genuine forks.
- Track the resolved decisions as you go.
## 2. Synthesize
- When no open branches remain, summarize the shared design concept in 36 bullets and
confirm it with the user.
- Then invoke the **write-spec** skill for issue `<n>`, handing it the resolved
decisions, so it renders `.agents/prompts/ready-spec.md` into the issue body and
sets Effort + Priority.
## 3. Hand back (human gate)
- Do **not** auto-advance to Ready — that's the user's call. Report that the spec is
written, Stage is `Research`, and they should review the issue and flip Stage →
`Ready` when satisfied (`set_stage <n> Ready`).
- Print the issue URL (`issue_url <n>`).
Stay focused on THIS issue's design throughout. Note but don't chase out-of-scope ideas.

View File

@@ -0,0 +1,32 @@
---
name: implement
description: Dispatch a Ready box-box issue to a coding harness (claude/codex/opencode/pi/cursor) in an isolated git worktree, run the build gate, and open a PR. Use when supervising implementation via "/implement <issue-number> --harness <name> [--dry-run]" or from any harness terminal with .agents/bin/dev.
argument-hint: <issue-number> --harness <claude|codex|opencode|pi|cursor> [--dry-run]
---
# /implement — dispatch an issue to a harness (supervised)
Supervise the implementation of issue **$ARGUMENTS**. You are SUPERVISING, not coding —
a fresh harness does the work in its own isolated worktree with clean context. Do not
edit project files yourself.
## Steps
1. Parse: `<issue-number> --harness <name> [--dry-run] [--base <branch>]`.
2. **Preflight (report, don't hard-block):** `source .agents/lib/gh.sh` and check
`get_field <n> Stage` is `Ready` and the issue body has an "## Acceptance Criteria"
section (a groomed spec). If it's not Ready or has no spec, say so and recommend
`/groom <n>` first — proceed only if the user confirms.
3. **Dispatch:** run `.agents/bin/dev implement <n> --harness <name> [flags]`. For a
first run against an unfamiliar harness, suggest `--dry-run` first so the user can
eyeball the prompt.
4. **Report the outcome:** branch, worktree path, gate result (pass/FAILED → draft PR),
and the PR URL. On success, Stage will be `In Review`.
5. If no PR was created (no changes, or push failed), surface exactly why and point at
the worktree (`.worktrees/issue-<n>`) so the user can inspect. Diagnose from the
dispatcher output; recommend a fix or re-run — don't silently take over the coding.
## Notes
- The harness adapters and the gate live in `.agents/harnesses.sh` — the single place
to tune per-tool flags.
- From any harness shell, run the same thing directly:
`.agents/bin/dev implement <n> --harness <name>`.

View File

@@ -0,0 +1,22 @@
---
name: lens-architect
description: Grill/analyze a box-box issue or epic from a software-architecture perspective and post the findings as a comment. Use standalone as "/lens-architect <issue-number>" for an on-the-fly architecture review, or let /groom compose it via "--lens architect". Reads the base grill + architect persona and focuses on reuse, data flow, seams, testability, and risk.
argument-hint: <issue-number>
---
# /lens-architect — architecture lens
Target: issue **$ARGUMENTS** (box-box, Project #2).
1. `source .agents/lib/gh.sh`; load context (`issue_json <n>`, `CLAUDE.md`, and the
relevant code paths).
2. Read `.agents/personas/grill.md` + `.agents/personas/architect.md` and run a focused
grill from the architecture lens: hybrid asking (AskUserQuestion for discrete
decisions, recommendation first; prose otherwise), recommend every answer, and
explore the code to self-answer before asking.
3. When aligned, write an **"## Architecture review"** summary (decisions taken,
files/seams affected, risks, the test seam) to a temp file and `add_comment <n> <file>`.
4. Print the issue URL.
If invoked from **within /groom**, skip the comment — instead return the architecture
decisions inline so groom can fold them into the spec.

View File

@@ -0,0 +1,53 @@
---
name: review
description: Locally review a box-box PR against its linked GitHub issue spec, run tests, capture visual screenshots when applicable, create a .review packet, and post a GitHub PR comment. Use when a ticket implementation is ready for independent local review before the human merge gate.
argument-hint: <pr-number> [--harness <name>] [--publish-screenshots]
---
# /review — local independent PR review
Review PR **$ARGUMENTS** for box-box using the local-only lifecycle. You are
reviewing, not implementing. Do not merge the PR.
## 0. Setup
- Parse the PR number. Optional `--harness <name>` records which reviewer harness is
acting; optional `--publish-screenshots` allows pushing visual artifacts to a
dedicated artifact branch.
- Read `.agents/prompts/review.md`, `AGENTS.md`, the PR metadata/diff, and the linked
issue body.
- Identify the implementer harness from the PR body when present. If it matches the
reviewer harness, call that out as a reduced independence caveat.
- Create `.review/issue-<issue>-pr-<pr>/logs`, `screenshots`, and `artifacts`.
## 1. Verify
- Run the smallest meaningful gates first, then broaden based on risk:
`go test` for touched Go packages, `npm run test` and `npm run build` for frontend
changes, and hermetic Playwright when user-facing routes changed.
- Save all command output to `.review/issue-<issue>-pr-<pr>/logs`.
- Inspect the diff for spec conformance, scope leaks, missing tests, and known project
conventions from `AGENTS.md`.
## 2. Visual Packet
- If the PR changes UI, start a local seeded or mocked preview and capture desktop and
mobile screenshots for affected routes.
- Prefer hermetic mocks/seeded data over live external state.
- Save screenshots under `.review/issue-<issue>-pr-<pr>/screenshots`.
- Create a concise `summary.md`; create `index.html` when screenshots exist.
## 3. Publish
- If `--publish-screenshots` is present, publish only review artifacts to a dedicated
branch such as `review-artifacts/pr-<pr>/` and use raw GitHub URLs in the comment.
- Post a PR comment with:
- reviewer harness and implementer harness
- result: pass, pass with caveats, or needs changes
- local packet path
- gates run and results
- acceptance-criteria checklist
- screenshots or artifact links when available
- caveats that the human must inspect
- If the result is pass/pass-with-caveats, set the linked issue's custom Project
`Stage` to `In Review` using `.agents/lib/gh.sh`. Do not set `Done`.
## 4. Hand Back
- Tell the human exactly what to open locally and what decision remains theirs.
- Do not merge or delete worktrees.

View File

@@ -0,0 +1,30 @@
---
name: write-spec
description: Render a groomed design into the box-box Ready-spec template and write it into a GitHub issue body, then set Effort and Priority. Called by /groom after a grill session, or run standalone as "/write-spec <issue-number>" to (re)write an issue's spec from agreed decisions. Does not change Stage.
argument-hint: <issue-number>
---
# /write-spec — write the Ready spec into an issue
Target issue: **$ARGUMENTS** (box-box, Project #2).
## Steps
1. `source .agents/lib/gh.sh`
2. Gather the agreed design decisions: from the current grooming conversation if one
is in progress; otherwise ask the user for the key points, or read the issue and
explore the code to draft them and confirm.
3. Read `.agents/prompts/ready-spec.md` and fill every placeholder:
- Concrete, **behavioural** acceptance criteria (checkboxes).
- Technical approach grounded in **real files/paths** and the chosen data source.
- Test plan mapped to the actual suites (`go test` · `vitest` · hermetic Playwright).
- Explicit out-of-scope.
- Keep the Definition of Done checklist verbatim.
- Stamp the footer: date (from the current-date context), Effort (S/M/L),
Priority (P0P2).
4. Write it into the issue body: save the filled template to a temp file under the
scratchpad and `set_issue_body <n> <file>`. The spec is the single source of truth —
only preserve prior body text that captures decisions the spec doesn't.
5. Set fields: `set_effort <n> <S|M|L>` and `set_priority <n> <P0|P1|P2>`.
6. Print the issue URL and a one-line summary of what was written.
Do **not** change Stage — `/groom` owns state transitions.

1
.claude/skills Symbolic link
View File

@@ -0,0 +1 @@
../.agents/skills

9
.gitignore vendored
View File

@@ -22,7 +22,8 @@ frontend/node_modules/
frontend/dist/ frontend/dist/
# Local Claude workspace settings # Local Claude workspace settings
.claude/ .claude/*
!.claude/skills
# Playwright # Playwright
node_modules/ node_modules/
@@ -31,3 +32,9 @@ node_modules/
/blob-report/ /blob-report/
/playwright/.cache/ /playwright/.cache/
/playwright/.auth/ /playwright/.auth/
# agentic dev: isolated implementation worktrees
.worktrees/
# agentic dev: local review packets
.review/

198
docs/PRODUCT_ROADMAP.md Normal file
View File

@@ -0,0 +1,198 @@
# box-box — Product & UX Roadmap
> An F1-geek + product/UX analysis of box-box as an F1 companion, and a
> prioritized plan to make it richer than "data in tables" — for both live race
> weekends and the quiet stretch between rounds.
## Where the app is today (honest read)
box-box has a strong data layer with two front doors (TUI + web) and six web
surfaces:
| Surface | What it does today | Where it's thin |
|---|---|---|
| **Command Center** (`/`) | Hero weekend + countdown, season calendar, championship snapshot, weekend schedule, news | Strong launcher, but tells you *when*, rarely *what to care about*. It's a menu, not a companion |
| **Race Hub** (`/race-hub`) | Per-session tabs: Overview, Race Story (position-evolution scrubber), Strategy (stints), Laps, Conditions, Race Control | `RaceStoryCanvas` is the best "beyond tables" work in the app. Everything else is still tables |
| **Live Timing** (`/live`) | SSE timing tower, battle chips, gap sparklines, pinned drivers, RC feed | Solid tower, but it's "a spreadsheet that updates" — no track map, no telemetry, no "what just happened" |
| **Championship** (`/championship`) | Standings + hub stats (wins/poles/form) + points simulator | Good; the simulator is a genuine differentiator |
| **Briefing** (`/briefing`) | 7 RSS feeds (FIA, BBC, Autosport, RaceFans, Guardian, RACER, F1 YouTube) + readability extraction | Passive reading list, not tied to the season narrative |
| **Data Library** (`/admin`) | Ingestion / coverage admin | Fine as-is |
**Biggest finding:** the app already *pulls* the richest data in F1 — GPS
`Location`, full car telemetry (throttle / brake / DRS / gear / rpm / speed),
team-radio audio, mini-sector `Segments`, speed traps, overtakes — and the web
app visualizes almost none of it. The **track map exists only in the TUI**
(`internal/ui/trackmap.go`); telemetry is fetched and shown to nobody. This is a
large latent asset.
## Core product thesis
An F1 companion serves two different jobs; the app currently treats them the same
(data, in tables, per session):
1. **On a race weekend — the "second screen."** The user is watching the
broadcast (or can't, and wants to *feel* it). They want: what's happening now,
why it matters, what to watch next. The broadcast supplies emotion; box-box
should supply **the data the broadcast doesn't show** — the delta the director
cut away from, the undercut developing, the tyre cliff approaching.
2. **Between races — "understand the season."** No live action. The user wants to
make sense of what happened and anticipate what's next: rewatch the story of
the last race, argue strategy, track the title fight, get smart before the next
round.
Everything below makes each job *feel* like a companion instead of a database.
---
## Race-weekend experience ("second screen")
### 1. Live Track Map — highest-leverage missing feature
Port the TUI's GPS outline (`trackmap.go`, `GetLocation`) to a web SVG dot-map:
cars as team-colored dots on the circuit outline, DRS zones highlighted, sectors
tinted by status. Turn mini-sectors purple/green live using the `Segments` data
already fetched. Tap a car → mini-telemetry readout. This is what turns "updating
spreadsheet" into "I'm watching the race."
### 2. "What just happened" synthesized event rail — the companion voice
Combine Race Control + overtakes + pit stops + position deltas into plain-English
beats: *"LAP 34 — VER pits (2.4s), rejoins P4 behind NOR — undercut on RUS is
live."* The inputs all exist; this is a synthesis layer, not new data. It's the
difference between *data* and *commentary*.
### 3. Telemetry compare overlay
On the tower / a driver panel, pick two drivers → overlaid speed/throttle/brake
traces + delta-time graph for their last comparable lap. *The* tifosi feature —
"where did Leclerc lose the lap." `/api/v1/laps/comparison` + car data already
exist; this needs a chart, not a table.
### 4. Live strategy / tyre-degradation view
Stints + tyre age + pit-lane times → live pit window and tyre-cliff panel, with
**undercut/overcut threat** indicators from gap-to-car-behind vs. pit-loss time.
The pit-window calculator exists in the TUI (`pitwindow.go`); bring it to the web
and make it live.
### 5. Team radio, surfaced
`GetTeamRadio` returns audio clips. Add a "Radio" ticker on the live page — play
button + driver + timestamp. Peak emotional content, and no rival dashboard has it
inline.
### 6. Session-aware Command Center
When a session is live, the home hero should pull the top battles, leader gap, and
last RC flag onto the front page instead of only saying "LIVE." Make the front
page reactive to the moment.
---
## Between-races experience ("understand the season")
### 7. Race replay as a first-class story
Grow `RaceStoryCanvas` from a scrubber into a narrative replay: auto-generated
"chapters" (start, first pit phase, VSC, decisive overtake, finish) with a
headline each (derived from RC + position swings), and scrub the position graph
and track map together. The "watch the race in 90 seconds" mode that makes people
open the app on a Tuesday.
### 8. Driver pages + rivalry view (a real gap)
No driver profile exists in the web app. Add a **driver page** (season form,
teammate H2H — already computed in the champ hub — quali vs. race pace, tyre
management, track-by-track) and a **rivalry view** (two drivers → cumulative
points, H2H, gap-over-season). Feeds the argument every fan has.
### 9. Next-race preview / "get smart" page
Between races the app goes quiet. Fill it: circuit characteristics, last year's
result (2023/24 is cached), typical strategy (1 vs 2 stop), DRS zones, weather
outlook, and the storylines (title-fight math to watch). Turn dead air into
anticipation.
### 10. Championship scenario narratives
Extend the simulator from "drag points around" to narrative permutations:
*"VER clinches if he outscores NOR by 9 this weekend"* / *"first race McLaren can
seal constructors'."* The stuff fans actually search for.
### 11. Briefing → calendar-aware season digest
Reframe the news reader into a paddock digest tied to the calendar: group by GP,
tag by team/driver, surface a "since last race" summary. Optionally an
LLM-generated weekly briefing (a `claude-haiku` summarization pass over ingested
RSS — cheap and on-brand).
---
## New / better data sources
- **Jolpica (Ergast successor, `api.jolpi.ca`)** — free historical results back to
1950: qualifying, pit stops, lap times, circuit metadata. Unlocks all-time
records, "best-ever at this track," and career stats OpenF1 (2023+) can't give.
High value for driver pages and the "get smart" preview.
- **OpenF1 weather timeseries** — already fetched; plot it as a session-long strip
(track temp / rain / wind) instead of a table. Weather narrates strategy.
- **Circuit metadata / DRS zones / corner names** — enriches the track map and
previews. Some is in OpenF1 circuit info; curate the rest once as static data.
- **Static per-race context** — a tiny curated JSON per round (tyre allocation,
notable stats) goes a long way for previews.
---
## Cross-cutting UX principles
1. **Replace tables with a shape wherever a shape carries the meaning.** Lap times
→ a trace with fastest lap marked. Stints → a horizontal tyre timeline. Gaps →
the sparkline already shipped. Keep tables only where data is genuinely tabular
(standings) — but annotate them.
2. **Always answer "so what?"** Every number sits next to its meaning (a gap next
to "undercut live," a tyre age next to "5 laps from the cliff").
3. **One primary "moment" per screen.** The Command Center should always have a
single obvious "here's what to watch/do now."
4. **Make between-races feel alive.** The app currently rewards you only on
Sundays. Previews, digests, replays, and rivalries give a reason to open it
midweek.
---
## Suggested sequencing (impact × effort)
**Phase 1 — turn live into a companion (highest impact; data already in hand)**
1. Web track map (SVG + GPS + mini-sectors)
2. Telemetry compare overlay (speed/throttle/brake + delta)
3. "What just happened" synthesized event rail
4. Team radio ticker
**Phase 2 — own the between-races window**
5. Narrative race replay (grow `RaceStoryCanvas`)
6. Driver pages + rivalry view
7. Next-race preview page
**Phase 3 — depth & reach**
8. Jolpica/Ergast historical integration + all-time records
9. Championship scenario narratives
10. Calendar-aware briefing digest (optional LLM summaries)
---
## Execution tracking (GitHub)
This roadmap is tracked on GitHub:
- **Project board:** https://github.com/users/AmanTahiliani/projects/2 ("box-box Roadmap")
- **Epics:** issues #2#8 (label `epic`), one per epic above, on Phase milestones
- **Stories:** issues #9#31, wired as native **sub-issues** under their epic (progress rolls up automatically)
- **Labels:** `epic`, `enabler`, `research`, `area:{live,viz,between-races,championship,data,ux}`
- **Milestones:** `Phase 1 — Live Companion`, `Phase 2 — Between-Races`, `Phase 3 — Depth & Reach`
**Board fields:** `Stage` (Icebox → Research → Ready → In Progress → In Review → Done),
`Priority` (P0P2), `Effort` (S/M/L), `Phase` (13).
**Working model — active vs. bank:**
- Only **12 epics active** at a time (currently **E1 Live Race Companion** + **E2 Viz
Primitives**); the other five epics are the theme-level idea bank.
- `Stage = Icebox` is the story-level bank. The 10 E1/E2 stories are seeded to
`Ready`; everything else is `Icebox`. Promote a handful to `Ready`/`In Progress`
per cycle; use `Research` to scope a vague idea before it's `Ready`.
**Board views** (built):
- **Backlog** — Table, all epics with expandable sub-issues: the full bank.
- **Board** — grouped by `Stage`, filtered `-stage:Icebox`: the active-WIP wall.
- **Roadmap** — Table grouped by `Phase` (Phase 1/2/3). Note: this is a
phase-grouped table, not a timeline. A true timeline roadmap needs a date or
iteration field (GitHub won't draw/persist a roadmap layout without one); add a
`Target date` field and populate it once phases have real target dates, then
switch this view to the Roadmap layout.