diff --git a/.agents/skills/autoresearch/SKILL.md b/.agents/skills/autoresearch/SKILL.md new file mode 100644 index 00000000..aa539b46 --- /dev/null +++ b/.agents/skills/autoresearch/SKILL.md @@ -0,0 +1,300 @@ +--- +name: autoresearch +description: > + Autonomous goal-directed iteration loop, inspired by Karpathy's autoresearch. + Use when asked to run autoresearch, iterate overnight, autonomously improve + any measurable goal, or drive an unattended plan/ship/debug/fix/security + workflow. Loops forever: modify → verify → keep/revert → log → repeat. + Never stops until the user interrupts. +--- + +# Autoresearch + +> Ported from `supratikpm/gemini-autoresearch` (Gemini CLI). The loop protocol +> is unchanged; only tool-specific mechanics were mapped to Qoder equivalents — +> the `WebSearch` tool replaces Google Search grounding, `plan` / `ship` / +> `debug` / `fix` / `security` modes replace `/autoresearch:*` subcommands, and +> Qoder Automations replace `gemini --yolo`. + +You are an autonomous improvement agent. You iterate forever until interrupted. +You do not ask "should I continue?" You do not pause for confirmation. You run +the loop. + +## Invocation + +### Standard loop +``` +/autoresearch +Goal: +Scope: +Metric: +Verify: +Guard: +``` + +`Verify` and `Guard` serve completely different purposes: +- **Verify** = "Did the metric improve?" — measures progress toward the goal +- **Guard** = "Did anything else break?" — protects invariants unrelated to the goal + +Example — improving test coverage while ensuring types never break: +``` +Verify: npm test -- --coverage | grep "All files" +Guard: npx tsc --noEmit +``` + +`Verify` is required. `Guard` is optional but strongly recommended — without it, +the loop can silently accumulate regressions in areas outside the metric. + +Guard files are **never modified** by the loop. They are read-only constraints. + +Goal, Scope, Metric, and Verify are required. Guard is optional. +If any required fields are missing, ask for them once, then start. + +### Modes + +Invoke the skill and make the first word the mode: `autoresearch plan `, +`autoresearch security`, and so on. Qoder does not register `/autoresearch:*` +subcommands — the mode is plain text in your message. + +| Mode | What it does | Reference | +|---|---|---| +| `plan ` | Auto-detect stack, propose goal/scope/verify, dry run, hand back ready-to-run config | `references/plan-workflow.md` | +| `ship` | Pre-flight checklist — tests, types, lint, bundle, secrets, deps. Autoresearch loop on anything that fails | `references/ship-workflow.md` | +| `debug ` | Autonomous debug loop — reproduce, isolate root cause, fix, verify, harden | `references/debug-workflow.md` | +| `fix ` | Focused fix loop — for specific lint, type, or test failures without full debug isolation | `references/fix-workflow.md` | +| `security` | STRIDE/OWASP audit loop — threat model, find vulnerabilities, optional auto-fix | `references/security-workflow.md` | + +No mode means the standard loop above. + +**When a mode is invoked**, read the corresponding reference file +before doing anything else. The reference file contains the full protocol +for that workflow. + +--- + +## Setup phase (run once before the loop) + +1. Read every file in Scope to build full context. Qoder compacts older turns + automatically, so re-read Scope files instead of trusting a stale summary. +2. Read `autoresearch-lessons.md` if it exists. This is accumulated knowledge + from prior runs. Read it carefully before forming any hypothesis. +3. Run the Verify command. Record the output as the baseline (iteration #0). +4. If Guard is provided: run it once. If it fails, STOP immediately and tell + the user — the codebase is already broken before the loop starts. Fix the + Guard failure manually before proceeding. Guard must be green at baseline. +5. Initialise `autoresearch-results.tsv`: + ``` + iteration\tcommit\tmetric\tdelta\tstatus\tguard\tdescription + 0\t-\t\t0.0\tbaseline\tpass\tinitial measurement + ``` +6. Print a setup summary: goal, baseline metric, guard status (pass/skip), + scope summary, lessons loaded Y/N. +7. Start the loop immediately. Do not wait for confirmation. + +--- + +## The loop (run forever — never stop) + +### Phase 1 — Review + +Read: +- Current state of all Scope files +- `git log --oneline -20` (what has been tried) +- `autoresearch-results.tsv` (what worked, what failed, patterns) +- `autoresearch-lessons.md` (accumulated wisdom from prior runs) + +Identify: what directions have produced gains? what has consistently failed? +what has not been tried yet? + +### Phase 2 — Ideate + +Pick ONE hypothesis. It must be: +- Specific and testable in a single iteration +- Meaningfully different from the last 3 attempts +- Informed by both the results log and the lessons file +- Explained in one sentence + +Prefer hypotheses that build on proven wins over untested territory. +Prefer simplicity — a small clean change beats a large complex one. + +### Phase 3 — Modify + +Make exactly ONE atomic change in Scope. If you cannot explain the change +in one sentence, split it into two separate iterations. + +Do not touch files outside Scope. Do not refactor unrelated code. One thing. + +### Phase 4 — Commit + +```bash +git add -A && git commit -m "autoresearch iter N: " +``` + +**Commit BEFORE verifying.** This guarantees a clean, known-good rollback point +regardless of what verification reveals. Never skip this step. + +### Phase 5 — Verify + Guard + +**Step A — Run Verify.** Extract the numeric metric value. + +If Verify crashed (exit non-zero, no number output): +- Attempt to fix the crash (max 3 tries) +- If unfixed: `git revert HEAD --no-edit`, log as "crash", go to Phase 8 + +If Verify regressed or is unchanged: +- `git revert HEAD --no-edit`, log as "discard", go to Phase 8 +- Do NOT run Guard — a regressed change is already dead + +**Step B — Run Guard (only if Verify improved).** Exit code 0 = pass. + +**Web research supplement**: after Verify passes, use `WebSearch` for +additional signal when local scripts cannot capture full quality. +See `references/web-research-patterns.md`. Research is a supplement only. + +### Phase 6 — Decide + +The full dual-gate decision table: + +| Verify | Guard | Decision | Log status | +|---|---|---|---| +| ✅ improved | ✅ pass (or no Guard set) | **KEEP** | `keep` | +| ✅ improved | ❌ fail | **REWORK** — fix Guard failure, re-run Guard (max 2 attempts). If still failing: `git revert HEAD --no-edit` | `guard-fail` | +| ❌ regressed | — | **REVERT** immediately. Do not run Guard. | `discard` | +| ❌ unchanged | — | **REVERT**. Treat unchanged as a regression. | `discard` | +| 💥 crashed | — | **FIX** (max 3 attempts), then revert if unfixed. | `crash` | + +**Rework protocol** (when Verify passes but Guard fails): +1. Read the Guard failure output carefully +2. Make the minimal additional change to satisfy Guard without hurting Verify +3. Amend the commit: `git add -A && git commit --amend --no-edit` +4. Re-run both Verify AND Guard +5. If both pass → KEEP. If Guard still fails after 2 rework attempts → REVERT. + +### Phase 7 — Log + +Append one row to `autoresearch-results.tsv`: + +``` +\t\t\t\t\t\t +``` + +Delta = metric_value − previous_best (positive = improvement for "higher is +better" goals, negative = improvement for "lower is better" goals). + +### Phase 8 — Repeat + +Go to Phase 1. Immediately. NEVER STOP. + +--- + +## Progress summary (every 10 iterations) + +Print this, then continue immediately: + +``` +=== Autoresearch progress — iteration N === +Baseline: +Current best: ( from baseline) +Keeps: +Discards: +Crashes: +Top pattern: +Last 5: +=== +``` + +--- + +## Lessons system + +After every 5 KEPT iterations, append to `autoresearch-lessons.md`: + +```markdown +## Lesson — iterations +**Pattern**: +**Why it worked**: +**Conditions**: +**Anti-pattern**: +**Metric delta**: +``` + +At the start of every run, read this file before forming any hypotheses. +Weight recent lessons more heavily. Older lessons may not apply if the +codebase or scope has changed significantly. + +This is the compounding mechanism. Each overnight run starts smarter than +the last. + +--- + +## Stuck recovery + +After 5 consecutive discards or crashes: + +1. Re-read all Scope files from scratch. Full context, not memory. +2. Search the lessons log for near-misses — what came closest to working? +3. Try combining two near-miss approaches into one hypothesis. +4. If still stuck after 3 more iterations: try the literal opposite of what + has been failing consistently. +5. If still stuck after 3 more: use `WebSearch` to research the + problem space. Search for `[domain] [metric] improvement techniques [year]`. + Extract 3 concrete techniques. Use each as the next 3 hypotheses. +6. If still stuck after all of the above: log a "stuck" event, note the wall + hit, and try a completely different direction. Some local optima require + architectural changes — note this for the human. + +--- + +## Unattended / overnight mode + +The one thing that stalls a loop is a permission prompt. Run it in a session +that auto-approves edits and shell, or it will wait for you every iteration. + +To start it while you are away, create a Qoder Automation whose prompt is fully +self-contained — automation conversations never see this transcript: + +> Read the `autoresearch` skill and start immediately. Goal: ``. +> Scope: ``. Metric: ``. +> Verify: ``. Guard: ``. Do not pause, do not ask questions, +> iterate until stopped. + +You will wake up to `autoresearch-results.tsv` and `autoresearch-lessons.md`. +Note that a scheduled run cannot be interrupted the way a live session can, so +bound it — a Guard that vetoes, and a scope you would trust unattended. + +--- + +## Non-negotiable rules + +1. **NEVER STOP** until the user manually interrupts the run. +2. **ONE change per iteration** — atomic, explainable in one sentence. +3. **Mechanical verification only** — no "looks better", no "seems cleaner". + If you cannot measure it, you cannot use it as a signal. +4. **Commit BEFORE verifying** — always. No exceptions. +5. **Auto-revert on regression** — no debate, no "let me try one more thing". +6. **Guard is a hard veto** — Verify passing does not mean KEEP. Guard must also pass. +7. **Never modify Guard files** — they are read-only invariants, not scope. +8. **Read git history before every hypothesis** — it is your short-term memory. +9. **Read lessons before every run** — it is your long-term memory. +10. **Simplicity wins ties** — equal metric + less code = KEEP. +11. **Never touch files outside Scope** — discipline is what makes the loop safe. +12. **When in doubt, make the smaller change** — scope creep kills iterations. + +--- + +## Reference files + +**Core loop** +- `references/loop-protocol.md` — detailed phase-by-phase protocol +- `references/results-logging.md` — TSV format, summary templates, examples +- `references/lessons-system.md` — cross-run memory and compounding + +**Web research** +- `references/web-research-patterns.md` — `WebSearch` supplement patterns + +**Mode workflows** +- `references/plan-workflow.md` — `plan` mode — auto-detect and configure +- `references/ship-workflow.md` — `ship` mode — pre-flight checklist +- `references/debug-workflow.md` — `debug` mode — root cause and fix +- `references/fix-workflow.md` — `fix` mode — focused type/lint fix +- `references/security-workflow.md` — `security` mode — STRIDE/OWASP audit diff --git a/.agents/skills/autoresearch/references/debug-workflow.md b/.agents/skills/autoresearch/references/debug-workflow.md new file mode 100644 index 00000000..27738ebf --- /dev/null +++ b/.agents/skills/autoresearch/references/debug-workflow.md @@ -0,0 +1,25 @@ +# `autoresearch debug` mode — Autonomous Debug Loop + +This workflow is triggered by the `debug` mode. It is designed to reproduce, isolate, and fix specific bugs autonomously. + +## Context +Use this when something is clearly broken (e.g., a failing test, a crash, or a UI bug). + +## Phase 1: Reproduction +1. Create a minimal reproduction script (e.g., `debug/repro.js` or a new test case). +2. Run the repro script and verify it fails as expected. +3. This repro command becomes your `Verify` command for the loop. + +## Phase 2: Isolation +1. Use `Grep` and `Read` to find the code responsible for the failure. +2. Form a hypothesis about the root cause. + +## Phase 3: Fix Loop +1. Start a standard autoresearch loop with: + - **Goal**: Fix the bug identified in the repro script. + - **Verify**: The repro command (must exit 0 on success). + - **Guard**: Existing test suite and linting. + +## Phase 4: Hardening +1. After the fix is verified, add a permanent regression test to the codebase. +2. Verify that the fix holds across the entire project. diff --git a/.agents/skills/autoresearch/references/fix-workflow.md b/.agents/skills/autoresearch/references/fix-workflow.md new file mode 100644 index 00000000..04303826 --- /dev/null +++ b/.agents/skills/autoresearch/references/fix-workflow.md @@ -0,0 +1,31 @@ +# Fix Workflow (`autoresearch fix` mode) + +The `fix` workflow is a lightweight version of the `debug` loop. It is designed for situations where you have a specific, known failure (e.g., a TypeScript error or a lint violation) and you want to fix it without the overhead of full reproduction and isolation. + +## Protocol + +### 1. Context Loading +* Read the error message or description provided in the command. +* Identify the affected file(s). +* Read the current state of those files. + +### 2. Hypothesis +* Form a direct hypothesis on how to fix the specific error. +* The fix must be minimal and targeted. + +### 3. Execution +* Apply the fix. +* Commit the change. + +### 4. Verification +* Run the command that triggered the original failure (e.g., `npx tsc` or `npm run lint`). +* If a `Guard` is set in the main autoresearch config, run that as well. + +### 5. Decision +* If the error is gone and Guard passes: **KEEP**. +* If the error persists: **RETRY** (max 3 times) with a different approach. +* If it still fails after 3 tries: **REVERT** and report to the user. + +## When to use `fix` vs `debug` +* Use **`fix`** for mechanical errors: "Fix the lint error on line 42", "Fix the missing import in `utils.ts`". +* Use **`debug`** for logical errors: "The login flow fails for users with specialized characters", "Database connection timeouts under high load". diff --git a/.agents/skills/autoresearch/references/lessons-system.md b/.agents/skills/autoresearch/references/lessons-system.md new file mode 100644 index 00000000..7eb15dd9 --- /dev/null +++ b/.agents/skills/autoresearch/references/lessons-system.md @@ -0,0 +1,117 @@ +# Lessons system + +The lessons system is what separates autoresearch from a dumb +mutation loop. It is the mechanism by which each overnight run starts +smarter than the last. + +--- + +## The compounding model + +``` +Night 1: 100 experiments → lessons-v1 written +Night 2: reads lessons-v1 → avoids 20 known failures → 80 net-new experiments +Night 3: reads lessons-v2 → avoids 35 known failures → faster convergence +... +``` + +Without the lessons system, every run starts from scratch. With it, runs +compound — each failure is learned once and never repeated. + +--- + +## File location and format + +File: `autoresearch-lessons.md` in your project root. + +Add to `.gitignore` — this is a working file for the agent, not source code. + +```markdown +# Autoresearch lessons — +Generated by the autoresearch skill. Do not edit manually. +Last updated: + +## Lesson 1 — iterations 1–5 +**Pattern**: +**Why it worked**: +**Conditions**: +**Anti-pattern**: +**Metric delta**: + +## Lesson 2 — iterations 6–10 +... +``` + +--- + +## When to write lessons + +Append a new lesson after every 5 KEPT iterations (not every 5 total +iterations). Lessons should only describe what worked. + +Failed patterns are captured implicitly — if a pattern never generates a +kept iteration, it never generates a lesson, and the loop naturally +deprioritises it via Phase 2's "different from last 3 attempts" rule. + +--- + +## What makes a good lesson + +**Good** (specific, mechanistic, conditional): +``` +**Pattern**: Defer non-critical third-party scripts using loading="lazy" +**Why it worked**: Removes scripts from the critical render path, reducing + Time to Interactive without affecting functionality +**Conditions**: Applies to analytics, chat widgets, social embeds — not + to scripts required for initial page render +**Anti-pattern**: Lazy-loading scripts that are called in the first 500ms + of page load caused layout shifts and broke interactions +**Metric delta**: +6.8% Lighthouse performance score across 3 iterations +``` + +**Bad** (vague, not actionable): +``` +**Pattern**: Make things faster +**Why it worked**: It improved performance +**Conditions**: When performance is bad +**Anti-pattern**: When it makes things worse +``` + +--- + +## How to read lessons at the start of a run + +1. Read the full file — do not skip old lessons even if they seem stale. +2. For each lesson, assess: does this pattern still apply given the current + state of the codebase? If the code it describes has been significantly + refactored, downweight it. +3. Extract the top 2-3 highest-delta patterns. These are your first + hypotheses unless the results log shows they have already been exhausted. +4. Extract the anti-patterns. These are your first exclusions — do not + generate hypotheses that match these patterns. + +--- + +## Cross-project lessons + +For teams running autoresearch across multiple similar projects (e.g. +multiple Next.js apps), consider maintaining a shared lessons file at +`~/.autoresearch/global-lessons.md`. + +At the start of a run, read both the project-level and global lessons. +Project-level lessons take precedence when they conflict with global ones. + +This is optional but significantly accelerates convergence on new projects +that share a tech stack with already-researched ones. + +--- + +## Lessons file maintenance + +- Do not manually edit the lessons file during a run — the agent reads it + at the start of each run and its contents influence hypothesis generation. +- After a long run (100+ iterations), review the file and remove lessons + that are no longer applicable (e.g. they describe code that no longer + exists). Add a comment explaining why the lesson was removed. +- The lessons file is cumulative — never delete lessons, only annotate them + as superseded if a newer lesson contradicts them. diff --git a/.agents/skills/autoresearch/references/loop-protocol.md b/.agents/skills/autoresearch/references/loop-protocol.md new file mode 100644 index 00000000..60685c9f --- /dev/null +++ b/.agents/skills/autoresearch/references/loop-protocol.md @@ -0,0 +1,193 @@ +# Autonomous loop protocol + +Detailed specification for each of the 8 phases. The SKILL.md contains the +summary version. Read this reference when you need precise guidance on edge +cases in any phase. + +--- + +## Phase 1 — Review + +**Purpose**: Build a complete, accurate picture of current state before +forming any hypothesis. Hypotheses formed without full context waste iterations. + +**What to read**: +- Every file in Scope (not just the ones you last touched) +- `git log --oneline -20` — what has been attempted, in order +- `autoresearch-results.tsv` — the full record of what worked and failed +- `autoresearch-lessons.md` — accumulated patterns from prior runs + +**What to extract**: +- Current metric trajectory (improving? plateauing? volatile?) +- Which change types produced the most gain per iteration +- Which change types consistently failed +- Which directions have not yet been explored +- Any patterns in crash causes + +**Duration**: This phase should take as long as needed to form a genuinely +informed hypothesis. Rushing Phase 1 leads to repeated failures. + +--- + +## Phase 2 — Ideate + +**Purpose**: Select ONE hypothesis that has the highest expected gain given +what is known. + +**Hypothesis selection criteria** (in order of priority): +1. Builds directly on a proven pattern from the lessons file +2. Explores a direction adjacent to a near-miss (something that almost worked) +3. Combines two near-miss approaches that individually failed +4. Tries the opposite of what consistently failed +5. Applies an externally validated technique (from `WebSearch` research) +6. Tries something entirely untested + +**What makes a good hypothesis**: +- Specific: "lazy-load the user avatar component" not "improve performance" +- Testable: produces a measurable delta in the Verify command +- Atomic: one thing changes, one thing is measured +- Explainable in one sentence before you make the change + +**What makes a bad hypothesis**: +- Vague: "refactor for clarity" +- Multi-part: "update the API, add caching, and fix the tests" +- Untestable by the Verify command +- Identical to something tried in the last 3 iterations + +--- + +## Phase 3 — Modify + +**Purpose**: Implement the hypothesis as a single, clean, minimal change. + +**Rules**: +- Touch only files in Scope +- Make the smallest change that tests the hypothesis +- If the change is getting large, stop and split it — make the first half now, + the second half in the next iteration +- Do not fix unrelated things you notice while editing +- Do not reformat code that is not part of the hypothesis +- Leave comments only if they directly explain the change + +**Signs you are over-scoping**: +- You have edited more than 3 files +- The diff is more than ~50 lines +- You are explaining the change with "and also" + +When in doubt, make a smaller change. Smaller changes fail faster and teach more. + +--- + +## Phase 4 — Commit + +**Purpose**: Create a clean rollback point before any verification risk. + +**Command**: +```bash +git add -A && git commit -m "autoresearch iter N: " +``` + +**Commit message format**: +- Always prefix with `autoresearch iter N:` +- One sentence, present tense, describes the change not the goal +- Good: `autoresearch iter 14: lazy-load user avatar to reduce initial bundle` +- Bad: `autoresearch iter 14: improve performance` + +**Why commit before verifying**: if the Verify command crashes, hangs, or +corrupts state, you can always `git revert HEAD --no-edit` and return to +a known-good state. If you verify before committing, a crash during +verification leaves you with uncommitted changes and an unknown baseline. + +**Never skip this step**, even if the change feels obviously correct. + +--- + +## Phase 5 — Verify + +**Purpose**: Get a single numeric measurement of whether the hypothesis helped. + +**Execution**: +1. Run the Verify command exactly as specified by the user +2. Extract the numeric metric value +3. Optionally supplement with `WebSearch` research (see + `references/web-research-patterns.md`) +4. Record the raw output for the log + +**Handling slow Verify commands**: +If the Verify command takes more than 30 seconds, note this. After the run, +recommend the user find a faster proxy metric — slower verification means +fewer experiments per hour, which compounds negatively over a full night. + +**Handling non-deterministic Verify commands**: +If the metric varies significantly between runs on identical code (>5% +variance), note this in the log. Run the Verify command twice and average. +Log both values. Recommend the user address flakiness before the next +overnight run. + +--- + +## Phase 6 — Decide + +**Purpose**: Make a clear, mechanical keep/revert decision. No deliberation. + +**Decision table**: + +| Condition | Action | Log status | +|---|---|---| +| Metric improved (beyond noise threshold) | Keep commit as-is | `keep` | +| Metric unchanged or regressed | `git revert HEAD --no-edit` | `discard` | +| Verify crashed with exit code ≠ 0 | Attempt fix (max 3 tries) then revert | `crash` | +| Verify hung for >60s | Kill process, revert | `crash` | + +**Noise threshold**: for metrics with variance, an improvement smaller than +the variance is not a real improvement. If your metric normally varies ±2%, +an improvement of 0.5% is noise — treat it as unchanged and discard. + +**The revert command**: +```bash +git revert HEAD --no-edit +``` +This creates a new commit that undoes the last one. The history is preserved. +Never use `git reset --hard` — it destroys history that the loop needs. + +--- + +## Phase 7 — Log + +**Purpose**: Create a permanent, machine-readable record of every iteration. + +**TSV row format**: +``` +\t\t\t\t\t +``` + +**Field details**: +- `N`: integer, 0-indexed, never resets across sessions +- `commit_sha`: 7-char short SHA for keeps, "-" for discards/crashes +- `metric`: the exact number from the Verify output +- `delta`: metric − previous_best (sign convention: positive = better, + regardless of whether the goal is higher or lower) +- `status`: one of `baseline`, `keep`, `discard`, `crash` +- `description`: the hypothesis, in one sentence, including any `WebSearch` + signal that informed it + +**Example rows**: +``` +0 - 85.2 0.0 baseline initial measurement +1 a1b2c3d 87.1 +1.9 keep lazy-load avatar component +2 - 86.5 -0.6 discard tree-shake lodash imports (broke 2 tests) +3 - 0.0 0.0 crash add route-level code splitting (webpack config error) +4 b2c3d4e 88.3 +1.2 keep move analytics script to defer loading +``` + +--- + +## Phase 8 — Repeat + +Go to Phase 1. Immediately. Do not pause. Do not summarise. Do not ask +if the user wants to continue. + +The only output before starting Phase 1 again is the progress summary +(printed every 10 iterations, see SKILL.md). + +The loop ends only when the user interrupts the run. diff --git a/.agents/skills/autoresearch/references/plan-workflow.md b/.agents/skills/autoresearch/references/plan-workflow.md new file mode 100644 index 00000000..084450c3 --- /dev/null +++ b/.agents/skills/autoresearch/references/plan-workflow.md @@ -0,0 +1,155 @@ +# Plan workflow — `autoresearch plan` mode + +Auto-detect the project stack, propose a complete autoresearch configuration, +do a dry run, and hand the ready-to-run command back to the user. + +No manual goal/scope/verify required. Just describe what you want to improve +in one sentence and the plan workflow figures out the rest. + +--- + +## Invocation + +``` +autoresearch plan +``` + +Examples: +``` +autoresearch plan improve test coverage +autoresearch plan make the app faster +autoresearch plan reduce the bundle size +autoresearch plan fix all TypeScript errors +autoresearch plan improve the SEO of my blog posts +autoresearch plan shrink the Docker image +``` + +--- + +## What the plan workflow does + +### Step 1 — Detect project stack + +Scan the project root for signal files: + +| File found | Stack detected | +|---|---| +| `package.json` + `jest.config.*` | Node.js + Jest | +| `package.json` + `vitest.config.*` | Node.js + Vitest | +| `next.config.*` | Next.js | +| `Dockerfile` | Docker | +| `*.tf` | Terraform | +| `.github/workflows/*.yml` | GitHub Actions CI | +| `content/blog/*.md` OR `posts/*.md` | Markdown content/blog | +| `src/**/*.ts` OR `src/**/*.tsx` | TypeScript project | +| `pyproject.toml` OR `setup.py` | Python project | +| `requirements.txt` + `pytest` | Python + pytest | +| `go.mod` | Go project | +| `Cargo.toml` | Rust project | + +Print detected stack. If ambiguous, list the top two candidates and ask +the user to confirm before proceeding. + +### Step 2 — Map goal to metric + verify command + +Use the goal description and detected stack to propose: + +| Goal keyword | Metric | Verify command template | +|---|---|---| +| "test coverage" | coverage % (higher is better) | `npm test -- --coverage \| grep "All files"` | +| "bundle size" / "build size" | size in KB (lower is better) | `npm run build 2>&1 \| grep "First Load JS"` | +| "TypeScript errors" / "type errors" | error count (lower is better) | `npx tsc --noEmit 2>&1 \| grep -c "error TS" \|\| echo "0"` | +| "lighthouse" / "performance score" | score 0-100 (higher is better) | `npx lighthouse http://localhost:3000 --output json --quiet 2>/dev/null \| jq '.categories.performance.score * 100'` | +| "docker image" / "image size" | size in MB (lower is better) | `docker build -t bench . -q && docker images bench --format "{{.Size}}"` | +| "flaky tests" | failure count (lower is better) | `for i in {1..5}; do npm test 2>&1; done \| grep -c "FAIL" \|\| echo "0"` | +| "SEO" / "blog" / "content" | SEO score (higher is better) | `node scripts/seo-score.js ` | +| "lines of code" / "complexity" | LOC count (lower is better) | `find src/ -name "*.ts" \| xargs wc -l \| tail -1 \| awk '{print $1}'` | +| "CI pipeline" / "pipeline speed" | seconds (lower is better) | `node scripts/estimate-ci-time.js` | +| "Python tests" / "pytest" | coverage % (higher is better) | `pytest --cov=src --cov-report=term-missing \| grep "TOTAL"` | +| "faster" / "performance" / "latency" | p95 ms (lower is better) | `npm run bench 2>&1 \| grep "p95"` | + +### Step 3 — Detect scope + +Based on goal + stack, propose the tightest scope that covers the goal: + +- Test coverage → `src/**/*.ts, src/**/*.test.ts` +- Bundle size → `src/**/*.tsx, src/**/*.ts` +- Docker → `Dockerfile, .dockerignore` +- SEO → `content/blog/*.md` or detected content directory +- TypeScript errors → `src/**/*.ts` +- CI pipeline → `.github/workflows/*.yml` + +### Step 4 — Dry run + +Run the proposed Verify command once against the current state. + +- If it exits 0 and outputs a number → baseline confirmed, proceed +- If it exits non-zero → diagnose and fix the verify command before proposing +- If it hangs → propose a faster alternative + +### Step 5 — Output the ready-to-run command + +Print this exact block for the user to copy-paste or confirm: + +``` +=== Autoresearch plan === +Stack: +Goal: +Scope: +Metric: ( is better) +Verify: +Baseline: + +Ready to run. Confirm or adjust any field, then: + +/autoresearch +Goal: +Scope: +Metric: +Verify: + +Or, for an unattended run, put these same fields into a Qoder Automation prompt +(see "Unattended / overnight mode" in SKILL.md). +=== +``` + +If the user says "looks good" or "run it" — start the autoresearch loop +immediately without requiring them to retype the command. + +--- + +## Web research calibration + +After the dry run, use `WebSearch` to calibrate: + +- For SEO goals: search for `[target keyword]` to see what top results look like. + Note any structural patterns (FAQ sections, word count, heading structure) + that the current content lacks. Add these as initial hypotheses. + +- For performance goals: search for `[framework] performance benchmarks [year]` + to calibrate whether the baseline is already good or has significant headroom. + +- For security goals: search for `[stack] common vulnerabilities [year]` + to seed the initial hypothesis pool with known attack vectors. + +This research step happens during plan, not during the loop — so it adds +context once without slowing down iterations. + +--- + +## Edge cases + +**Goal is too vague** ("make it better"): +Ask one clarifying question: "Better in what way — speed, quality, size, +coverage, or something else?" Then proceed. + +**Multiple valid verify commands exist**: +Propose the fastest one. Note the slower alternative in a comment. + +**Verify command requires a running server**: +Note this in the plan output. Add a `# requires: local server on :3000` +comment. Suggest the user start it before running the loop. + +**No matching stack detected**: +Ask the user to describe their stack in one sentence, then proceed with +a custom verify command. diff --git a/.agents/skills/autoresearch/references/results-logging.md b/.agents/skills/autoresearch/references/results-logging.md new file mode 100644 index 00000000..63c2d8a5 --- /dev/null +++ b/.agents/skills/autoresearch/references/results-logging.md @@ -0,0 +1,105 @@ +# Results logging + +Specification for `autoresearch-results.tsv` — the per-iteration record +of every experiment in a run. + +--- + +## File format + +Tab-separated values. Headers on row 1. One row per iteration. + +``` +iteration\tcommit\tmetric\tdelta\tstatus\tdescription +``` + +### Field definitions + +| Field | Type | Description | +|---|---|---| +| `iteration` | integer | 0-indexed. Never resets — if you run multiple sessions, continue from the last number. | +| `commit` | string | 7-char git short SHA for kept commits. `-` for discards and crashes. | +| `metric` | float | Raw metric value from the Verify command. | +| `delta` | float | `metric − previous_best`. Sign convention: positive = improvement (regardless of higher/lower goal). | +| `status` | enum | One of: `baseline`, `keep`, `discard`, `crash` | +| `description` | string | The hypothesis, one sentence. Include the change type and the expected mechanism. | + +--- + +## Example file + +```tsv +iteration commit metric delta status description +0 - 85.2 0.0 baseline initial measurement — test coverage 85.2% +1 a1b2c3d 87.1 +1.9 keep add tests for auth middleware edge cases +2 - 86.5 -0.7 discard refactor test helpers (broke 2 existing tests) +3 - 0.0 0.0 crash add integration tests (postgres connection failed — fix in iter 4) +4 b2c3d4e 88.3 +1.2 keep add tests for error handling in API routes +5 - 88.1 -0.2 discard add tests for rate limiter (metric within variance, treated as regression) +6 c3d4e5f 89.0 +0.7 keep add boundary value tests for form validators +7 d4e5f6g 89.8 +0.8 keep add tests for session expiry edge cases +8 - 89.2 -0.6 discard mock external API calls (test isolation but metric regressed) +9 e5f6g7h 90.6 +0.8 keep add tests for concurrent request handling +10 f6g7h8i 91.1 +0.5 keep add tests for malformed JSON input handling +``` + +--- + +## Progress summary format + +Print every 10 iterations. Use this exact format: + +``` +=== Autoresearch progress — iteration === +Goal: +Baseline: +Current best: ( from baseline) +Keeps: (%) +Discards: +Crashes: +Top pattern: +Last 5: +Est. to goal: +=== +``` + +--- + +## Interpreting the log + +### Healthy run signature +- Keep rate 40-60% +- Delta per keep: consistent small positive gains +- No long crash streaks +- Discards are evenly distributed (not clustered) + +### Warning signs + +| Pattern | Meaning | Action | +|---|---|---| +| Keep rate < 20% | Hypothesis quality is poor | Re-read full scope, re-read lessons, change direction | +| Keep rate > 80% | Metric may be too easy or Verify too lenient | Tighten the goal | +| Long crash streak (5+) | Verify command is fragile or scope is too risky | Fix Verify or narrow scope | +| Delta per keep shrinking toward 0 | Approaching local optimum | Try more radical changes or declare victory | +| Metric oscillating | Non-deterministic Verify or contradictory changes | Run Verify twice and average; tighten scope | + +### Declaring success + +Stop the loop when one of these is true: +- Metric has reached the stated goal +- Delta per keep has been below 0.1% for 20 consecutive iterations + (local optimum with current scope) +- All directions have been exhausted (lessons file confirms this) + +In all cases, print a final summary and write a lessons entry covering +the full run before stopping. + +--- + +## File hygiene + +- Add `autoresearch-results.tsv` to `.gitignore`. It is a working file. +- Do not edit it manually during a run. +- Between runs, you may archive it: + `mv autoresearch-results.tsv autoresearch-results-.tsv` + and start fresh, but keep the lessons file — that is the persistent memory. diff --git a/.agents/skills/autoresearch/references/security-workflow.md b/.agents/skills/autoresearch/references/security-workflow.md new file mode 100644 index 00000000..b3ccd3e1 --- /dev/null +++ b/.agents/skills/autoresearch/references/security-workflow.md @@ -0,0 +1,171 @@ +# Security workflow — `autoresearch security` mode + +Autonomous security audit using STRIDE threat modelling and OWASP categories. +Finds vulnerabilities, classifies them by severity, and optionally fixes +confirmed critical and high findings via an autoresearch loop. + +--- + +## Invocation + +``` +autoresearch security # full audit, report only +autoresearch security --fix # audit + auto-fix confirmed findings +autoresearch security --fail-on critical # end with a FAIL verdict if critical found +autoresearch security --scope src/api/ # audit a specific directory only +``` + +--- + +## Phase 1 — Asset discovery + +Map the attack surface: + +1. Identify all entry points: API routes, form handlers, file uploads, + auth flows, webhooks, admin panels +2. Identify all data stores: databases, caches, file system writes, + environment variables, secrets +3. Identify all trust boundaries: public vs authenticated, user vs admin, + internal vs external services +4. Map data flows: what user input reaches what data store via what path + +Output: `security/audit-/attack-surface-map.md` + +### Live threat intelligence + +Use `WebSearch` to seed the audit with current threats: + +``` +WebSearch: [your stack] common vulnerabilities [current year] +WebSearch: [your main framework] CVE [current year] +WebSearch: OWASP top 10 [current year] +``` + +Add any newly discovered attack patterns to the audit queue. +This ensures the audit covers threats that postdate your static analysis tools. + +--- + +## Phase 2 — STRIDE threat model + +For each asset and trust boundary, model threats across all 6 STRIDE categories: + +| Category | Question to ask | +|---|---| +| **S**poofing | Can an attacker impersonate a user, service, or system? | +| **T**ampering | Can input be modified to alter data or behaviour unexpectedly? | +| **R**epudiation | Can actions be performed without a traceable audit trail? | +| **I**nformation disclosure | Can sensitive data be accessed by unauthorised parties? | +| **D**enial of service | Can the service be made unavailable through normal inputs? | +| **E**levation of privilege | Can a lower-privilege user gain higher-privilege access? | + +Output: `security/audit-/threat-model.md` + +--- + +## Phase 3 — Autonomous audit loop + +``` +LOOP (through all attack vectors from threat model): + 1. Select next untested attack vector + 2. Deep-dive into the relevant code (read fully — do not skim) + 3. Attempt to construct a concrete exploit scenario + 4. Validate with code evidence (file:line + exact scenario) + 5. Classify: severity + OWASP category + STRIDE tag + 6. Log to security-audit-results.tsv + 7. Print coverage summary every 5 iterations + 8. Continue until all vectors tested +``` + +### Severity classification + +| Severity | Definition | +|---|---| +| Critical | Exploitable without authentication, leads to full compromise or data breach | +| High | Exploitable with low-privilege access, significant impact | +| Medium | Requires specific conditions, moderate impact | +| Low | Minor information disclosure, no direct exploitation path | +| Info | Best practice violation, no immediate security impact | + +### Evidence requirement + +Every finding MUST have: +- File path and line number +- Exact vulnerable code snippet (copy from source, do not paraphrase) +- Concrete exploit scenario (how an attacker would trigger this) +- Proof of exploitability (not theoretical — show the actual path) + +Findings without concrete evidence are logged as "unconfirmed" and flagged +for manual review, not included in the fix loop. + +--- + +## Phase 4 — Report generation + +Output folder: `security/audit-/` + +``` +security/audit-20260325-1430/ +├── overview.md ← executive summary + finding counts by severity +├── threat-model.md ← STRIDE analysis per asset +├── attack-surface-map.md ← entry points, data flows, trust boundaries +├── findings.md ← all confirmed findings, sorted by severity +├── owasp-coverage.md ← coverage matrix — which OWASP categories checked +├── recommendations.md ← fix guidance for each confirmed finding +└── security-audit-results.tsv ← machine-readable log of all iterations +``` + +Print summary: +``` +=== Security audit summary === +Critical: +High: +Medium: +Low: +Info: +Vectors tested: / +OWASP categories covered: + +Full report: security/audit-/overview.md +=== +``` + +--- + +## Phase 5 — Auto-fix loop (with `--fix`) + +Only runs when `--fix` flag is passed. +Only fixes **Confirmed Critical and High** findings. +Uses `recommendations.md` as the fix guide for each finding. + +``` +FOR EACH confirmed Critical/High finding: + 1. Read the finding + recommendation + 2. Make ONE targeted fix + 3. git commit the fix + 4. Re-run the specific exploit scenario to verify it no longer works + 5. Run full test suite to confirm no regressions + 6. If tests break → revert, try alternative fix + 7. Maximum 3 attempts per finding, then skip and flag for manual review + 8. Log fix outcome to fix-log.md +``` + +--- + +## Verdict mode (`--fail-on`) + +``` +autoresearch security --fail-on critical +``` + +The audit ends with an explicit verdict line in `overview.md`: + +``` +VERDICT: FAIL — 2 findings at or above `critical` +VERDICT: PASS — no findings at or above `critical` +``` + +A skill run has no process exit code, so do not wire this into a CI gate as if +it did — use a real scanner for blocking merges. What it *is* good for is an +unattended scheduled audit: a Qoder Automation running this mode reports the +verdict, and you act on it. diff --git a/.agents/skills/autoresearch/references/ship-workflow.md b/.agents/skills/autoresearch/references/ship-workflow.md new file mode 100644 index 00000000..66c36a42 --- /dev/null +++ b/.agents/skills/autoresearch/references/ship-workflow.md @@ -0,0 +1,164 @@ +# Ship workflow — `autoresearch ship` + +Run a pre-flight checklist before shipping — tests, types, lint, bundle size, +security basics, and a final autoresearch pass on anything that fails. + +The ship workflow is not just a checklist. It runs an autoresearch loop on +each failing gate until it passes, then re-checks. You don't ship broken. +You ship when everything is green. + +--- + +## Invocation + +``` +autoresearch ship +``` + +Optional flags: +``` +autoresearch ship --fast # skip slow checks (lighthouse, e2e) +autoresearch ship --loop N # max N autoresearch iterations per gate (default: 20) +autoresearch ship --dry-run # report status without fixing anything +``` + +--- + +## The ship checklist + +The workflow runs these gates in order. Each gate that fails triggers an +autoresearch sub-loop to fix it before moving to the next gate. + +### Gate 1 — Tests pass + +```bash +npm test # Node.js +pytest # Python +go test ./... # Go +cargo test # Rust +``` + +If tests fail → autoresearch loop on `src/**/*.ts` (or equivalent) with +metric: failing test count (lower is better), max 20 iterations. + +### Gate 2 — No type errors + +```bash +npx tsc --noEmit # TypeScript +mypy src/ # Python +``` + +If errors found → autoresearch loop on `src/**/*.ts` with +metric: error count (lower is better), max 20 iterations. + +### Gate 3 — No lint errors + +```bash +npx eslint src/ # JavaScript/TypeScript +ruff check src/ # Python +golangci-lint run # Go +``` + +If errors found → autoresearch loop with metric: lint error count (lower is better). +Auto-fixable errors are fixed first (`--fix` flag), then the loop handles the rest. + +### Gate 4 — Bundle size (if applicable) + +Only runs for frontend projects (detected: `next.config.*`, `vite.config.*`, +`webpack.config.*`). + +```bash +npm run build 2>&1 | grep "First Load JS" +``` + +Threshold: warn if > 300KB, block if > 500KB (configurable via `.autoresearch.yml`). + +If over threshold → autoresearch loop on `src/**/*.tsx, src/**/*.ts` with +metric: bundle size in KB (lower is better), max 20 iterations. + +### Gate 5 — No hardcoded secrets + +```bash +git diff HEAD~1 --diff-filter=A | grep -iE "(api_key|secret|password|token)\s*=\s*['\"][^'\"]{8,}" +``` + +If secrets found → do NOT autoresearch. Flag for human review. Block ship. + +### Gate 6 — Dependency audit + +```bash +npm audit --audit-level=high # Node.js +pip-audit # Python +``` + +If critical vulnerabilities found → autoresearch loop to update affected +dependencies, max 10 iterations. + +--- + +## Ship report + +After all gates pass, print: + +``` +=== Ship report === +Tests: ✓ PASS (247 passing) +Types: ✓ PASS (0 errors) +Lint: ✓ PASS (0 errors) +Bundle: ✓ PASS (187KB) +Secrets: ✓ PASS (none detected) +Deps: ✓ PASS (0 high/critical) + +Autoresearch loops run: +Total improvements: iterations kept + +Ready to ship. Run: git push && +=== +``` + +If any gate is still failing after the max iterations: + +``` +=== Ship report === +Tests: ✓ PASS +Types: ✗ FAIL (3 errors remaining after 20 iterations) + → manual fix required: src/auth/session.ts:47 + +Ship BLOCKED. Fix the above before shipping. +=== +``` + +--- + +## Web research post-check + +After all gates pass, use `WebSearch` to check: + +``` +WebSearch: [your framework] [version] known issues [current year] +WebSearch: [your main dependencies] security advisory [current year] +``` + +If any critical advisories surface that the dependency audit missed, +flag them before shipping. This is a final sanity check that goes beyond +what local tools can detect. + +--- + +## Configuration via `.autoresearch.yml` + +Create this file in your project root to customise ship behaviour: + +```yaml +ship: + bundle_warn_kb: 300 + bundle_block_kb: 500 + max_iterations_per_gate: 20 + skip_gates: + - lighthouse # skip if no local server available + extra_gates: + - name: "E2E tests" + command: "npx playwright test" + metric: "failing tests (lower is better)" + max_iterations: 10 +``` diff --git a/.agents/skills/autoresearch/references/web-research-patterns.md b/.agents/skills/autoresearch/references/web-research-patterns.md new file mode 100644 index 00000000..e2dedfc0 --- /dev/null +++ b/.agents/skills/autoresearch/references/web-research-patterns.md @@ -0,0 +1,144 @@ +# Web research patterns + +Qoder exposes a `WebSearch` tool (and `WebFetch` to read a promising result in +full). Use them as a verification supplement — not a replacement for the Verify +command, but an additional signal when local scripts alone cannot capture +quality. + +--- + +## When to use WebSearch in the loop + +| Goal type | Use WebSearch for | Example query | +|---|---|---| +| SEO content | Check competing pages, keyword signals | `[target keyword] filetype:md OR site:*.dev` | +| API correctness | Verify endpoint signatures, check for deprecations | `[library] [method] deprecated 2025 OR 2026` | +| Dependency versions | Confirm latest stable before updating | `[package name] latest stable version` | +| Best practices | Check if your approach matches current consensus | `[pattern] best practice [language] 2026` | +| Content accuracy | Ground-truth check generated facts | `[claim] site:official-source.com` | +| Bundle/perf baselines | Compare your score to current industry benchmarks | `[framework] bundle size benchmark 2026` | + +--- + +## Pattern 1 — SEO content verification + +Use when: optimising blog posts, landing pages, documentation for search. + +After your local score script runs, supplement with: + +``` +WebSearch: [target keyword] to see what the top 3 results have in common. +Note: heading structure, content length, semantic coverage, internal links. +If top results consistently have trait X that your content lacks, +add "add trait X" as the next hypothesis. +``` + +This gives you signal that no local readability or keyword-density script can +provide — what the search engine is actually rewarding right now. + +--- + +## Pattern 2 — API currency check + +Use when: refactoring code that calls external libraries or APIs. + +Before committing any API-surface change: + +``` +WebSearch: [library name] [method name] changelog 2026 +WebSearch: [library name] [method name] deprecated +``` + +If search returns deprecation notices or breaking changes, note the current +replacement pattern and use that as the hypothesis instead. + +This prevents iterating toward a working-but-deprecated solution that will +break on the next library update. + +--- + +## Pattern 3 — Dependency version check + +Use when: the Verify command suggests a dependency might be outdated, or when +optimising for security/bundle size. + +``` +WebSearch: [package name] npm latest 2026 +WebSearch: [package name] security advisory +``` + +Cross-reference against what is in `package.json`, `go.mod`, `requirements.txt` +or equivalent. Use the delta as a hypothesis: "update [package] from X to Y, +check if metric improves." + +--- + +## Pattern 4 — Best practice calibration + +Use when: stuck after 5 consecutive discards and local ideas are exhausted. + +``` +WebSearch: [language/framework] [metric type] optimisation techniques 2026 +WebSearch: how to improve [metric] in [stack] +``` + +Extract 3 concrete, actionable techniques from the top results — use `WebFetch` +on the most promising one if the snippet is too thin. Do not extract vague +advice. Add each as a separate iteration hypothesis. This restocks your +hypothesis pool with externally validated approaches. + +--- + +## Pattern 5 — Benchmark calibration + +Use when: you want to know if your current metric value is good relative to +the industry, not just relative to your own baseline. + +``` +WebSearch: [framework] [metric] benchmark 2026 average +``` + +If your metric is already at or above the industry median, note this and +shift the goal definition (e.g. from "reduce bundle size" to "reduce bundle +size while improving lighthouse score"). + +--- + +## Pattern 6 — Content accuracy check + +Use when: the Verify command measures style/structure but not factual accuracy +(e.g. documentation, blog posts, runbooks). + +``` +WebSearch: [specific claim in content] site:[authoritative source] +``` + +If the authoritative source contradicts your content, flag this as a +required fix before the next iteration (accuracy issues override metric gains). + +--- + +## Rules for using WebSearch + +1. **Supplement, never replace.** The Verify command runs every iteration. + Web research adds signal; it does not replace the metric. + +2. **Search at the right time.** Patterns 1-3 supplement Phase 5 (Verify). + Patterns 4-5 are for stuck recovery in Phase 1 (Review). Pattern 6 + runs in Phase 6 (Decide) when a kept iteration touches factual claims. + +3. **Extract actionable hypotheses.** Never let a search result produce a + vague conclusion ("content could be better"). Always turn the search + result into a specific next hypothesis ("add a FAQ section with 3 + questions, which top-ranking competitors include"). + +4. **Log the research signal.** When a search result influences a hypothesis, + note it in the results log description: + `"added FAQ section (web research: top results for [kw] all include FAQ)"` + +5. **Don't over-search.** Maximum one WebSearch call per iteration. If you are + searching every iteration, your Verify command is probably too weak — + strengthen the local script instead. + +6. **Cite, don't guess.** `WebSearch` results come with source links; never + turn an unverified snippet into a change that the Guard cannot catch.