skills: autoresearch

This commit is contained in:
ryan
2026-08-28 20:36:15 +08:00
parent b68060255f
commit 85b383a4e0
10 changed files with 1405 additions and 0 deletions
+300
View File
@@ -0,0 +1,300 @@
---
name: autoresearch
description: >
Autonomous goal-directed iteration loop, inspired by Karpathy's autoresearch.
Use when asked to run autoresearch, iterate overnight, autonomously improve
any measurable goal, or drive an unattended plan/ship/debug/fix/security
workflow. Loops forever: modify → verify → keep/revert → log → repeat.
Never stops until the user interrupts.
---
# Autoresearch
> Ported from `supratikpm/gemini-autoresearch` (Gemini CLI). The loop protocol
> is unchanged; only tool-specific mechanics were mapped to Qoder equivalents —
> the `WebSearch` tool replaces Google Search grounding, `plan` / `ship` /
> `debug` / `fix` / `security` modes replace `/autoresearch:*` subcommands, and
> Qoder Automations replace `gemini --yolo`.
You are an autonomous improvement agent. You iterate forever until interrupted.
You do not ask "should I continue?" You do not pause for confirmation. You run
the loop.
## Invocation
### Standard loop
```
/autoresearch
Goal: <what to improve — be specific>
Scope: <files or directories you may modify>
Metric: <the number you are optimising, and whether higher or lower is better>
Verify: <shell command that measures progress — must output a number in under 10s>
Guard: <shell command that must always pass — optional but strongly recommended>
```
`Verify` and `Guard` serve completely different purposes:
- **Verify** = "Did the metric improve?" — measures progress toward the goal
- **Guard** = "Did anything else break?" — protects invariants unrelated to the goal
Example — improving test coverage while ensuring types never break:
```
Verify: npm test -- --coverage | grep "All files"
Guard: npx tsc --noEmit
```
`Verify` is required. `Guard` is optional but strongly recommended — without it,
the loop can silently accumulate regressions in areas outside the metric.
Guard files are **never modified** by the loop. They are read-only constraints.
Goal, Scope, Metric, and Verify are required. Guard is optional.
If any required fields are missing, ask for them once, then start.
### Modes
Invoke the skill and make the first word the mode: `autoresearch plan <goal>`,
`autoresearch security`, and so on. Qoder does not register `/autoresearch:*`
subcommands — the mode is plain text in your message.
| Mode | What it does | Reference |
|---|---|---|
| `plan <goal>` | Auto-detect stack, propose goal/scope/verify, dry run, hand back ready-to-run config | `references/plan-workflow.md` |
| `ship` | Pre-flight checklist — tests, types, lint, bundle, secrets, deps. Autoresearch loop on anything that fails | `references/ship-workflow.md` |
| `debug <description>` | Autonomous debug loop — reproduce, isolate root cause, fix, verify, harden | `references/debug-workflow.md` |
| `fix <description>` | Focused fix loop — for specific lint, type, or test failures without full debug isolation | `references/fix-workflow.md` |
| `security` | STRIDE/OWASP audit loop — threat model, find vulnerabilities, optional auto-fix | `references/security-workflow.md` |
No mode means the standard loop above.
**When a mode is invoked**, read the corresponding reference file
before doing anything else. The reference file contains the full protocol
for that workflow.
---
## Setup phase (run once before the loop)
1. Read every file in Scope to build full context. Qoder compacts older turns
automatically, so re-read Scope files instead of trusting a stale summary.
2. Read `autoresearch-lessons.md` if it exists. This is accumulated knowledge
from prior runs. Read it carefully before forming any hypothesis.
3. Run the Verify command. Record the output as the baseline (iteration #0).
4. If Guard is provided: run it once. If it fails, STOP immediately and tell
the user — the codebase is already broken before the loop starts. Fix the
Guard failure manually before proceeding. Guard must be green at baseline.
5. Initialise `autoresearch-results.tsv`:
```
iteration\tcommit\tmetric\tdelta\tstatus\tguard\tdescription
0\t-\t<baseline>\t0.0\tbaseline\tpass\tinitial measurement
```
6. Print a setup summary: goal, baseline metric, guard status (pass/skip),
scope summary, lessons loaded Y/N.
7. Start the loop immediately. Do not wait for confirmation.
---
## The loop (run forever — never stop)
### Phase 1 — Review
Read:
- Current state of all Scope files
- `git log --oneline -20` (what has been tried)
- `autoresearch-results.tsv` (what worked, what failed, patterns)
- `autoresearch-lessons.md` (accumulated wisdom from prior runs)
Identify: what directions have produced gains? what has consistently failed?
what has not been tried yet?
### Phase 2 — Ideate
Pick ONE hypothesis. It must be:
- Specific and testable in a single iteration
- Meaningfully different from the last 3 attempts
- Informed by both the results log and the lessons file
- Explained in one sentence
Prefer hypotheses that build on proven wins over untested territory.
Prefer simplicity — a small clean change beats a large complex one.
### Phase 3 — Modify
Make exactly ONE atomic change in Scope. If you cannot explain the change
in one sentence, split it into two separate iterations.
Do not touch files outside Scope. Do not refactor unrelated code. One thing.
### Phase 4 — Commit
```bash
git add -A && git commit -m "autoresearch iter N: <one-sentence description>"
```
**Commit BEFORE verifying.** This guarantees a clean, known-good rollback point
regardless of what verification reveals. Never skip this step.
### Phase 5 — Verify + Guard
**Step A — Run Verify.** Extract the numeric metric value.
If Verify crashed (exit non-zero, no number output):
- Attempt to fix the crash (max 3 tries)
- If unfixed: `git revert HEAD --no-edit`, log as "crash", go to Phase 8
If Verify regressed or is unchanged:
- `git revert HEAD --no-edit`, log as "discard", go to Phase 8
- Do NOT run Guard — a regressed change is already dead
**Step B — Run Guard (only if Verify improved).** Exit code 0 = pass.
**Web research supplement**: after Verify passes, use `WebSearch` for
additional signal when local scripts cannot capture full quality.
See `references/web-research-patterns.md`. Research is a supplement only.
### Phase 6 — Decide
The full dual-gate decision table:
| Verify | Guard | Decision | Log status |
|---|---|---|---|
| ✅ improved | ✅ pass (or no Guard set) | **KEEP** | `keep` |
| ✅ improved | ❌ fail | **REWORK** — fix Guard failure, re-run Guard (max 2 attempts). If still failing: `git revert HEAD --no-edit` | `guard-fail` |
| ❌ regressed | — | **REVERT** immediately. Do not run Guard. | `discard` |
| ❌ unchanged | — | **REVERT**. Treat unchanged as a regression. | `discard` |
| 💥 crashed | — | **FIX** (max 3 attempts), then revert if unfixed. | `crash` |
**Rework protocol** (when Verify passes but Guard fails):
1. Read the Guard failure output carefully
2. Make the minimal additional change to satisfy Guard without hurting Verify
3. Amend the commit: `git add -A && git commit --amend --no-edit`
4. Re-run both Verify AND Guard
5. If both pass → KEEP. If Guard still fails after 2 rework attempts → REVERT.
### Phase 7 — Log
Append one row to `autoresearch-results.tsv`:
```
<N>\t<commit_sha or "-">\t<metric_value>\t<delta>\t<keep|discard|guard-fail|crash>\t<guard:pass|fail|skip>\t<description>
```
Delta = metric_value − previous_best (positive = improvement for "higher is
better" goals, negative = improvement for "lower is better" goals).
### Phase 8 — Repeat
Go to Phase 1. Immediately. NEVER STOP.
---
## Progress summary (every 10 iterations)
Print this, then continue immediately:
```
=== Autoresearch progress — iteration N ===
Baseline: <value>
Current best: <value> (<delta> from baseline)
Keeps: <count>
Discards: <count>
Crashes: <count>
Top pattern: <what has worked most consistently>
Last 5: <keep/discard/crash sequence>
===
```
---
## Lessons system
After every 5 KEPT iterations, append to `autoresearch-lessons.md`:
```markdown
## Lesson <N> — iterations <range>
**Pattern**: <what change type produced gains>
**Why it worked**: <mechanistic hypothesis>
**Conditions**: <when to apply — be specific about codebase state>
**Anti-pattern**: <what failed when trying similar things>
**Metric delta**: <how much the metric moved, cumulative>
```
At the start of every run, read this file before forming any hypotheses.
Weight recent lessons more heavily. Older lessons may not apply if the
codebase or scope has changed significantly.
This is the compounding mechanism. Each overnight run starts smarter than
the last.
---
## Stuck recovery
After 5 consecutive discards or crashes:
1. Re-read all Scope files from scratch. Full context, not memory.
2. Search the lessons log for near-misses — what came closest to working?
3. Try combining two near-miss approaches into one hypothesis.
4. If still stuck after 3 more iterations: try the literal opposite of what
has been failing consistently.
5. If still stuck after 3 more: use `WebSearch` to research the
problem space. Search for `[domain] [metric] improvement techniques [year]`.
Extract 3 concrete techniques. Use each as the next 3 hypotheses.
6. If still stuck after all of the above: log a "stuck" event, note the wall
hit, and try a completely different direction. Some local optima require
architectural changes — note this for the human.
---
## Unattended / overnight mode
The one thing that stalls a loop is a permission prompt. Run it in a session
that auto-approves edits and shell, or it will wait for you every iteration.
To start it while you are away, create a Qoder Automation whose prompt is fully
self-contained — automation conversations never see this transcript:
> Read the `autoresearch` skill and start immediately. Goal: `<goal>`.
> Scope: `<scope>`. Metric: `<metric — higher/lower is better>`.
> Verify: `<command>`. Guard: `<command>`. Do not pause, do not ask questions,
> iterate until stopped.
You will wake up to `autoresearch-results.tsv` and `autoresearch-lessons.md`.
Note that a scheduled run cannot be interrupted the way a live session can, so
bound it — a Guard that vetoes, and a scope you would trust unattended.
---
## Non-negotiable rules
1. **NEVER STOP** until the user manually interrupts the run.
2. **ONE change per iteration** — atomic, explainable in one sentence.
3. **Mechanical verification only** — no "looks better", no "seems cleaner".
If you cannot measure it, you cannot use it as a signal.
4. **Commit BEFORE verifying** — always. No exceptions.
5. **Auto-revert on regression** — no debate, no "let me try one more thing".
6. **Guard is a hard veto** — Verify passing does not mean KEEP. Guard must also pass.
7. **Never modify Guard files** — they are read-only invariants, not scope.
8. **Read git history before every hypothesis** — it is your short-term memory.
9. **Read lessons before every run** — it is your long-term memory.
10. **Simplicity wins ties** — equal metric + less code = KEEP.
11. **Never touch files outside Scope** — discipline is what makes the loop safe.
12. **When in doubt, make the smaller change** — scope creep kills iterations.
---
## Reference files
**Core loop**
- `references/loop-protocol.md` — detailed phase-by-phase protocol
- `references/results-logging.md` — TSV format, summary templates, examples
- `references/lessons-system.md` — cross-run memory and compounding
**Web research**
- `references/web-research-patterns.md` — `WebSearch` supplement patterns
**Mode workflows**
- `references/plan-workflow.md` — `plan` mode — auto-detect and configure
- `references/ship-workflow.md` — `ship` mode — pre-flight checklist
- `references/debug-workflow.md` — `debug` mode — root cause and fix
- `references/fix-workflow.md` — `fix` mode — focused type/lint fix
- `references/security-workflow.md` — `security` mode — STRIDE/OWASP audit
@@ -0,0 +1,25 @@
# `autoresearch debug` mode — Autonomous Debug Loop
This workflow is triggered by the `debug` mode. It is designed to reproduce, isolate, and fix specific bugs autonomously.
## Context
Use this when something is clearly broken (e.g., a failing test, a crash, or a UI bug).
## Phase 1: Reproduction
1. Create a minimal reproduction script (e.g., `debug/repro.js` or a new test case).
2. Run the repro script and verify it fails as expected.
3. This repro command becomes your `Verify` command for the loop.
## Phase 2: Isolation
1. Use `Grep` and `Read` to find the code responsible for the failure.
2. Form a hypothesis about the root cause.
## Phase 3: Fix Loop
1. Start a standard autoresearch loop with:
- **Goal**: Fix the bug identified in the repro script.
- **Verify**: The repro command (must exit 0 on success).
- **Guard**: Existing test suite and linting.
## Phase 4: Hardening
1. After the fix is verified, add a permanent regression test to the codebase.
2. Verify that the fix holds across the entire project.
@@ -0,0 +1,31 @@
# Fix Workflow (`autoresearch fix` mode)
The `fix` workflow is a lightweight version of the `debug` loop. It is designed for situations where you have a specific, known failure (e.g., a TypeScript error or a lint violation) and you want to fix it without the overhead of full reproduction and isolation.
## Protocol
### 1. Context Loading
* Read the error message or description provided in the command.
* Identify the affected file(s).
* Read the current state of those files.
### 2. Hypothesis
* Form a direct hypothesis on how to fix the specific error.
* The fix must be minimal and targeted.
### 3. Execution
* Apply the fix.
* Commit the change.
### 4. Verification
* Run the command that triggered the original failure (e.g., `npx tsc` or `npm run lint`).
* If a `Guard` is set in the main autoresearch config, run that as well.
### 5. Decision
* If the error is gone and Guard passes: **KEEP**.
* If the error persists: **RETRY** (max 3 times) with a different approach.
* If it still fails after 3 tries: **REVERT** and report to the user.
## When to use `fix` vs `debug`
* Use **`fix`** for mechanical errors: "Fix the lint error on line 42", "Fix the missing import in `utils.ts`".
* Use **`debug`** for logical errors: "The login flow fails for users with specialized characters", "Database connection timeouts under high load".
@@ -0,0 +1,117 @@
# Lessons system
The lessons system is what separates autoresearch from a dumb
mutation loop. It is the mechanism by which each overnight run starts
smarter than the last.
---
## The compounding model
```
Night 1: 100 experiments → lessons-v1 written
Night 2: reads lessons-v1 → avoids 20 known failures → 80 net-new experiments
Night 3: reads lessons-v2 → avoids 35 known failures → faster convergence
...
```
Without the lessons system, every run starts from scratch. With it, runs
compound — each failure is learned once and never repeated.
---
## File location and format
File: `autoresearch-lessons.md` in your project root.
Add to `.gitignore` — this is a working file for the agent, not source code.
```markdown
# Autoresearch lessons — <project name>
Generated by the autoresearch skill. Do not edit manually.
Last updated: <ISO date>
## Lesson 1 — iterations 1–5
**Pattern**: <the type of change that produced gains>
**Why it worked**: <mechanistic hypothesis — be specific>
**Conditions**: <codebase state where this applies>
**Anti-pattern**: <what failed when trying similar approaches>
**Metric delta**: <cumulative gain from this pattern, e.g. "+4.2%">
## Lesson 2 — iterations 6–10
...
```
---
## When to write lessons
Append a new lesson after every 5 KEPT iterations (not every 5 total
iterations). Lessons should only describe what worked.
Failed patterns are captured implicitly — if a pattern never generates a
kept iteration, it never generates a lesson, and the loop naturally
deprioritises it via Phase 2's "different from last 3 attempts" rule.
---
## What makes a good lesson
**Good** (specific, mechanistic, conditional):
```
**Pattern**: Defer non-critical third-party scripts using loading="lazy"
**Why it worked**: Removes scripts from the critical render path, reducing
Time to Interactive without affecting functionality
**Conditions**: Applies to analytics, chat widgets, social embeds — not
to scripts required for initial page render
**Anti-pattern**: Lazy-loading scripts that are called in the first 500ms
of page load caused layout shifts and broke interactions
**Metric delta**: +6.8% Lighthouse performance score across 3 iterations
```
**Bad** (vague, not actionable):
```
**Pattern**: Make things faster
**Why it worked**: It improved performance
**Conditions**: When performance is bad
**Anti-pattern**: When it makes things worse
```
---
## How to read lessons at the start of a run
1. Read the full file — do not skip old lessons even if they seem stale.
2. For each lesson, assess: does this pattern still apply given the current
state of the codebase? If the code it describes has been significantly
refactored, downweight it.
3. Extract the top 2-3 highest-delta patterns. These are your first
hypotheses unless the results log shows they have already been exhausted.
4. Extract the anti-patterns. These are your first exclusions — do not
generate hypotheses that match these patterns.
---
## Cross-project lessons
For teams running autoresearch across multiple similar projects (e.g.
multiple Next.js apps), consider maintaining a shared lessons file at
`~/.autoresearch/global-lessons.md`.
At the start of a run, read both the project-level and global lessons.
Project-level lessons take precedence when they conflict with global ones.
This is optional but significantly accelerates convergence on new projects
that share a tech stack with already-researched ones.
---
## Lessons file maintenance
- Do not manually edit the lessons file during a run — the agent reads it
at the start of each run and its contents influence hypothesis generation.
- After a long run (100+ iterations), review the file and remove lessons
that are no longer applicable (e.g. they describe code that no longer
exists). Add a comment explaining why the lesson was removed.
- The lessons file is cumulative — never delete lessons, only annotate them
as superseded if a newer lesson contradicts them.
@@ -0,0 +1,193 @@
# Autonomous loop protocol
Detailed specification for each of the 8 phases. The SKILL.md contains the
summary version. Read this reference when you need precise guidance on edge
cases in any phase.
---
## Phase 1 — Review
**Purpose**: Build a complete, accurate picture of current state before
forming any hypothesis. Hypotheses formed without full context waste iterations.
**What to read**:
- Every file in Scope (not just the ones you last touched)
- `git log --oneline -20` — what has been attempted, in order
- `autoresearch-results.tsv` — the full record of what worked and failed
- `autoresearch-lessons.md` — accumulated patterns from prior runs
**What to extract**:
- Current metric trajectory (improving? plateauing? volatile?)
- Which change types produced the most gain per iteration
- Which change types consistently failed
- Which directions have not yet been explored
- Any patterns in crash causes
**Duration**: This phase should take as long as needed to form a genuinely
informed hypothesis. Rushing Phase 1 leads to repeated failures.
---
## Phase 2 — Ideate
**Purpose**: Select ONE hypothesis that has the highest expected gain given
what is known.
**Hypothesis selection criteria** (in order of priority):
1. Builds directly on a proven pattern from the lessons file
2. Explores a direction adjacent to a near-miss (something that almost worked)
3. Combines two near-miss approaches that individually failed
4. Tries the opposite of what consistently failed
5. Applies an externally validated technique (from `WebSearch` research)
6. Tries something entirely untested
**What makes a good hypothesis**:
- Specific: "lazy-load the user avatar component" not "improve performance"
- Testable: produces a measurable delta in the Verify command
- Atomic: one thing changes, one thing is measured
- Explainable in one sentence before you make the change
**What makes a bad hypothesis**:
- Vague: "refactor for clarity"
- Multi-part: "update the API, add caching, and fix the tests"
- Untestable by the Verify command
- Identical to something tried in the last 3 iterations
---
## Phase 3 — Modify
**Purpose**: Implement the hypothesis as a single, clean, minimal change.
**Rules**:
- Touch only files in Scope
- Make the smallest change that tests the hypothesis
- If the change is getting large, stop and split it — make the first half now,
the second half in the next iteration
- Do not fix unrelated things you notice while editing
- Do not reformat code that is not part of the hypothesis
- Leave comments only if they directly explain the change
**Signs you are over-scoping**:
- You have edited more than 3 files
- The diff is more than ~50 lines
- You are explaining the change with "and also"
When in doubt, make a smaller change. Smaller changes fail faster and teach more.
---
## Phase 4 — Commit
**Purpose**: Create a clean rollback point before any verification risk.
**Command**:
```bash
git add -A && git commit -m "autoresearch iter N: <one-sentence description>"
```
**Commit message format**:
- Always prefix with `autoresearch iter N:`
- One sentence, present tense, describes the change not the goal
- Good: `autoresearch iter 14: lazy-load user avatar to reduce initial bundle`
- Bad: `autoresearch iter 14: improve performance`
**Why commit before verifying**: if the Verify command crashes, hangs, or
corrupts state, you can always `git revert HEAD --no-edit` and return to
a known-good state. If you verify before committing, a crash during
verification leaves you with uncommitted changes and an unknown baseline.
**Never skip this step**, even if the change feels obviously correct.
---
## Phase 5 — Verify
**Purpose**: Get a single numeric measurement of whether the hypothesis helped.
**Execution**:
1. Run the Verify command exactly as specified by the user
2. Extract the numeric metric value
3. Optionally supplement with `WebSearch` research (see
`references/web-research-patterns.md`)
4. Record the raw output for the log
**Handling slow Verify commands**:
If the Verify command takes more than 30 seconds, note this. After the run,
recommend the user find a faster proxy metric — slower verification means
fewer experiments per hour, which compounds negatively over a full night.
**Handling non-deterministic Verify commands**:
If the metric varies significantly between runs on identical code (>5%
variance), note this in the log. Run the Verify command twice and average.
Log both values. Recommend the user address flakiness before the next
overnight run.
---
## Phase 6 — Decide
**Purpose**: Make a clear, mechanical keep/revert decision. No deliberation.
**Decision table**:
| Condition | Action | Log status |
|---|---|---|
| Metric improved (beyond noise threshold) | Keep commit as-is | `keep` |
| Metric unchanged or regressed | `git revert HEAD --no-edit` | `discard` |
| Verify crashed with exit code ≠ 0 | Attempt fix (max 3 tries) then revert | `crash` |
| Verify hung for >60s | Kill process, revert | `crash` |
**Noise threshold**: for metrics with variance, an improvement smaller than
the variance is not a real improvement. If your metric normally varies ±2%,
an improvement of 0.5% is noise — treat it as unchanged and discard.
**The revert command**:
```bash
git revert HEAD --no-edit
```
This creates a new commit that undoes the last one. The history is preserved.
Never use `git reset --hard` — it destroys history that the loop needs.
---
## Phase 7 — Log
**Purpose**: Create a permanent, machine-readable record of every iteration.
**TSV row format**:
```
<N>\t<commit_sha or "-">\t<metric>\t<delta>\t<status>\t<description>
```
**Field details**:
- `N`: integer, 0-indexed, never resets across sessions
- `commit_sha`: 7-char short SHA for keeps, "-" for discards/crashes
- `metric`: the exact number from the Verify output
- `delta`: metric − previous_best (sign convention: positive = better,
regardless of whether the goal is higher or lower)
- `status`: one of `baseline`, `keep`, `discard`, `crash`
- `description`: the hypothesis, in one sentence, including any `WebSearch`
signal that informed it
**Example rows**:
```
0 - 85.2 0.0 baseline initial measurement
1 a1b2c3d 87.1 +1.9 keep lazy-load avatar component
2 - 86.5 -0.6 discard tree-shake lodash imports (broke 2 tests)
3 - 0.0 0.0 crash add route-level code splitting (webpack config error)
4 b2c3d4e 88.3 +1.2 keep move analytics script to defer loading
```
---
## Phase 8 — Repeat
Go to Phase 1. Immediately. Do not pause. Do not summarise. Do not ask
if the user wants to continue.
The only output before starting Phase 1 again is the progress summary
(printed every 10 iterations, see SKILL.md).
The loop ends only when the user interrupts the run.
@@ -0,0 +1,155 @@
# Plan workflow — `autoresearch plan` mode
Auto-detect the project stack, propose a complete autoresearch configuration,
do a dry run, and hand the ready-to-run command back to the user.
No manual goal/scope/verify required. Just describe what you want to improve
in one sentence and the plan workflow figures out the rest.
---
## Invocation
```
autoresearch plan <goal in plain english>
```
Examples:
```
autoresearch plan improve test coverage
autoresearch plan make the app faster
autoresearch plan reduce the bundle size
autoresearch plan fix all TypeScript errors
autoresearch plan improve the SEO of my blog posts
autoresearch plan shrink the Docker image
```
---
## What the plan workflow does
### Step 1 — Detect project stack
Scan the project root for signal files:
| File found | Stack detected |
|---|---|
| `package.json` + `jest.config.*` | Node.js + Jest |
| `package.json` + `vitest.config.*` | Node.js + Vitest |
| `next.config.*` | Next.js |
| `Dockerfile` | Docker |
| `*.tf` | Terraform |
| `.github/workflows/*.yml` | GitHub Actions CI |
| `content/blog/*.md` OR `posts/*.md` | Markdown content/blog |
| `src/**/*.ts` OR `src/**/*.tsx` | TypeScript project |
| `pyproject.toml` OR `setup.py` | Python project |
| `requirements.txt` + `pytest` | Python + pytest |
| `go.mod` | Go project |
| `Cargo.toml` | Rust project |
Print detected stack. If ambiguous, list the top two candidates and ask
the user to confirm before proceeding.
### Step 2 — Map goal to metric + verify command
Use the goal description and detected stack to propose:
| Goal keyword | Metric | Verify command template |
|---|---|---|
| "test coverage" | coverage % (higher is better) | `npm test -- --coverage \| grep "All files"` |
| "bundle size" / "build size" | size in KB (lower is better) | `npm run build 2>&1 \| grep "First Load JS"` |
| "TypeScript errors" / "type errors" | error count (lower is better) | `npx tsc --noEmit 2>&1 \| grep -c "error TS" \|\| echo "0"` |
| "lighthouse" / "performance score" | score 0-100 (higher is better) | `npx lighthouse http://localhost:3000 --output json --quiet 2>/dev/null \| jq '.categories.performance.score * 100'` |
| "docker image" / "image size" | size in MB (lower is better) | `docker build -t bench . -q && docker images bench --format "{{.Size}}"` |
| "flaky tests" | failure count (lower is better) | `for i in {1..5}; do npm test 2>&1; done \| grep -c "FAIL" \|\| echo "0"` |
| "SEO" / "blog" / "content" | SEO score (higher is better) | `node scripts/seo-score.js <detected content path>` |
| "lines of code" / "complexity" | LOC count (lower is better) | `find src/ -name "*.ts" \| xargs wc -l \| tail -1 \| awk '{print $1}'` |
| "CI pipeline" / "pipeline speed" | seconds (lower is better) | `node scripts/estimate-ci-time.js` |
| "Python tests" / "pytest" | coverage % (higher is better) | `pytest --cov=src --cov-report=term-missing \| grep "TOTAL"` |
| "faster" / "performance" / "latency" | p95 ms (lower is better) | `npm run bench 2>&1 \| grep "p95"` |
### Step 3 — Detect scope
Based on goal + stack, propose the tightest scope that covers the goal:
- Test coverage → `src/**/*.ts, src/**/*.test.ts`
- Bundle size → `src/**/*.tsx, src/**/*.ts`
- Docker → `Dockerfile, .dockerignore`
- SEO → `content/blog/*.md` or detected content directory
- TypeScript errors → `src/**/*.ts`
- CI pipeline → `.github/workflows/*.yml`
### Step 4 — Dry run
Run the proposed Verify command once against the current state.
- If it exits 0 and outputs a number → baseline confirmed, proceed
- If it exits non-zero → diagnose and fix the verify command before proposing
- If it hangs → propose a faster alternative
### Step 5 — Output the ready-to-run command
Print this exact block for the user to copy-paste or confirm:
```
=== Autoresearch plan ===
Stack: <detected stack>
Goal: <interpreted goal>
Scope: <proposed scope>
Metric: <metric name> (<higher/lower> is better)
Verify: <verify command>
Baseline: <dry run result>
Ready to run. Confirm or adjust any field, then:
/autoresearch
Goal: <goal>
Scope: <scope>
Metric: <metric>
Verify: <verify command>
Or, for an unattended run, put these same fields into a Qoder Automation prompt
(see "Unattended / overnight mode" in SKILL.md).
===
```
If the user says "looks good" or "run it" — start the autoresearch loop
immediately without requiring them to retype the command.
---
## Web research calibration
After the dry run, use `WebSearch` to calibrate:
- For SEO goals: search for `[target keyword]` to see what top results look like.
Note any structural patterns (FAQ sections, word count, heading structure)
that the current content lacks. Add these as initial hypotheses.
- For performance goals: search for `[framework] performance benchmarks [year]`
to calibrate whether the baseline is already good or has significant headroom.
- For security goals: search for `[stack] common vulnerabilities [year]`
to seed the initial hypothesis pool with known attack vectors.
This research step happens during plan, not during the loop — so it adds
context once without slowing down iterations.
---
## Edge cases
**Goal is too vague** ("make it better"):
Ask one clarifying question: "Better in what way — speed, quality, size,
coverage, or something else?" Then proceed.
**Multiple valid verify commands exist**:
Propose the fastest one. Note the slower alternative in a comment.
**Verify command requires a running server**:
Note this in the plan output. Add a `# requires: local server on :3000`
comment. Suggest the user start it before running the loop.
**No matching stack detected**:
Ask the user to describe their stack in one sentence, then proceed with
a custom verify command.
@@ -0,0 +1,105 @@
# Results logging
Specification for `autoresearch-results.tsv` — the per-iteration record
of every experiment in a run.
---
## File format
Tab-separated values. Headers on row 1. One row per iteration.
```
iteration\tcommit\tmetric\tdelta\tstatus\tdescription
```
### Field definitions
| Field | Type | Description |
|---|---|---|
| `iteration` | integer | 0-indexed. Never resets — if you run multiple sessions, continue from the last number. |
| `commit` | string | 7-char git short SHA for kept commits. `-` for discards and crashes. |
| `metric` | float | Raw metric value from the Verify command. |
| `delta` | float | `metric − previous_best`. Sign convention: positive = improvement (regardless of higher/lower goal). |
| `status` | enum | One of: `baseline`, `keep`, `discard`, `crash` |
| `description` | string | The hypothesis, one sentence. Include the change type and the expected mechanism. |
---
## Example file
```tsv
iteration commit metric delta status description
0 - 85.2 0.0 baseline initial measurement — test coverage 85.2%
1 a1b2c3d 87.1 +1.9 keep add tests for auth middleware edge cases
2 - 86.5 -0.7 discard refactor test helpers (broke 2 existing tests)
3 - 0.0 0.0 crash add integration tests (postgres connection failed — fix in iter 4)
4 b2c3d4e 88.3 +1.2 keep add tests for error handling in API routes
5 - 88.1 -0.2 discard add tests for rate limiter (metric within variance, treated as regression)
6 c3d4e5f 89.0 +0.7 keep add boundary value tests for form validators
7 d4e5f6g 89.8 +0.8 keep add tests for session expiry edge cases
8 - 89.2 -0.6 discard mock external API calls (test isolation but metric regressed)
9 e5f6g7h 90.6 +0.8 keep add tests for concurrent request handling
10 f6g7h8i 91.1 +0.5 keep add tests for malformed JSON input handling
```
---
## Progress summary format
Print every 10 iterations. Use this exact format:
```
=== Autoresearch progress — iteration <N> ===
Goal: <original goal statement>
Baseline: <iteration 0 metric>
Current best: <best metric so far> (<total delta> from baseline)
Keeps: <count> (<keeps/total * 100>%)
Discards: <count>
Crashes: <count>
Top pattern: <the change type that has produced the most total delta>
Last 5: <sequence of keep/discard/crash for iterations N-4 through N>
Est. to goal: <if goal metric is known, N iterations at current rate>
===
```
---
## Interpreting the log
### Healthy run signature
- Keep rate 40-60%
- Delta per keep: consistent small positive gains
- No long crash streaks
- Discards are evenly distributed (not clustered)
### Warning signs
| Pattern | Meaning | Action |
|---|---|---|
| Keep rate < 20% | Hypothesis quality is poor | Re-read full scope, re-read lessons, change direction |
| Keep rate > 80% | Metric may be too easy or Verify too lenient | Tighten the goal |
| Long crash streak (5+) | Verify command is fragile or scope is too risky | Fix Verify or narrow scope |
| Delta per keep shrinking toward 0 | Approaching local optimum | Try more radical changes or declare victory |
| Metric oscillating | Non-deterministic Verify or contradictory changes | Run Verify twice and average; tighten scope |
### Declaring success
Stop the loop when one of these is true:
- Metric has reached the stated goal
- Delta per keep has been below 0.1% for 20 consecutive iterations
(local optimum with current scope)
- All directions have been exhausted (lessons file confirms this)
In all cases, print a final summary and write a lessons entry covering
the full run before stopping.
---
## File hygiene
- Add `autoresearch-results.tsv` to `.gitignore`. It is a working file.
- Do not edit it manually during a run.
- Between runs, you may archive it:
`mv autoresearch-results.tsv autoresearch-results-<date>.tsv`
and start fresh, but keep the lessons file — that is the persistent memory.
@@ -0,0 +1,171 @@
# Security workflow — `autoresearch security` mode
Autonomous security audit using STRIDE threat modelling and OWASP categories.
Finds vulnerabilities, classifies them by severity, and optionally fixes
confirmed critical and high findings via an autoresearch loop.
---
## Invocation
```
autoresearch security # full audit, report only
autoresearch security --fix # audit + auto-fix confirmed findings
autoresearch security --fail-on critical # end with a FAIL verdict if critical found
autoresearch security --scope src/api/ # audit a specific directory only
```
---
## Phase 1 — Asset discovery
Map the attack surface:
1. Identify all entry points: API routes, form handlers, file uploads,
auth flows, webhooks, admin panels
2. Identify all data stores: databases, caches, file system writes,
environment variables, secrets
3. Identify all trust boundaries: public vs authenticated, user vs admin,
internal vs external services
4. Map data flows: what user input reaches what data store via what path
Output: `security/audit-<timestamp>/attack-surface-map.md`
### Live threat intelligence
Use `WebSearch` to seed the audit with current threats:
```
WebSearch: [your stack] common vulnerabilities [current year]
WebSearch: [your main framework] CVE [current year]
WebSearch: OWASP top 10 [current year]
```
Add any newly discovered attack patterns to the audit queue.
This ensures the audit covers threats that postdate your static analysis tools.
---
## Phase 2 — STRIDE threat model
For each asset and trust boundary, model threats across all 6 STRIDE categories:
| Category | Question to ask |
|---|---|
| **S**poofing | Can an attacker impersonate a user, service, or system? |
| **T**ampering | Can input be modified to alter data or behaviour unexpectedly? |
| **R**epudiation | Can actions be performed without a traceable audit trail? |
| **I**nformation disclosure | Can sensitive data be accessed by unauthorised parties? |
| **D**enial of service | Can the service be made unavailable through normal inputs? |
| **E**levation of privilege | Can a lower-privilege user gain higher-privilege access? |
Output: `security/audit-<timestamp>/threat-model.md`
---
## Phase 3 — Autonomous audit loop
```
LOOP (through all attack vectors from threat model):
1. Select next untested attack vector
2. Deep-dive into the relevant code (read fully — do not skim)
3. Attempt to construct a concrete exploit scenario
4. Validate with code evidence (file:line + exact scenario)
5. Classify: severity + OWASP category + STRIDE tag
6. Log to security-audit-results.tsv
7. Print coverage summary every 5 iterations
8. Continue until all vectors tested
```
### Severity classification
| Severity | Definition |
|---|---|
| Critical | Exploitable without authentication, leads to full compromise or data breach |
| High | Exploitable with low-privilege access, significant impact |
| Medium | Requires specific conditions, moderate impact |
| Low | Minor information disclosure, no direct exploitation path |
| Info | Best practice violation, no immediate security impact |
### Evidence requirement
Every finding MUST have:
- File path and line number
- Exact vulnerable code snippet (copy from source, do not paraphrase)
- Concrete exploit scenario (how an attacker would trigger this)
- Proof of exploitability (not theoretical — show the actual path)
Findings without concrete evidence are logged as "unconfirmed" and flagged
for manual review, not included in the fix loop.
---
## Phase 4 — Report generation
Output folder: `security/audit-<timestamp>/`
```
security/audit-20260325-1430/
├── overview.md ← executive summary + finding counts by severity
├── threat-model.md ← STRIDE analysis per asset
├── attack-surface-map.md ← entry points, data flows, trust boundaries
├── findings.md ← all confirmed findings, sorted by severity
├── owasp-coverage.md ← coverage matrix — which OWASP categories checked
├── recommendations.md ← fix guidance for each confirmed finding
└── security-audit-results.tsv ← machine-readable log of all iterations
```
Print summary:
```
=== Security audit summary ===
Critical: <N>
High: <N>
Medium: <N>
Low: <N>
Info: <N>
Vectors tested: <N> / <total>
OWASP categories covered: <list>
Full report: security/audit-<timestamp>/overview.md
===
```
---
## Phase 5 — Auto-fix loop (with `--fix`)
Only runs when `--fix` flag is passed.
Only fixes **Confirmed Critical and High** findings.
Uses `recommendations.md` as the fix guide for each finding.
```
FOR EACH confirmed Critical/High finding:
1. Read the finding + recommendation
2. Make ONE targeted fix
3. git commit the fix
4. Re-run the specific exploit scenario to verify it no longer works
5. Run full test suite to confirm no regressions
6. If tests break → revert, try alternative fix
7. Maximum 3 attempts per finding, then skip and flag for manual review
8. Log fix outcome to fix-log.md
```
---
## Verdict mode (`--fail-on`)
```
autoresearch security --fail-on critical
```
The audit ends with an explicit verdict line in `overview.md`:
```
VERDICT: FAIL — 2 findings at or above `critical`
VERDICT: PASS — no findings at or above `critical`
```
A skill run has no process exit code, so do not wire this into a CI gate as if
it did — use a real scanner for blocking merges. What it *is* good for is an
unattended scheduled audit: a Qoder Automation running this mode reports the
verdict, and you act on it.
@@ -0,0 +1,164 @@
# Ship workflow — `autoresearch ship`
Run a pre-flight checklist before shipping — tests, types, lint, bundle size,
security basics, and a final autoresearch pass on anything that fails.
The ship workflow is not just a checklist. It runs an autoresearch loop on
each failing gate until it passes, then re-checks. You don't ship broken.
You ship when everything is green.
---
## Invocation
```
autoresearch ship
```
Optional flags:
```
autoresearch ship --fast # skip slow checks (lighthouse, e2e)
autoresearch ship --loop N # max N autoresearch iterations per gate (default: 20)
autoresearch ship --dry-run # report status without fixing anything
```
---
## The ship checklist
The workflow runs these gates in order. Each gate that fails triggers an
autoresearch sub-loop to fix it before moving to the next gate.
### Gate 1 — Tests pass
```bash
npm test # Node.js
pytest # Python
go test ./... # Go
cargo test # Rust
```
If tests fail → autoresearch loop on `src/**/*.ts` (or equivalent) with
metric: failing test count (lower is better), max 20 iterations.
### Gate 2 — No type errors
```bash
npx tsc --noEmit # TypeScript
mypy src/ # Python
```
If errors found → autoresearch loop on `src/**/*.ts` with
metric: error count (lower is better), max 20 iterations.
### Gate 3 — No lint errors
```bash
npx eslint src/ # JavaScript/TypeScript
ruff check src/ # Python
golangci-lint run # Go
```
If errors found → autoresearch loop with metric: lint error count (lower is better).
Auto-fixable errors are fixed first (`--fix` flag), then the loop handles the rest.
### Gate 4 — Bundle size (if applicable)
Only runs for frontend projects (detected: `next.config.*`, `vite.config.*`,
`webpack.config.*`).
```bash
npm run build 2>&1 | grep "First Load JS"
```
Threshold: warn if > 300KB, block if > 500KB (configurable via `.autoresearch.yml`).
If over threshold → autoresearch loop on `src/**/*.tsx, src/**/*.ts` with
metric: bundle size in KB (lower is better), max 20 iterations.
### Gate 5 — No hardcoded secrets
```bash
git diff HEAD~1 --diff-filter=A | grep -iE "(api_key|secret|password|token)\s*=\s*['\"][^'\"]{8,}"
```
If secrets found → do NOT autoresearch. Flag for human review. Block ship.
### Gate 6 — Dependency audit
```bash
npm audit --audit-level=high # Node.js
pip-audit # Python
```
If critical vulnerabilities found → autoresearch loop to update affected
dependencies, max 10 iterations.
---
## Ship report
After all gates pass, print:
```
=== Ship report ===
Tests: ✓ PASS (247 passing)
Types: ✓ PASS (0 errors)
Lint: ✓ PASS (0 errors)
Bundle: ✓ PASS (187KB)
Secrets: ✓ PASS (none detected)
Deps: ✓ PASS (0 high/critical)
Autoresearch loops run: <N>
Total improvements: <M> iterations kept
Ready to ship. Run: git push && <your deploy command>
===
```
If any gate is still failing after the max iterations:
```
=== Ship report ===
Tests: ✓ PASS
Types: ✗ FAIL (3 errors remaining after 20 iterations)
→ manual fix required: src/auth/session.ts:47
Ship BLOCKED. Fix the above before shipping.
===
```
---
## Web research post-check
After all gates pass, use `WebSearch` to check:
```
WebSearch: [your framework] [version] known issues [current year]
WebSearch: [your main dependencies] security advisory [current year]
```
If any critical advisories surface that the dependency audit missed,
flag them before shipping. This is a final sanity check that goes beyond
what local tools can detect.
---
## Configuration via `.autoresearch.yml`
Create this file in your project root to customise ship behaviour:
```yaml
ship:
bundle_warn_kb: 300
bundle_block_kb: 500
max_iterations_per_gate: 20
skip_gates:
- lighthouse # skip if no local server available
extra_gates:
- name: "E2E tests"
command: "npx playwright test"
metric: "failing tests (lower is better)"
max_iterations: 10
```
@@ -0,0 +1,144 @@
# Web research patterns
Qoder exposes a `WebSearch` tool (and `WebFetch` to read a promising result in
full). Use them as a verification supplement — not a replacement for the Verify
command, but an additional signal when local scripts alone cannot capture
quality.
---
## When to use WebSearch in the loop
| Goal type | Use WebSearch for | Example query |
|---|---|---|
| SEO content | Check competing pages, keyword signals | `[target keyword] filetype:md OR site:*.dev` |
| API correctness | Verify endpoint signatures, check for deprecations | `[library] [method] deprecated 2025 OR 2026` |
| Dependency versions | Confirm latest stable before updating | `[package name] latest stable version` |
| Best practices | Check if your approach matches current consensus | `[pattern] best practice [language] 2026` |
| Content accuracy | Ground-truth check generated facts | `[claim] site:official-source.com` |
| Bundle/perf baselines | Compare your score to current industry benchmarks | `[framework] bundle size benchmark 2026` |
---
## Pattern 1 — SEO content verification
Use when: optimising blog posts, landing pages, documentation for search.
After your local score script runs, supplement with:
```
WebSearch: [target keyword] to see what the top 3 results have in common.
Note: heading structure, content length, semantic coverage, internal links.
If top results consistently have trait X that your content lacks,
add "add trait X" as the next hypothesis.
```
This gives you signal that no local readability or keyword-density script can
provide — what the search engine is actually rewarding right now.
---
## Pattern 2 — API currency check
Use when: refactoring code that calls external libraries or APIs.
Before committing any API-surface change:
```
WebSearch: [library name] [method name] changelog 2026
WebSearch: [library name] [method name] deprecated
```
If search returns deprecation notices or breaking changes, note the current
replacement pattern and use that as the hypothesis instead.
This prevents iterating toward a working-but-deprecated solution that will
break on the next library update.
---
## Pattern 3 — Dependency version check
Use when: the Verify command suggests a dependency might be outdated, or when
optimising for security/bundle size.
```
WebSearch: [package name] npm latest 2026
WebSearch: [package name] security advisory
```
Cross-reference against what is in `package.json`, `go.mod`, `requirements.txt`
or equivalent. Use the delta as a hypothesis: "update [package] from X to Y,
check if metric improves."
---
## Pattern 4 — Best practice calibration
Use when: stuck after 5 consecutive discards and local ideas are exhausted.
```
WebSearch: [language/framework] [metric type] optimisation techniques 2026
WebSearch: how to improve [metric] in [stack]
```
Extract 3 concrete, actionable techniques from the top results — use `WebFetch`
on the most promising one if the snippet is too thin. Do not extract vague
advice. Add each as a separate iteration hypothesis. This restocks your
hypothesis pool with externally validated approaches.
---
## Pattern 5 — Benchmark calibration
Use when: you want to know if your current metric value is good relative to
the industry, not just relative to your own baseline.
```
WebSearch: [framework] [metric] benchmark 2026 average
```
If your metric is already at or above the industry median, note this and
shift the goal definition (e.g. from "reduce bundle size" to "reduce bundle
size while improving lighthouse score").
---
## Pattern 6 — Content accuracy check
Use when: the Verify command measures style/structure but not factual accuracy
(e.g. documentation, blog posts, runbooks).
```
WebSearch: [specific claim in content] site:[authoritative source]
```
If the authoritative source contradicts your content, flag this as a
required fix before the next iteration (accuracy issues override metric gains).
---
## Rules for using WebSearch
1. **Supplement, never replace.** The Verify command runs every iteration.
Web research adds signal; it does not replace the metric.
2. **Search at the right time.** Patterns 1-3 supplement Phase 5 (Verify).
Patterns 4-5 are for stuck recovery in Phase 1 (Review). Pattern 6
runs in Phase 6 (Decide) when a kept iteration touches factual claims.
3. **Extract actionable hypotheses.** Never let a search result produce a
vague conclusion ("content could be better"). Always turn the search
result into a specific next hypothesis ("add a FAQ section with 3
questions, which top-ranking competitors include").
4. **Log the research signal.** When a search result influences a hypothesis,
note it in the results log description:
`"added FAQ section (web research: top results for [kw] all include FAQ)"`
5. **Don't over-search.** Maximum one WebSearch call per iteration. If you are
searching every iteration, your Verify command is probably too weak —
strengthen the local script instead.
6. **Cite, don't guess.** `WebSearch` results come with source links; never
turn an unverified snippet into a change that the Guard cannot catch.