mirror of
https://github.com/Rain-kl/OpenFlare.git
synced 2026-09-28 05:46:36 +08:00
skills: autoresearch
This commit is contained in:
@@ -0,0 +1,300 @@
|
||||
---
|
||||
name: autoresearch
|
||||
description: >
|
||||
Autonomous goal-directed iteration loop, inspired by Karpathy's autoresearch.
|
||||
Use when asked to run autoresearch, iterate overnight, autonomously improve
|
||||
any measurable goal, or drive an unattended plan/ship/debug/fix/security
|
||||
workflow. Loops forever: modify → verify → keep/revert → log → repeat.
|
||||
Never stops until the user interrupts.
|
||||
---
|
||||
|
||||
# Autoresearch
|
||||
|
||||
> Ported from `supratikpm/gemini-autoresearch` (Gemini CLI). The loop protocol
|
||||
> is unchanged; only tool-specific mechanics were mapped to Qoder equivalents —
|
||||
> the `WebSearch` tool replaces Google Search grounding, `plan` / `ship` /
|
||||
> `debug` / `fix` / `security` modes replace `/autoresearch:*` subcommands, and
|
||||
> Qoder Automations replace `gemini --yolo`.
|
||||
|
||||
You are an autonomous improvement agent. You iterate forever until interrupted.
|
||||
You do not ask "should I continue?" You do not pause for confirmation. You run
|
||||
the loop.
|
||||
|
||||
## Invocation
|
||||
|
||||
### Standard loop
|
||||
```
|
||||
/autoresearch
|
||||
Goal: <what to improve — be specific>
|
||||
Scope: <files or directories you may modify>
|
||||
Metric: <the number you are optimising, and whether higher or lower is better>
|
||||
Verify: <shell command that measures progress — must output a number in under 10s>
|
||||
Guard: <shell command that must always pass — optional but strongly recommended>
|
||||
```
|
||||
|
||||
`Verify` and `Guard` serve completely different purposes:
|
||||
- **Verify** = "Did the metric improve?" — measures progress toward the goal
|
||||
- **Guard** = "Did anything else break?" — protects invariants unrelated to the goal
|
||||
|
||||
Example — improving test coverage while ensuring types never break:
|
||||
```
|
||||
Verify: npm test -- --coverage | grep "All files"
|
||||
Guard: npx tsc --noEmit
|
||||
```
|
||||
|
||||
`Verify` is required. `Guard` is optional but strongly recommended — without it,
|
||||
the loop can silently accumulate regressions in areas outside the metric.
|
||||
|
||||
Guard files are **never modified** by the loop. They are read-only constraints.
|
||||
|
||||
Goal, Scope, Metric, and Verify are required. Guard is optional.
|
||||
If any required fields are missing, ask for them once, then start.
|
||||
|
||||
### Modes
|
||||
|
||||
Invoke the skill and make the first word the mode: `autoresearch plan <goal>`,
|
||||
`autoresearch security`, and so on. Qoder does not register `/autoresearch:*`
|
||||
subcommands — the mode is plain text in your message.
|
||||
|
||||
| Mode | What it does | Reference |
|
||||
|---|---|---|
|
||||
| `plan <goal>` | Auto-detect stack, propose goal/scope/verify, dry run, hand back ready-to-run config | `references/plan-workflow.md` |
|
||||
| `ship` | Pre-flight checklist — tests, types, lint, bundle, secrets, deps. Autoresearch loop on anything that fails | `references/ship-workflow.md` |
|
||||
| `debug <description>` | Autonomous debug loop — reproduce, isolate root cause, fix, verify, harden | `references/debug-workflow.md` |
|
||||
| `fix <description>` | Focused fix loop — for specific lint, type, or test failures without full debug isolation | `references/fix-workflow.md` |
|
||||
| `security` | STRIDE/OWASP audit loop — threat model, find vulnerabilities, optional auto-fix | `references/security-workflow.md` |
|
||||
|
||||
No mode means the standard loop above.
|
||||
|
||||
**When a mode is invoked**, read the corresponding reference file
|
||||
before doing anything else. The reference file contains the full protocol
|
||||
for that workflow.
|
||||
|
||||
---
|
||||
|
||||
## Setup phase (run once before the loop)
|
||||
|
||||
1. Read every file in Scope to build full context. Qoder compacts older turns
|
||||
automatically, so re-read Scope files instead of trusting a stale summary.
|
||||
2. Read `autoresearch-lessons.md` if it exists. This is accumulated knowledge
|
||||
from prior runs. Read it carefully before forming any hypothesis.
|
||||
3. Run the Verify command. Record the output as the baseline (iteration #0).
|
||||
4. If Guard is provided: run it once. If it fails, STOP immediately and tell
|
||||
the user — the codebase is already broken before the loop starts. Fix the
|
||||
Guard failure manually before proceeding. Guard must be green at baseline.
|
||||
5. Initialise `autoresearch-results.tsv`:
|
||||
```
|
||||
iteration\tcommit\tmetric\tdelta\tstatus\tguard\tdescription
|
||||
0\t-\t<baseline>\t0.0\tbaseline\tpass\tinitial measurement
|
||||
```
|
||||
6. Print a setup summary: goal, baseline metric, guard status (pass/skip),
|
||||
scope summary, lessons loaded Y/N.
|
||||
7. Start the loop immediately. Do not wait for confirmation.
|
||||
|
||||
---
|
||||
|
||||
## The loop (run forever — never stop)
|
||||
|
||||
### Phase 1 — Review
|
||||
|
||||
Read:
|
||||
- Current state of all Scope files
|
||||
- `git log --oneline -20` (what has been tried)
|
||||
- `autoresearch-results.tsv` (what worked, what failed, patterns)
|
||||
- `autoresearch-lessons.md` (accumulated wisdom from prior runs)
|
||||
|
||||
Identify: what directions have produced gains? what has consistently failed?
|
||||
what has not been tried yet?
|
||||
|
||||
### Phase 2 — Ideate
|
||||
|
||||
Pick ONE hypothesis. It must be:
|
||||
- Specific and testable in a single iteration
|
||||
- Meaningfully different from the last 3 attempts
|
||||
- Informed by both the results log and the lessons file
|
||||
- Explained in one sentence
|
||||
|
||||
Prefer hypotheses that build on proven wins over untested territory.
|
||||
Prefer simplicity — a small clean change beats a large complex one.
|
||||
|
||||
### Phase 3 — Modify
|
||||
|
||||
Make exactly ONE atomic change in Scope. If you cannot explain the change
|
||||
in one sentence, split it into two separate iterations.
|
||||
|
||||
Do not touch files outside Scope. Do not refactor unrelated code. One thing.
|
||||
|
||||
### Phase 4 — Commit
|
||||
|
||||
```bash
|
||||
git add -A && git commit -m "autoresearch iter N: <one-sentence description>"
|
||||
```
|
||||
|
||||
**Commit BEFORE verifying.** This guarantees a clean, known-good rollback point
|
||||
regardless of what verification reveals. Never skip this step.
|
||||
|
||||
### Phase 5 — Verify + Guard
|
||||
|
||||
**Step A — Run Verify.** Extract the numeric metric value.
|
||||
|
||||
If Verify crashed (exit non-zero, no number output):
|
||||
- Attempt to fix the crash (max 3 tries)
|
||||
- If unfixed: `git revert HEAD --no-edit`, log as "crash", go to Phase 8
|
||||
|
||||
If Verify regressed or is unchanged:
|
||||
- `git revert HEAD --no-edit`, log as "discard", go to Phase 8
|
||||
- Do NOT run Guard — a regressed change is already dead
|
||||
|
||||
**Step B — Run Guard (only if Verify improved).** Exit code 0 = pass.
|
||||
|
||||
**Web research supplement**: after Verify passes, use `WebSearch` for
|
||||
additional signal when local scripts cannot capture full quality.
|
||||
See `references/web-research-patterns.md`. Research is a supplement only.
|
||||
|
||||
### Phase 6 — Decide
|
||||
|
||||
The full dual-gate decision table:
|
||||
|
||||
| Verify | Guard | Decision | Log status |
|
||||
|---|---|---|---|
|
||||
| ✅ improved | ✅ pass (or no Guard set) | **KEEP** | `keep` |
|
||||
| ✅ improved | ❌ fail | **REWORK** — fix Guard failure, re-run Guard (max 2 attempts). If still failing: `git revert HEAD --no-edit` | `guard-fail` |
|
||||
| ❌ regressed | — | **REVERT** immediately. Do not run Guard. | `discard` |
|
||||
| ❌ unchanged | — | **REVERT**. Treat unchanged as a regression. | `discard` |
|
||||
| 💥 crashed | — | **FIX** (max 3 attempts), then revert if unfixed. | `crash` |
|
||||
|
||||
**Rework protocol** (when Verify passes but Guard fails):
|
||||
1. Read the Guard failure output carefully
|
||||
2. Make the minimal additional change to satisfy Guard without hurting Verify
|
||||
3. Amend the commit: `git add -A && git commit --amend --no-edit`
|
||||
4. Re-run both Verify AND Guard
|
||||
5. If both pass → KEEP. If Guard still fails after 2 rework attempts → REVERT.
|
||||
|
||||
### Phase 7 — Log
|
||||
|
||||
Append one row to `autoresearch-results.tsv`:
|
||||
|
||||
```
|
||||
<N>\t<commit_sha or "-">\t<metric_value>\t<delta>\t<keep|discard|guard-fail|crash>\t<guard:pass|fail|skip>\t<description>
|
||||
```
|
||||
|
||||
Delta = metric_value − previous_best (positive = improvement for "higher is
|
||||
better" goals, negative = improvement for "lower is better" goals).
|
||||
|
||||
### Phase 8 — Repeat
|
||||
|
||||
Go to Phase 1. Immediately. NEVER STOP.
|
||||
|
||||
---
|
||||
|
||||
## Progress summary (every 10 iterations)
|
||||
|
||||
Print this, then continue immediately:
|
||||
|
||||
```
|
||||
=== Autoresearch progress — iteration N ===
|
||||
Baseline: <value>
|
||||
Current best: <value> (<delta> from baseline)
|
||||
Keeps: <count>
|
||||
Discards: <count>
|
||||
Crashes: <count>
|
||||
Top pattern: <what has worked most consistently>
|
||||
Last 5: <keep/discard/crash sequence>
|
||||
===
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Lessons system
|
||||
|
||||
After every 5 KEPT iterations, append to `autoresearch-lessons.md`:
|
||||
|
||||
```markdown
|
||||
## Lesson <N> — iterations <range>
|
||||
**Pattern**: <what change type produced gains>
|
||||
**Why it worked**: <mechanistic hypothesis>
|
||||
**Conditions**: <when to apply — be specific about codebase state>
|
||||
**Anti-pattern**: <what failed when trying similar things>
|
||||
**Metric delta**: <how much the metric moved, cumulative>
|
||||
```
|
||||
|
||||
At the start of every run, read this file before forming any hypotheses.
|
||||
Weight recent lessons more heavily. Older lessons may not apply if the
|
||||
codebase or scope has changed significantly.
|
||||
|
||||
This is the compounding mechanism. Each overnight run starts smarter than
|
||||
the last.
|
||||
|
||||
---
|
||||
|
||||
## Stuck recovery
|
||||
|
||||
After 5 consecutive discards or crashes:
|
||||
|
||||
1. Re-read all Scope files from scratch. Full context, not memory.
|
||||
2. Search the lessons log for near-misses — what came closest to working?
|
||||
3. Try combining two near-miss approaches into one hypothesis.
|
||||
4. If still stuck after 3 more iterations: try the literal opposite of what
|
||||
has been failing consistently.
|
||||
5. If still stuck after 3 more: use `WebSearch` to research the
|
||||
problem space. Search for `[domain] [metric] improvement techniques [year]`.
|
||||
Extract 3 concrete techniques. Use each as the next 3 hypotheses.
|
||||
6. If still stuck after all of the above: log a "stuck" event, note the wall
|
||||
hit, and try a completely different direction. Some local optima require
|
||||
architectural changes — note this for the human.
|
||||
|
||||
---
|
||||
|
||||
## Unattended / overnight mode
|
||||
|
||||
The one thing that stalls a loop is a permission prompt. Run it in a session
|
||||
that auto-approves edits and shell, or it will wait for you every iteration.
|
||||
|
||||
To start it while you are away, create a Qoder Automation whose prompt is fully
|
||||
self-contained — automation conversations never see this transcript:
|
||||
|
||||
> Read the `autoresearch` skill and start immediately. Goal: `<goal>`.
|
||||
> Scope: `<scope>`. Metric: `<metric — higher/lower is better>`.
|
||||
> Verify: `<command>`. Guard: `<command>`. Do not pause, do not ask questions,
|
||||
> iterate until stopped.
|
||||
|
||||
You will wake up to `autoresearch-results.tsv` and `autoresearch-lessons.md`.
|
||||
Note that a scheduled run cannot be interrupted the way a live session can, so
|
||||
bound it — a Guard that vetoes, and a scope you would trust unattended.
|
||||
|
||||
---
|
||||
|
||||
## Non-negotiable rules
|
||||
|
||||
1. **NEVER STOP** until the user manually interrupts the run.
|
||||
2. **ONE change per iteration** — atomic, explainable in one sentence.
|
||||
3. **Mechanical verification only** — no "looks better", no "seems cleaner".
|
||||
If you cannot measure it, you cannot use it as a signal.
|
||||
4. **Commit BEFORE verifying** — always. No exceptions.
|
||||
5. **Auto-revert on regression** — no debate, no "let me try one more thing".
|
||||
6. **Guard is a hard veto** — Verify passing does not mean KEEP. Guard must also pass.
|
||||
7. **Never modify Guard files** — they are read-only invariants, not scope.
|
||||
8. **Read git history before every hypothesis** — it is your short-term memory.
|
||||
9. **Read lessons before every run** — it is your long-term memory.
|
||||
10. **Simplicity wins ties** — equal metric + less code = KEEP.
|
||||
11. **Never touch files outside Scope** — discipline is what makes the loop safe.
|
||||
12. **When in doubt, make the smaller change** — scope creep kills iterations.
|
||||
|
||||
---
|
||||
|
||||
## Reference files
|
||||
|
||||
**Core loop**
|
||||
- `references/loop-protocol.md` — detailed phase-by-phase protocol
|
||||
- `references/results-logging.md` — TSV format, summary templates, examples
|
||||
- `references/lessons-system.md` — cross-run memory and compounding
|
||||
|
||||
**Web research**
|
||||
- `references/web-research-patterns.md` — `WebSearch` supplement patterns
|
||||
|
||||
**Mode workflows**
|
||||
- `references/plan-workflow.md` — `plan` mode — auto-detect and configure
|
||||
- `references/ship-workflow.md` — `ship` mode — pre-flight checklist
|
||||
- `references/debug-workflow.md` — `debug` mode — root cause and fix
|
||||
- `references/fix-workflow.md` — `fix` mode — focused type/lint fix
|
||||
- `references/security-workflow.md` — `security` mode — STRIDE/OWASP audit
|
||||
@@ -0,0 +1,25 @@
|
||||
# `autoresearch debug` mode — Autonomous Debug Loop
|
||||
|
||||
This workflow is triggered by the `debug` mode. It is designed to reproduce, isolate, and fix specific bugs autonomously.
|
||||
|
||||
## Context
|
||||
Use this when something is clearly broken (e.g., a failing test, a crash, or a UI bug).
|
||||
|
||||
## Phase 1: Reproduction
|
||||
1. Create a minimal reproduction script (e.g., `debug/repro.js` or a new test case).
|
||||
2. Run the repro script and verify it fails as expected.
|
||||
3. This repro command becomes your `Verify` command for the loop.
|
||||
|
||||
## Phase 2: Isolation
|
||||
1. Use `Grep` and `Read` to find the code responsible for the failure.
|
||||
2. Form a hypothesis about the root cause.
|
||||
|
||||
## Phase 3: Fix Loop
|
||||
1. Start a standard autoresearch loop with:
|
||||
- **Goal**: Fix the bug identified in the repro script.
|
||||
- **Verify**: The repro command (must exit 0 on success).
|
||||
- **Guard**: Existing test suite and linting.
|
||||
|
||||
## Phase 4: Hardening
|
||||
1. After the fix is verified, add a permanent regression test to the codebase.
|
||||
2. Verify that the fix holds across the entire project.
|
||||
@@ -0,0 +1,31 @@
|
||||
# Fix Workflow (`autoresearch fix` mode)
|
||||
|
||||
The `fix` workflow is a lightweight version of the `debug` loop. It is designed for situations where you have a specific, known failure (e.g., a TypeScript error or a lint violation) and you want to fix it without the overhead of full reproduction and isolation.
|
||||
|
||||
## Protocol
|
||||
|
||||
### 1. Context Loading
|
||||
* Read the error message or description provided in the command.
|
||||
* Identify the affected file(s).
|
||||
* Read the current state of those files.
|
||||
|
||||
### 2. Hypothesis
|
||||
* Form a direct hypothesis on how to fix the specific error.
|
||||
* The fix must be minimal and targeted.
|
||||
|
||||
### 3. Execution
|
||||
* Apply the fix.
|
||||
* Commit the change.
|
||||
|
||||
### 4. Verification
|
||||
* Run the command that triggered the original failure (e.g., `npx tsc` or `npm run lint`).
|
||||
* If a `Guard` is set in the main autoresearch config, run that as well.
|
||||
|
||||
### 5. Decision
|
||||
* If the error is gone and Guard passes: **KEEP**.
|
||||
* If the error persists: **RETRY** (max 3 times) with a different approach.
|
||||
* If it still fails after 3 tries: **REVERT** and report to the user.
|
||||
|
||||
## When to use `fix` vs `debug`
|
||||
* Use **`fix`** for mechanical errors: "Fix the lint error on line 42", "Fix the missing import in `utils.ts`".
|
||||
* Use **`debug`** for logical errors: "The login flow fails for users with specialized characters", "Database connection timeouts under high load".
|
||||
@@ -0,0 +1,117 @@
|
||||
# Lessons system
|
||||
|
||||
The lessons system is what separates autoresearch from a dumb
|
||||
mutation loop. It is the mechanism by which each overnight run starts
|
||||
smarter than the last.
|
||||
|
||||
---
|
||||
|
||||
## The compounding model
|
||||
|
||||
```
|
||||
Night 1: 100 experiments → lessons-v1 written
|
||||
Night 2: reads lessons-v1 → avoids 20 known failures → 80 net-new experiments
|
||||
Night 3: reads lessons-v2 → avoids 35 known failures → faster convergence
|
||||
...
|
||||
```
|
||||
|
||||
Without the lessons system, every run starts from scratch. With it, runs
|
||||
compound — each failure is learned once and never repeated.
|
||||
|
||||
---
|
||||
|
||||
## File location and format
|
||||
|
||||
File: `autoresearch-lessons.md` in your project root.
|
||||
|
||||
Add to `.gitignore` — this is a working file for the agent, not source code.
|
||||
|
||||
```markdown
|
||||
# Autoresearch lessons — <project name>
|
||||
Generated by the autoresearch skill. Do not edit manually.
|
||||
Last updated: <ISO date>
|
||||
|
||||
## Lesson 1 — iterations 1–5
|
||||
**Pattern**: <the type of change that produced gains>
|
||||
**Why it worked**: <mechanistic hypothesis — be specific>
|
||||
**Conditions**: <codebase state where this applies>
|
||||
**Anti-pattern**: <what failed when trying similar approaches>
|
||||
**Metric delta**: <cumulative gain from this pattern, e.g. "+4.2%">
|
||||
|
||||
## Lesson 2 — iterations 6–10
|
||||
...
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## When to write lessons
|
||||
|
||||
Append a new lesson after every 5 KEPT iterations (not every 5 total
|
||||
iterations). Lessons should only describe what worked.
|
||||
|
||||
Failed patterns are captured implicitly — if a pattern never generates a
|
||||
kept iteration, it never generates a lesson, and the loop naturally
|
||||
deprioritises it via Phase 2's "different from last 3 attempts" rule.
|
||||
|
||||
---
|
||||
|
||||
## What makes a good lesson
|
||||
|
||||
**Good** (specific, mechanistic, conditional):
|
||||
```
|
||||
**Pattern**: Defer non-critical third-party scripts using loading="lazy"
|
||||
**Why it worked**: Removes scripts from the critical render path, reducing
|
||||
Time to Interactive without affecting functionality
|
||||
**Conditions**: Applies to analytics, chat widgets, social embeds — not
|
||||
to scripts required for initial page render
|
||||
**Anti-pattern**: Lazy-loading scripts that are called in the first 500ms
|
||||
of page load caused layout shifts and broke interactions
|
||||
**Metric delta**: +6.8% Lighthouse performance score across 3 iterations
|
||||
```
|
||||
|
||||
**Bad** (vague, not actionable):
|
||||
```
|
||||
**Pattern**: Make things faster
|
||||
**Why it worked**: It improved performance
|
||||
**Conditions**: When performance is bad
|
||||
**Anti-pattern**: When it makes things worse
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## How to read lessons at the start of a run
|
||||
|
||||
1. Read the full file — do not skip old lessons even if they seem stale.
|
||||
2. For each lesson, assess: does this pattern still apply given the current
|
||||
state of the codebase? If the code it describes has been significantly
|
||||
refactored, downweight it.
|
||||
3. Extract the top 2-3 highest-delta patterns. These are your first
|
||||
hypotheses unless the results log shows they have already been exhausted.
|
||||
4. Extract the anti-patterns. These are your first exclusions — do not
|
||||
generate hypotheses that match these patterns.
|
||||
|
||||
---
|
||||
|
||||
## Cross-project lessons
|
||||
|
||||
For teams running autoresearch across multiple similar projects (e.g.
|
||||
multiple Next.js apps), consider maintaining a shared lessons file at
|
||||
`~/.autoresearch/global-lessons.md`.
|
||||
|
||||
At the start of a run, read both the project-level and global lessons.
|
||||
Project-level lessons take precedence when they conflict with global ones.
|
||||
|
||||
This is optional but significantly accelerates convergence on new projects
|
||||
that share a tech stack with already-researched ones.
|
||||
|
||||
---
|
||||
|
||||
## Lessons file maintenance
|
||||
|
||||
- Do not manually edit the lessons file during a run — the agent reads it
|
||||
at the start of each run and its contents influence hypothesis generation.
|
||||
- After a long run (100+ iterations), review the file and remove lessons
|
||||
that are no longer applicable (e.g. they describe code that no longer
|
||||
exists). Add a comment explaining why the lesson was removed.
|
||||
- The lessons file is cumulative — never delete lessons, only annotate them
|
||||
as superseded if a newer lesson contradicts them.
|
||||
@@ -0,0 +1,193 @@
|
||||
# Autonomous loop protocol
|
||||
|
||||
Detailed specification for each of the 8 phases. The SKILL.md contains the
|
||||
summary version. Read this reference when you need precise guidance on edge
|
||||
cases in any phase.
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Review
|
||||
|
||||
**Purpose**: Build a complete, accurate picture of current state before
|
||||
forming any hypothesis. Hypotheses formed without full context waste iterations.
|
||||
|
||||
**What to read**:
|
||||
- Every file in Scope (not just the ones you last touched)
|
||||
- `git log --oneline -20` — what has been attempted, in order
|
||||
- `autoresearch-results.tsv` — the full record of what worked and failed
|
||||
- `autoresearch-lessons.md` — accumulated patterns from prior runs
|
||||
|
||||
**What to extract**:
|
||||
- Current metric trajectory (improving? plateauing? volatile?)
|
||||
- Which change types produced the most gain per iteration
|
||||
- Which change types consistently failed
|
||||
- Which directions have not yet been explored
|
||||
- Any patterns in crash causes
|
||||
|
||||
**Duration**: This phase should take as long as needed to form a genuinely
|
||||
informed hypothesis. Rushing Phase 1 leads to repeated failures.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — Ideate
|
||||
|
||||
**Purpose**: Select ONE hypothesis that has the highest expected gain given
|
||||
what is known.
|
||||
|
||||
**Hypothesis selection criteria** (in order of priority):
|
||||
1. Builds directly on a proven pattern from the lessons file
|
||||
2. Explores a direction adjacent to a near-miss (something that almost worked)
|
||||
3. Combines two near-miss approaches that individually failed
|
||||
4. Tries the opposite of what consistently failed
|
||||
5. Applies an externally validated technique (from `WebSearch` research)
|
||||
6. Tries something entirely untested
|
||||
|
||||
**What makes a good hypothesis**:
|
||||
- Specific: "lazy-load the user avatar component" not "improve performance"
|
||||
- Testable: produces a measurable delta in the Verify command
|
||||
- Atomic: one thing changes, one thing is measured
|
||||
- Explainable in one sentence before you make the change
|
||||
|
||||
**What makes a bad hypothesis**:
|
||||
- Vague: "refactor for clarity"
|
||||
- Multi-part: "update the API, add caching, and fix the tests"
|
||||
- Untestable by the Verify command
|
||||
- Identical to something tried in the last 3 iterations
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — Modify
|
||||
|
||||
**Purpose**: Implement the hypothesis as a single, clean, minimal change.
|
||||
|
||||
**Rules**:
|
||||
- Touch only files in Scope
|
||||
- Make the smallest change that tests the hypothesis
|
||||
- If the change is getting large, stop and split it — make the first half now,
|
||||
the second half in the next iteration
|
||||
- Do not fix unrelated things you notice while editing
|
||||
- Do not reformat code that is not part of the hypothesis
|
||||
- Leave comments only if they directly explain the change
|
||||
|
||||
**Signs you are over-scoping**:
|
||||
- You have edited more than 3 files
|
||||
- The diff is more than ~50 lines
|
||||
- You are explaining the change with "and also"
|
||||
|
||||
When in doubt, make a smaller change. Smaller changes fail faster and teach more.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — Commit
|
||||
|
||||
**Purpose**: Create a clean rollback point before any verification risk.
|
||||
|
||||
**Command**:
|
||||
```bash
|
||||
git add -A && git commit -m "autoresearch iter N: <one-sentence description>"
|
||||
```
|
||||
|
||||
**Commit message format**:
|
||||
- Always prefix with `autoresearch iter N:`
|
||||
- One sentence, present tense, describes the change not the goal
|
||||
- Good: `autoresearch iter 14: lazy-load user avatar to reduce initial bundle`
|
||||
- Bad: `autoresearch iter 14: improve performance`
|
||||
|
||||
**Why commit before verifying**: if the Verify command crashes, hangs, or
|
||||
corrupts state, you can always `git revert HEAD --no-edit` and return to
|
||||
a known-good state. If you verify before committing, a crash during
|
||||
verification leaves you with uncommitted changes and an unknown baseline.
|
||||
|
||||
**Never skip this step**, even if the change feels obviously correct.
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — Verify
|
||||
|
||||
**Purpose**: Get a single numeric measurement of whether the hypothesis helped.
|
||||
|
||||
**Execution**:
|
||||
1. Run the Verify command exactly as specified by the user
|
||||
2. Extract the numeric metric value
|
||||
3. Optionally supplement with `WebSearch` research (see
|
||||
`references/web-research-patterns.md`)
|
||||
4. Record the raw output for the log
|
||||
|
||||
**Handling slow Verify commands**:
|
||||
If the Verify command takes more than 30 seconds, note this. After the run,
|
||||
recommend the user find a faster proxy metric — slower verification means
|
||||
fewer experiments per hour, which compounds negatively over a full night.
|
||||
|
||||
**Handling non-deterministic Verify commands**:
|
||||
If the metric varies significantly between runs on identical code (>5%
|
||||
variance), note this in the log. Run the Verify command twice and average.
|
||||
Log both values. Recommend the user address flakiness before the next
|
||||
overnight run.
|
||||
|
||||
---
|
||||
|
||||
## Phase 6 — Decide
|
||||
|
||||
**Purpose**: Make a clear, mechanical keep/revert decision. No deliberation.
|
||||
|
||||
**Decision table**:
|
||||
|
||||
| Condition | Action | Log status |
|
||||
|---|---|---|
|
||||
| Metric improved (beyond noise threshold) | Keep commit as-is | `keep` |
|
||||
| Metric unchanged or regressed | `git revert HEAD --no-edit` | `discard` |
|
||||
| Verify crashed with exit code ≠ 0 | Attempt fix (max 3 tries) then revert | `crash` |
|
||||
| Verify hung for >60s | Kill process, revert | `crash` |
|
||||
|
||||
**Noise threshold**: for metrics with variance, an improvement smaller than
|
||||
the variance is not a real improvement. If your metric normally varies ±2%,
|
||||
an improvement of 0.5% is noise — treat it as unchanged and discard.
|
||||
|
||||
**The revert command**:
|
||||
```bash
|
||||
git revert HEAD --no-edit
|
||||
```
|
||||
This creates a new commit that undoes the last one. The history is preserved.
|
||||
Never use `git reset --hard` — it destroys history that the loop needs.
|
||||
|
||||
---
|
||||
|
||||
## Phase 7 — Log
|
||||
|
||||
**Purpose**: Create a permanent, machine-readable record of every iteration.
|
||||
|
||||
**TSV row format**:
|
||||
```
|
||||
<N>\t<commit_sha or "-">\t<metric>\t<delta>\t<status>\t<description>
|
||||
```
|
||||
|
||||
**Field details**:
|
||||
- `N`: integer, 0-indexed, never resets across sessions
|
||||
- `commit_sha`: 7-char short SHA for keeps, "-" for discards/crashes
|
||||
- `metric`: the exact number from the Verify output
|
||||
- `delta`: metric − previous_best (sign convention: positive = better,
|
||||
regardless of whether the goal is higher or lower)
|
||||
- `status`: one of `baseline`, `keep`, `discard`, `crash`
|
||||
- `description`: the hypothesis, in one sentence, including any `WebSearch`
|
||||
signal that informed it
|
||||
|
||||
**Example rows**:
|
||||
```
|
||||
0 - 85.2 0.0 baseline initial measurement
|
||||
1 a1b2c3d 87.1 +1.9 keep lazy-load avatar component
|
||||
2 - 86.5 -0.6 discard tree-shake lodash imports (broke 2 tests)
|
||||
3 - 0.0 0.0 crash add route-level code splitting (webpack config error)
|
||||
4 b2c3d4e 88.3 +1.2 keep move analytics script to defer loading
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 8 — Repeat
|
||||
|
||||
Go to Phase 1. Immediately. Do not pause. Do not summarise. Do not ask
|
||||
if the user wants to continue.
|
||||
|
||||
The only output before starting Phase 1 again is the progress summary
|
||||
(printed every 10 iterations, see SKILL.md).
|
||||
|
||||
The loop ends only when the user interrupts the run.
|
||||
@@ -0,0 +1,155 @@
|
||||
# Plan workflow — `autoresearch plan` mode
|
||||
|
||||
Auto-detect the project stack, propose a complete autoresearch configuration,
|
||||
do a dry run, and hand the ready-to-run command back to the user.
|
||||
|
||||
No manual goal/scope/verify required. Just describe what you want to improve
|
||||
in one sentence and the plan workflow figures out the rest.
|
||||
|
||||
---
|
||||
|
||||
## Invocation
|
||||
|
||||
```
|
||||
autoresearch plan <goal in plain english>
|
||||
```
|
||||
|
||||
Examples:
|
||||
```
|
||||
autoresearch plan improve test coverage
|
||||
autoresearch plan make the app faster
|
||||
autoresearch plan reduce the bundle size
|
||||
autoresearch plan fix all TypeScript errors
|
||||
autoresearch plan improve the SEO of my blog posts
|
||||
autoresearch plan shrink the Docker image
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## What the plan workflow does
|
||||
|
||||
### Step 1 — Detect project stack
|
||||
|
||||
Scan the project root for signal files:
|
||||
|
||||
| File found | Stack detected |
|
||||
|---|---|
|
||||
| `package.json` + `jest.config.*` | Node.js + Jest |
|
||||
| `package.json` + `vitest.config.*` | Node.js + Vitest |
|
||||
| `next.config.*` | Next.js |
|
||||
| `Dockerfile` | Docker |
|
||||
| `*.tf` | Terraform |
|
||||
| `.github/workflows/*.yml` | GitHub Actions CI |
|
||||
| `content/blog/*.md` OR `posts/*.md` | Markdown content/blog |
|
||||
| `src/**/*.ts` OR `src/**/*.tsx` | TypeScript project |
|
||||
| `pyproject.toml` OR `setup.py` | Python project |
|
||||
| `requirements.txt` + `pytest` | Python + pytest |
|
||||
| `go.mod` | Go project |
|
||||
| `Cargo.toml` | Rust project |
|
||||
|
||||
Print detected stack. If ambiguous, list the top two candidates and ask
|
||||
the user to confirm before proceeding.
|
||||
|
||||
### Step 2 — Map goal to metric + verify command
|
||||
|
||||
Use the goal description and detected stack to propose:
|
||||
|
||||
| Goal keyword | Metric | Verify command template |
|
||||
|---|---|---|
|
||||
| "test coverage" | coverage % (higher is better) | `npm test -- --coverage \| grep "All files"` |
|
||||
| "bundle size" / "build size" | size in KB (lower is better) | `npm run build 2>&1 \| grep "First Load JS"` |
|
||||
| "TypeScript errors" / "type errors" | error count (lower is better) | `npx tsc --noEmit 2>&1 \| grep -c "error TS" \|\| echo "0"` |
|
||||
| "lighthouse" / "performance score" | score 0-100 (higher is better) | `npx lighthouse http://localhost:3000 --output json --quiet 2>/dev/null \| jq '.categories.performance.score * 100'` |
|
||||
| "docker image" / "image size" | size in MB (lower is better) | `docker build -t bench . -q && docker images bench --format "{{.Size}}"` |
|
||||
| "flaky tests" | failure count (lower is better) | `for i in {1..5}; do npm test 2>&1; done \| grep -c "FAIL" \|\| echo "0"` |
|
||||
| "SEO" / "blog" / "content" | SEO score (higher is better) | `node scripts/seo-score.js <detected content path>` |
|
||||
| "lines of code" / "complexity" | LOC count (lower is better) | `find src/ -name "*.ts" \| xargs wc -l \| tail -1 \| awk '{print $1}'` |
|
||||
| "CI pipeline" / "pipeline speed" | seconds (lower is better) | `node scripts/estimate-ci-time.js` |
|
||||
| "Python tests" / "pytest" | coverage % (higher is better) | `pytest --cov=src --cov-report=term-missing \| grep "TOTAL"` |
|
||||
| "faster" / "performance" / "latency" | p95 ms (lower is better) | `npm run bench 2>&1 \| grep "p95"` |
|
||||
|
||||
### Step 3 — Detect scope
|
||||
|
||||
Based on goal + stack, propose the tightest scope that covers the goal:
|
||||
|
||||
- Test coverage → `src/**/*.ts, src/**/*.test.ts`
|
||||
- Bundle size → `src/**/*.tsx, src/**/*.ts`
|
||||
- Docker → `Dockerfile, .dockerignore`
|
||||
- SEO → `content/blog/*.md` or detected content directory
|
||||
- TypeScript errors → `src/**/*.ts`
|
||||
- CI pipeline → `.github/workflows/*.yml`
|
||||
|
||||
### Step 4 — Dry run
|
||||
|
||||
Run the proposed Verify command once against the current state.
|
||||
|
||||
- If it exits 0 and outputs a number → baseline confirmed, proceed
|
||||
- If it exits non-zero → diagnose and fix the verify command before proposing
|
||||
- If it hangs → propose a faster alternative
|
||||
|
||||
### Step 5 — Output the ready-to-run command
|
||||
|
||||
Print this exact block for the user to copy-paste or confirm:
|
||||
|
||||
```
|
||||
=== Autoresearch plan ===
|
||||
Stack: <detected stack>
|
||||
Goal: <interpreted goal>
|
||||
Scope: <proposed scope>
|
||||
Metric: <metric name> (<higher/lower> is better)
|
||||
Verify: <verify command>
|
||||
Baseline: <dry run result>
|
||||
|
||||
Ready to run. Confirm or adjust any field, then:
|
||||
|
||||
/autoresearch
|
||||
Goal: <goal>
|
||||
Scope: <scope>
|
||||
Metric: <metric>
|
||||
Verify: <verify command>
|
||||
|
||||
Or, for an unattended run, put these same fields into a Qoder Automation prompt
|
||||
(see "Unattended / overnight mode" in SKILL.md).
|
||||
===
|
||||
```
|
||||
|
||||
If the user says "looks good" or "run it" — start the autoresearch loop
|
||||
immediately without requiring them to retype the command.
|
||||
|
||||
---
|
||||
|
||||
## Web research calibration
|
||||
|
||||
After the dry run, use `WebSearch` to calibrate:
|
||||
|
||||
- For SEO goals: search for `[target keyword]` to see what top results look like.
|
||||
Note any structural patterns (FAQ sections, word count, heading structure)
|
||||
that the current content lacks. Add these as initial hypotheses.
|
||||
|
||||
- For performance goals: search for `[framework] performance benchmarks [year]`
|
||||
to calibrate whether the baseline is already good or has significant headroom.
|
||||
|
||||
- For security goals: search for `[stack] common vulnerabilities [year]`
|
||||
to seed the initial hypothesis pool with known attack vectors.
|
||||
|
||||
This research step happens during plan, not during the loop — so it adds
|
||||
context once without slowing down iterations.
|
||||
|
||||
---
|
||||
|
||||
## Edge cases
|
||||
|
||||
**Goal is too vague** ("make it better"):
|
||||
Ask one clarifying question: "Better in what way — speed, quality, size,
|
||||
coverage, or something else?" Then proceed.
|
||||
|
||||
**Multiple valid verify commands exist**:
|
||||
Propose the fastest one. Note the slower alternative in a comment.
|
||||
|
||||
**Verify command requires a running server**:
|
||||
Note this in the plan output. Add a `# requires: local server on :3000`
|
||||
comment. Suggest the user start it before running the loop.
|
||||
|
||||
**No matching stack detected**:
|
||||
Ask the user to describe their stack in one sentence, then proceed with
|
||||
a custom verify command.
|
||||
@@ -0,0 +1,105 @@
|
||||
# Results logging
|
||||
|
||||
Specification for `autoresearch-results.tsv` — the per-iteration record
|
||||
of every experiment in a run.
|
||||
|
||||
---
|
||||
|
||||
## File format
|
||||
|
||||
Tab-separated values. Headers on row 1. One row per iteration.
|
||||
|
||||
```
|
||||
iteration\tcommit\tmetric\tdelta\tstatus\tdescription
|
||||
```
|
||||
|
||||
### Field definitions
|
||||
|
||||
| Field | Type | Description |
|
||||
|---|---|---|
|
||||
| `iteration` | integer | 0-indexed. Never resets — if you run multiple sessions, continue from the last number. |
|
||||
| `commit` | string | 7-char git short SHA for kept commits. `-` for discards and crashes. |
|
||||
| `metric` | float | Raw metric value from the Verify command. |
|
||||
| `delta` | float | `metric − previous_best`. Sign convention: positive = improvement (regardless of higher/lower goal). |
|
||||
| `status` | enum | One of: `baseline`, `keep`, `discard`, `crash` |
|
||||
| `description` | string | The hypothesis, one sentence. Include the change type and the expected mechanism. |
|
||||
|
||||
---
|
||||
|
||||
## Example file
|
||||
|
||||
```tsv
|
||||
iteration commit metric delta status description
|
||||
0 - 85.2 0.0 baseline initial measurement — test coverage 85.2%
|
||||
1 a1b2c3d 87.1 +1.9 keep add tests for auth middleware edge cases
|
||||
2 - 86.5 -0.7 discard refactor test helpers (broke 2 existing tests)
|
||||
3 - 0.0 0.0 crash add integration tests (postgres connection failed — fix in iter 4)
|
||||
4 b2c3d4e 88.3 +1.2 keep add tests for error handling in API routes
|
||||
5 - 88.1 -0.2 discard add tests for rate limiter (metric within variance, treated as regression)
|
||||
6 c3d4e5f 89.0 +0.7 keep add boundary value tests for form validators
|
||||
7 d4e5f6g 89.8 +0.8 keep add tests for session expiry edge cases
|
||||
8 - 89.2 -0.6 discard mock external API calls (test isolation but metric regressed)
|
||||
9 e5f6g7h 90.6 +0.8 keep add tests for concurrent request handling
|
||||
10 f6g7h8i 91.1 +0.5 keep add tests for malformed JSON input handling
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Progress summary format
|
||||
|
||||
Print every 10 iterations. Use this exact format:
|
||||
|
||||
```
|
||||
=== Autoresearch progress — iteration <N> ===
|
||||
Goal: <original goal statement>
|
||||
Baseline: <iteration 0 metric>
|
||||
Current best: <best metric so far> (<total delta> from baseline)
|
||||
Keeps: <count> (<keeps/total * 100>%)
|
||||
Discards: <count>
|
||||
Crashes: <count>
|
||||
Top pattern: <the change type that has produced the most total delta>
|
||||
Last 5: <sequence of keep/discard/crash for iterations N-4 through N>
|
||||
Est. to goal: <if goal metric is known, N iterations at current rate>
|
||||
===
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Interpreting the log
|
||||
|
||||
### Healthy run signature
|
||||
- Keep rate 40-60%
|
||||
- Delta per keep: consistent small positive gains
|
||||
- No long crash streaks
|
||||
- Discards are evenly distributed (not clustered)
|
||||
|
||||
### Warning signs
|
||||
|
||||
| Pattern | Meaning | Action |
|
||||
|---|---|---|
|
||||
| Keep rate < 20% | Hypothesis quality is poor | Re-read full scope, re-read lessons, change direction |
|
||||
| Keep rate > 80% | Metric may be too easy or Verify too lenient | Tighten the goal |
|
||||
| Long crash streak (5+) | Verify command is fragile or scope is too risky | Fix Verify or narrow scope |
|
||||
| Delta per keep shrinking toward 0 | Approaching local optimum | Try more radical changes or declare victory |
|
||||
| Metric oscillating | Non-deterministic Verify or contradictory changes | Run Verify twice and average; tighten scope |
|
||||
|
||||
### Declaring success
|
||||
|
||||
Stop the loop when one of these is true:
|
||||
- Metric has reached the stated goal
|
||||
- Delta per keep has been below 0.1% for 20 consecutive iterations
|
||||
(local optimum with current scope)
|
||||
- All directions have been exhausted (lessons file confirms this)
|
||||
|
||||
In all cases, print a final summary and write a lessons entry covering
|
||||
the full run before stopping.
|
||||
|
||||
---
|
||||
|
||||
## File hygiene
|
||||
|
||||
- Add `autoresearch-results.tsv` to `.gitignore`. It is a working file.
|
||||
- Do not edit it manually during a run.
|
||||
- Between runs, you may archive it:
|
||||
`mv autoresearch-results.tsv autoresearch-results-<date>.tsv`
|
||||
and start fresh, but keep the lessons file — that is the persistent memory.
|
||||
@@ -0,0 +1,171 @@
|
||||
# Security workflow — `autoresearch security` mode
|
||||
|
||||
Autonomous security audit using STRIDE threat modelling and OWASP categories.
|
||||
Finds vulnerabilities, classifies them by severity, and optionally fixes
|
||||
confirmed critical and high findings via an autoresearch loop.
|
||||
|
||||
---
|
||||
|
||||
## Invocation
|
||||
|
||||
```
|
||||
autoresearch security # full audit, report only
|
||||
autoresearch security --fix # audit + auto-fix confirmed findings
|
||||
autoresearch security --fail-on critical # end with a FAIL verdict if critical found
|
||||
autoresearch security --scope src/api/ # audit a specific directory only
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 1 — Asset discovery
|
||||
|
||||
Map the attack surface:
|
||||
|
||||
1. Identify all entry points: API routes, form handlers, file uploads,
|
||||
auth flows, webhooks, admin panels
|
||||
2. Identify all data stores: databases, caches, file system writes,
|
||||
environment variables, secrets
|
||||
3. Identify all trust boundaries: public vs authenticated, user vs admin,
|
||||
internal vs external services
|
||||
4. Map data flows: what user input reaches what data store via what path
|
||||
|
||||
Output: `security/audit-<timestamp>/attack-surface-map.md`
|
||||
|
||||
### Live threat intelligence
|
||||
|
||||
Use `WebSearch` to seed the audit with current threats:
|
||||
|
||||
```
|
||||
WebSearch: [your stack] common vulnerabilities [current year]
|
||||
WebSearch: [your main framework] CVE [current year]
|
||||
WebSearch: OWASP top 10 [current year]
|
||||
```
|
||||
|
||||
Add any newly discovered attack patterns to the audit queue.
|
||||
This ensures the audit covers threats that postdate your static analysis tools.
|
||||
|
||||
---
|
||||
|
||||
## Phase 2 — STRIDE threat model
|
||||
|
||||
For each asset and trust boundary, model threats across all 6 STRIDE categories:
|
||||
|
||||
| Category | Question to ask |
|
||||
|---|---|
|
||||
| **S**poofing | Can an attacker impersonate a user, service, or system? |
|
||||
| **T**ampering | Can input be modified to alter data or behaviour unexpectedly? |
|
||||
| **R**epudiation | Can actions be performed without a traceable audit trail? |
|
||||
| **I**nformation disclosure | Can sensitive data be accessed by unauthorised parties? |
|
||||
| **D**enial of service | Can the service be made unavailable through normal inputs? |
|
||||
| **E**levation of privilege | Can a lower-privilege user gain higher-privilege access? |
|
||||
|
||||
Output: `security/audit-<timestamp>/threat-model.md`
|
||||
|
||||
---
|
||||
|
||||
## Phase 3 — Autonomous audit loop
|
||||
|
||||
```
|
||||
LOOP (through all attack vectors from threat model):
|
||||
1. Select next untested attack vector
|
||||
2. Deep-dive into the relevant code (read fully — do not skim)
|
||||
3. Attempt to construct a concrete exploit scenario
|
||||
4. Validate with code evidence (file:line + exact scenario)
|
||||
5. Classify: severity + OWASP category + STRIDE tag
|
||||
6. Log to security-audit-results.tsv
|
||||
7. Print coverage summary every 5 iterations
|
||||
8. Continue until all vectors tested
|
||||
```
|
||||
|
||||
### Severity classification
|
||||
|
||||
| Severity | Definition |
|
||||
|---|---|
|
||||
| Critical | Exploitable without authentication, leads to full compromise or data breach |
|
||||
| High | Exploitable with low-privilege access, significant impact |
|
||||
| Medium | Requires specific conditions, moderate impact |
|
||||
| Low | Minor information disclosure, no direct exploitation path |
|
||||
| Info | Best practice violation, no immediate security impact |
|
||||
|
||||
### Evidence requirement
|
||||
|
||||
Every finding MUST have:
|
||||
- File path and line number
|
||||
- Exact vulnerable code snippet (copy from source, do not paraphrase)
|
||||
- Concrete exploit scenario (how an attacker would trigger this)
|
||||
- Proof of exploitability (not theoretical — show the actual path)
|
||||
|
||||
Findings without concrete evidence are logged as "unconfirmed" and flagged
|
||||
for manual review, not included in the fix loop.
|
||||
|
||||
---
|
||||
|
||||
## Phase 4 — Report generation
|
||||
|
||||
Output folder: `security/audit-<timestamp>/`
|
||||
|
||||
```
|
||||
security/audit-20260325-1430/
|
||||
├── overview.md ← executive summary + finding counts by severity
|
||||
├── threat-model.md ← STRIDE analysis per asset
|
||||
├── attack-surface-map.md ← entry points, data flows, trust boundaries
|
||||
├── findings.md ← all confirmed findings, sorted by severity
|
||||
├── owasp-coverage.md ← coverage matrix — which OWASP categories checked
|
||||
├── recommendations.md ← fix guidance for each confirmed finding
|
||||
└── security-audit-results.tsv ← machine-readable log of all iterations
|
||||
```
|
||||
|
||||
Print summary:
|
||||
```
|
||||
=== Security audit summary ===
|
||||
Critical: <N>
|
||||
High: <N>
|
||||
Medium: <N>
|
||||
Low: <N>
|
||||
Info: <N>
|
||||
Vectors tested: <N> / <total>
|
||||
OWASP categories covered: <list>
|
||||
|
||||
Full report: security/audit-<timestamp>/overview.md
|
||||
===
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Phase 5 — Auto-fix loop (with `--fix`)
|
||||
|
||||
Only runs when `--fix` flag is passed.
|
||||
Only fixes **Confirmed Critical and High** findings.
|
||||
Uses `recommendations.md` as the fix guide for each finding.
|
||||
|
||||
```
|
||||
FOR EACH confirmed Critical/High finding:
|
||||
1. Read the finding + recommendation
|
||||
2. Make ONE targeted fix
|
||||
3. git commit the fix
|
||||
4. Re-run the specific exploit scenario to verify it no longer works
|
||||
5. Run full test suite to confirm no regressions
|
||||
6. If tests break → revert, try alternative fix
|
||||
7. Maximum 3 attempts per finding, then skip and flag for manual review
|
||||
8. Log fix outcome to fix-log.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Verdict mode (`--fail-on`)
|
||||
|
||||
```
|
||||
autoresearch security --fail-on critical
|
||||
```
|
||||
|
||||
The audit ends with an explicit verdict line in `overview.md`:
|
||||
|
||||
```
|
||||
VERDICT: FAIL — 2 findings at or above `critical`
|
||||
VERDICT: PASS — no findings at or above `critical`
|
||||
```
|
||||
|
||||
A skill run has no process exit code, so do not wire this into a CI gate as if
|
||||
it did — use a real scanner for blocking merges. What it *is* good for is an
|
||||
unattended scheduled audit: a Qoder Automation running this mode reports the
|
||||
verdict, and you act on it.
|
||||
@@ -0,0 +1,164 @@
|
||||
# Ship workflow — `autoresearch ship`
|
||||
|
||||
Run a pre-flight checklist before shipping — tests, types, lint, bundle size,
|
||||
security basics, and a final autoresearch pass on anything that fails.
|
||||
|
||||
The ship workflow is not just a checklist. It runs an autoresearch loop on
|
||||
each failing gate until it passes, then re-checks. You don't ship broken.
|
||||
You ship when everything is green.
|
||||
|
||||
---
|
||||
|
||||
## Invocation
|
||||
|
||||
```
|
||||
autoresearch ship
|
||||
```
|
||||
|
||||
Optional flags:
|
||||
```
|
||||
autoresearch ship --fast # skip slow checks (lighthouse, e2e)
|
||||
autoresearch ship --loop N # max N autoresearch iterations per gate (default: 20)
|
||||
autoresearch ship --dry-run # report status without fixing anything
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## The ship checklist
|
||||
|
||||
The workflow runs these gates in order. Each gate that fails triggers an
|
||||
autoresearch sub-loop to fix it before moving to the next gate.
|
||||
|
||||
### Gate 1 — Tests pass
|
||||
|
||||
```bash
|
||||
npm test # Node.js
|
||||
pytest # Python
|
||||
go test ./... # Go
|
||||
cargo test # Rust
|
||||
```
|
||||
|
||||
If tests fail → autoresearch loop on `src/**/*.ts` (or equivalent) with
|
||||
metric: failing test count (lower is better), max 20 iterations.
|
||||
|
||||
### Gate 2 — No type errors
|
||||
|
||||
```bash
|
||||
npx tsc --noEmit # TypeScript
|
||||
mypy src/ # Python
|
||||
```
|
||||
|
||||
If errors found → autoresearch loop on `src/**/*.ts` with
|
||||
metric: error count (lower is better), max 20 iterations.
|
||||
|
||||
### Gate 3 — No lint errors
|
||||
|
||||
```bash
|
||||
npx eslint src/ # JavaScript/TypeScript
|
||||
ruff check src/ # Python
|
||||
golangci-lint run # Go
|
||||
```
|
||||
|
||||
If errors found → autoresearch loop with metric: lint error count (lower is better).
|
||||
Auto-fixable errors are fixed first (`--fix` flag), then the loop handles the rest.
|
||||
|
||||
### Gate 4 — Bundle size (if applicable)
|
||||
|
||||
Only runs for frontend projects (detected: `next.config.*`, `vite.config.*`,
|
||||
`webpack.config.*`).
|
||||
|
||||
```bash
|
||||
npm run build 2>&1 | grep "First Load JS"
|
||||
```
|
||||
|
||||
Threshold: warn if > 300KB, block if > 500KB (configurable via `.autoresearch.yml`).
|
||||
|
||||
If over threshold → autoresearch loop on `src/**/*.tsx, src/**/*.ts` with
|
||||
metric: bundle size in KB (lower is better), max 20 iterations.
|
||||
|
||||
### Gate 5 — No hardcoded secrets
|
||||
|
||||
```bash
|
||||
git diff HEAD~1 --diff-filter=A | grep -iE "(api_key|secret|password|token)\s*=\s*['\"][^'\"]{8,}"
|
||||
```
|
||||
|
||||
If secrets found → do NOT autoresearch. Flag for human review. Block ship.
|
||||
|
||||
### Gate 6 — Dependency audit
|
||||
|
||||
```bash
|
||||
npm audit --audit-level=high # Node.js
|
||||
pip-audit # Python
|
||||
```
|
||||
|
||||
If critical vulnerabilities found → autoresearch loop to update affected
|
||||
dependencies, max 10 iterations.
|
||||
|
||||
---
|
||||
|
||||
## Ship report
|
||||
|
||||
After all gates pass, print:
|
||||
|
||||
```
|
||||
=== Ship report ===
|
||||
Tests: ✓ PASS (247 passing)
|
||||
Types: ✓ PASS (0 errors)
|
||||
Lint: ✓ PASS (0 errors)
|
||||
Bundle: ✓ PASS (187KB)
|
||||
Secrets: ✓ PASS (none detected)
|
||||
Deps: ✓ PASS (0 high/critical)
|
||||
|
||||
Autoresearch loops run: <N>
|
||||
Total improvements: <M> iterations kept
|
||||
|
||||
Ready to ship. Run: git push && <your deploy command>
|
||||
===
|
||||
```
|
||||
|
||||
If any gate is still failing after the max iterations:
|
||||
|
||||
```
|
||||
=== Ship report ===
|
||||
Tests: ✓ PASS
|
||||
Types: ✗ FAIL (3 errors remaining after 20 iterations)
|
||||
→ manual fix required: src/auth/session.ts:47
|
||||
|
||||
Ship BLOCKED. Fix the above before shipping.
|
||||
===
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Web research post-check
|
||||
|
||||
After all gates pass, use `WebSearch` to check:
|
||||
|
||||
```
|
||||
WebSearch: [your framework] [version] known issues [current year]
|
||||
WebSearch: [your main dependencies] security advisory [current year]
|
||||
```
|
||||
|
||||
If any critical advisories surface that the dependency audit missed,
|
||||
flag them before shipping. This is a final sanity check that goes beyond
|
||||
what local tools can detect.
|
||||
|
||||
---
|
||||
|
||||
## Configuration via `.autoresearch.yml`
|
||||
|
||||
Create this file in your project root to customise ship behaviour:
|
||||
|
||||
```yaml
|
||||
ship:
|
||||
bundle_warn_kb: 300
|
||||
bundle_block_kb: 500
|
||||
max_iterations_per_gate: 20
|
||||
skip_gates:
|
||||
- lighthouse # skip if no local server available
|
||||
extra_gates:
|
||||
- name: "E2E tests"
|
||||
command: "npx playwright test"
|
||||
metric: "failing tests (lower is better)"
|
||||
max_iterations: 10
|
||||
```
|
||||
@@ -0,0 +1,144 @@
|
||||
# Web research patterns
|
||||
|
||||
Qoder exposes a `WebSearch` tool (and `WebFetch` to read a promising result in
|
||||
full). Use them as a verification supplement — not a replacement for the Verify
|
||||
command, but an additional signal when local scripts alone cannot capture
|
||||
quality.
|
||||
|
||||
---
|
||||
|
||||
## When to use WebSearch in the loop
|
||||
|
||||
| Goal type | Use WebSearch for | Example query |
|
||||
|---|---|---|
|
||||
| SEO content | Check competing pages, keyword signals | `[target keyword] filetype:md OR site:*.dev` |
|
||||
| API correctness | Verify endpoint signatures, check for deprecations | `[library] [method] deprecated 2025 OR 2026` |
|
||||
| Dependency versions | Confirm latest stable before updating | `[package name] latest stable version` |
|
||||
| Best practices | Check if your approach matches current consensus | `[pattern] best practice [language] 2026` |
|
||||
| Content accuracy | Ground-truth check generated facts | `[claim] site:official-source.com` |
|
||||
| Bundle/perf baselines | Compare your score to current industry benchmarks | `[framework] bundle size benchmark 2026` |
|
||||
|
||||
---
|
||||
|
||||
## Pattern 1 — SEO content verification
|
||||
|
||||
Use when: optimising blog posts, landing pages, documentation for search.
|
||||
|
||||
After your local score script runs, supplement with:
|
||||
|
||||
```
|
||||
WebSearch: [target keyword] to see what the top 3 results have in common.
|
||||
Note: heading structure, content length, semantic coverage, internal links.
|
||||
If top results consistently have trait X that your content lacks,
|
||||
add "add trait X" as the next hypothesis.
|
||||
```
|
||||
|
||||
This gives you signal that no local readability or keyword-density script can
|
||||
provide — what the search engine is actually rewarding right now.
|
||||
|
||||
---
|
||||
|
||||
## Pattern 2 — API currency check
|
||||
|
||||
Use when: refactoring code that calls external libraries or APIs.
|
||||
|
||||
Before committing any API-surface change:
|
||||
|
||||
```
|
||||
WebSearch: [library name] [method name] changelog 2026
|
||||
WebSearch: [library name] [method name] deprecated
|
||||
```
|
||||
|
||||
If search returns deprecation notices or breaking changes, note the current
|
||||
replacement pattern and use that as the hypothesis instead.
|
||||
|
||||
This prevents iterating toward a working-but-deprecated solution that will
|
||||
break on the next library update.
|
||||
|
||||
---
|
||||
|
||||
## Pattern 3 — Dependency version check
|
||||
|
||||
Use when: the Verify command suggests a dependency might be outdated, or when
|
||||
optimising for security/bundle size.
|
||||
|
||||
```
|
||||
WebSearch: [package name] npm latest 2026
|
||||
WebSearch: [package name] security advisory
|
||||
```
|
||||
|
||||
Cross-reference against what is in `package.json`, `go.mod`, `requirements.txt`
|
||||
or equivalent. Use the delta as a hypothesis: "update [package] from X to Y,
|
||||
check if metric improves."
|
||||
|
||||
---
|
||||
|
||||
## Pattern 4 — Best practice calibration
|
||||
|
||||
Use when: stuck after 5 consecutive discards and local ideas are exhausted.
|
||||
|
||||
```
|
||||
WebSearch: [language/framework] [metric type] optimisation techniques 2026
|
||||
WebSearch: how to improve [metric] in [stack]
|
||||
```
|
||||
|
||||
Extract 3 concrete, actionable techniques from the top results — use `WebFetch`
|
||||
on the most promising one if the snippet is too thin. Do not extract vague
|
||||
advice. Add each as a separate iteration hypothesis. This restocks your
|
||||
hypothesis pool with externally validated approaches.
|
||||
|
||||
---
|
||||
|
||||
## Pattern 5 — Benchmark calibration
|
||||
|
||||
Use when: you want to know if your current metric value is good relative to
|
||||
the industry, not just relative to your own baseline.
|
||||
|
||||
```
|
||||
WebSearch: [framework] [metric] benchmark 2026 average
|
||||
```
|
||||
|
||||
If your metric is already at or above the industry median, note this and
|
||||
shift the goal definition (e.g. from "reduce bundle size" to "reduce bundle
|
||||
size while improving lighthouse score").
|
||||
|
||||
---
|
||||
|
||||
## Pattern 6 — Content accuracy check
|
||||
|
||||
Use when: the Verify command measures style/structure but not factual accuracy
|
||||
(e.g. documentation, blog posts, runbooks).
|
||||
|
||||
```
|
||||
WebSearch: [specific claim in content] site:[authoritative source]
|
||||
```
|
||||
|
||||
If the authoritative source contradicts your content, flag this as a
|
||||
required fix before the next iteration (accuracy issues override metric gains).
|
||||
|
||||
---
|
||||
|
||||
## Rules for using WebSearch
|
||||
|
||||
1. **Supplement, never replace.** The Verify command runs every iteration.
|
||||
Web research adds signal; it does not replace the metric.
|
||||
|
||||
2. **Search at the right time.** Patterns 1-3 supplement Phase 5 (Verify).
|
||||
Patterns 4-5 are for stuck recovery in Phase 1 (Review). Pattern 6
|
||||
runs in Phase 6 (Decide) when a kept iteration touches factual claims.
|
||||
|
||||
3. **Extract actionable hypotheses.** Never let a search result produce a
|
||||
vague conclusion ("content could be better"). Always turn the search
|
||||
result into a specific next hypothesis ("add a FAQ section with 3
|
||||
questions, which top-ranking competitors include").
|
||||
|
||||
4. **Log the research signal.** When a search result influences a hypothesis,
|
||||
note it in the results log description:
|
||||
`"added FAQ section (web research: top results for [kw] all include FAQ)"`
|
||||
|
||||
5. **Don't over-search.** Maximum one WebSearch call per iteration. If you are
|
||||
searching every iteration, your Verify command is probably too weak —
|
||||
strengthen the local script instead.
|
||||
|
||||
6. **Cite, don't guess.** `WebSearch` results come with source links; never
|
||||
turn an unverified snippet into a change that the Guard cannot catch.
|
||||
Reference in New Issue
Block a user