Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Tutorial: run a whole-tracker audit

This walks /audit:issues end to end against QuantEcon/action-translation — 228 items, the repo the runbook was first executed against by hand.

It differs from the evaluation tutorial in one important way. That one reproduces a committed reference, so every number you produce can be checked. Here there is no reference: /audit:issues has been run as a skill exactly once — run 1, against this same repo on 2026-07-28, which found seven plugin defects and is recorded here. A method generalised from one execution is still a hypothesis, so your run is the next data point in the validation program (skills#16), and the part no automation can supply is your judgement of the output. Step 6 is therefore not optional garnish — it is the result.

Canonical references (this tutorial points, never restates): the procedure in SKILL.md, the method in doctrine.md, the org conventions in quantecon-context.md, the output contract in deliverables.md.

What you need

Step 0 — install the plugin

claude plugin marketplace add QuantEcon/skills
claude plugin install audit@quantecon

Then restart your session — plugins register at startup, so the skill does not appear until you reopen.

The /plugin marketplace add … slash form does the same job, but it is a terminal-CLI built-in: the VS Code extension and the web app answer /plugin isn't available in this environment, while the claude plugin CLI above works from any shell. Confirm with claude plugin list — the version it reports should match the audit entry in marketplace.json. (Naming a number here would go stale on the next release; if the two disagree, the install did not pick up the latest — claude plugin update audit@quantecon.)

If /audit:issues is still unrecognised after restarting, the plugin-prefixed slash form needs Claude Code 2.1.216+; the bare /issues works on older builds, and natural-language invocation (“audit every issue in this repo, output to …”) works on any version (using-skills § troubleshooting).

Step 1 — put the working directory where the repo already ignores it

cd ~/work/quantecon/action-translation
git checkout main && git pull --ff-only   # phase 2 verifies against the DEFAULT branch
git status --short                        # must be empty — the read-only baseline
git check-ignore -v .dev/scratch/x        # → .gitignore:… .dev/scratch/*

Be on the default branch, not merely clean. Phase 2’s core question is whether an issue still reproduces on main; run from a feature branch and every answer is measured against your unmerged work instead. A clean tree on the wrong branch passes the git status check and silently invalidates the phase the whole run exists to test — so check the branch, not just the status.

action-translation has a .dev/ notes system whose .dev/scratch/* is already gitignored, which makes it the first-choice working directory: the run leaves git status completely clean, and no .gitignore edit is needed — that would itself be a change to a tracked file. Repos without one fall back to an untracked .audit/ at the root (SKILL.md § Working directory).

Step 2 — invoke

/audit:issues QuantEcon/action-translation --out .dev/scratch/audit-2026-07-28

Before phase 1 the skill discovers its inputs rather than asking for them: the notes system (here .dev/STATE.md, PLAN.md, FUTURE.md, decisions/), the label policy, the work-plan anchor to tier against, and any prior audits. It should ask you only where discovery is genuinely ambiguous — two plausible plan anchors, say — and never merely because something came up empty. Every resolved input must appear in the report’s method section; that is the first thing to check in Step 6.

Step 3 — phase 1, the snapshot (~11 seconds)

Deterministic, and the only phase driven by a script rather than judgement. Measured 2026-07-27:

fetching issues from QuantEcon/action-translation …
  116 issues
fetching pull requests from QuantEcon/action-translation …
  112 pull requests

  numbers 1..228: 228 accounted, 0 unaccounted
  open issues     55    49 comments across 30
  closed issues   61    80 comments across 32
  open PRs         2    2 comments across 1, 5 reviews across 2
  closed PRs     110    28 comments across 18, 377 reviews across 105

Three things to read off it. 0 unaccounted means every number in 1..228 was captured — a non-zero count is not automatically wrong (deleted or transferred items, PR numbers burned by branches that never opened) but each one now owes an explanation in Step 5. No truncation warning — a stream returning exactly at --limit is indistinguishable from a truncated one. And meta.json’s fetched_by should be you, since visibility is per-account.

Your counts will differ from the ones above — the tracker moves (#11 measured 111/110 in July). That is expected and is why the report quotes its own snapshot timestamp rather than an earlier count.

Step 4 — phase 2, verify — and interrupt it

The long phase: 116 items checked against the default branch rather than against what their threads claim. Findings are appended to findings.md one entry per item, as each is verified — both passes, the open issues under ## Open and the closed ones under ## Closed.

This is the test. Once 20–30 entries exist, interrupt the session — close it, or press Esc twice. Then open a new session in the same directory and re-invoke the same command. What should happen: it partitions issues.json by state and, for each side independently, resumes at the lowest number with no entry under the matching heading, re-verifying only the last entry in each (which may have been half-written). What would be a failure: restarting from item 1, skipping the item it died on, duplicating entries, or resuming the open set correctly while re-doing the closed set from scratch.

Interrupt during the closed pass too, if you get the chance — that is the half that was not checkpointed at all before #34, so nothing has ever resumed from it. Resumability is asserted in three separate files and has never been tested. Two fixes have gone in ahead of this run and neither has been exercised: #17 named the artifacts, since before it phases 2 and 3 named none and a resumed session could only work by inventing the same filename; and #34 made the checkpoint cover both passes, since run 1 wrote only the open set and sent 62 closed issues straight to the catalog. This run is what checks both.

While it runs, tail findings.md occasionally. Every status claim should carry [verified], [stated] or [inferred], and a [verified] should cite file:line, a merged PR, a tag, or a commit — and whatever it cites must resolve on the ref the audit named, never a comment. A citation that only resolves in the author’s working tree or on an unmerged branch is the defect doctrine §2 now rules out; run 1’s headline finding had exactly that shape.

Step 5 — phases 3 to 5

Phase 3 writes the cross-link graph to links.md. Phase 4 tiers into the plan it discovered in Step 2 and writes the bundle. Phase 5 reconciles against coverage.json and folds any correction back into the documents rather than appending an erratum.

With 55 open issues this run should produce the full four-document bundle; a repo under about 30 open issues should instead fold the catalog and links into the report. That threshold is new and untested, so note whether four documents felt right at 55 or merely dutiful. The bundle’s destination is .dev/audits/<date>-issues/ — the audit writes it there, but committing it is your call, not the run’s.

Step 6 — review the output

The run cannot check any of this about itself. Ten checks, the last two of which only you can make:

#CheckWhere to lookWhat failure looks like
1Method section is completereport §1The plan anchor, label policy and prior-audit search are not all named
2Every claim is taggedcatalog entriesAn untagged status claim — a defect by the doctrine’s own rule
3[verified] means verifiedsample 5 “fixed” calls, and resolve each on the ref named in the headerCitation is a thread comment rather than code — or a citation of any form that looks right and does not resolve on that ref. git merge-base --is-ancestor <sha> <ref> on any commit cited. This is the check run 1 passed and should not have
4Code beat the threadsample 5 moreThe report repeats “fixed in #204” without saying it checked the branch
5The closed side was readclosed-set verification, and findings.mdClosed issues summarised from title and stateReason only; no deferred remainders surfaced. Also check the checkpoint, not just the output: closed entries reaching the catalog with no matching ## Closed block is how run 1 passed this check in its report while failing it in its log
6Coverage is honestREADME.md indexUnaccounted numbers waved at collectively; residue (inline review comments, GraphQL-only data) not stated
7Siblings were checkedexternal cross-link registryNo sibling considered, in an org where the same fix lands in several repos as SYNC: PRs
8Drafted comments are safeGC tierA closing keyword immediately before an owner/repo#N reference — that closes the upstream item when it lands
9Is the tiering right?report tiering sectionT0 that isn’t this week’s work, or a tier list not actually tied to .dev/PLAN.md. You are the authority; the run is guessing
10Would you act on it?the whole bundleThe only check that matters, and the only one no self-audit can make

Then confirm the boundary held: git status --short shows nothing but your ignored working directory, and the tracker is unchanged — no comments, no labels, no closures.

Step 7 — record the run

Findings belong in this repo; the bundle does not. Write reviews/audit-run-<repo>-<date>.md — run 1’s is audit-run-action-translation-2026-07-28.md — following the shape of the ge_arrow validation run:

Post the summary to skills#16. Anything that broke becomes an issue against the plugin, not a note in the margin.