Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
138 commits
Select commit Hold shift + click to select a range
1cc7b37
ci: jail merge routing to main behind the staging promotion path
modelmirror Jul 26, 2026
008406b
Merge pull request #853 from ModelMirrorAI/ci/main-base-jail
modelmirror Jul 26, 2026
d1e4c0a
docs(budget): correct the per-case cost model — evaluate fans out per…
modelmirror Jul 27, 2026
80407bc
docs: record four operational lessons that lived only in session memory
modelmirror Jul 27, 2026
5468d58
Merge pull request #865 from ModelMirrorAI/docs/budget-cost-model-cel…
modelmirror Jul 27, 2026
9773346
Merge pull request #866 from ModelMirrorAI/docs/operational-lessons
modelmirror Jul 27, 2026
e48970b
fix(evaluate): derive the backlog from resolved events, not a whole-c…
modelmirror Jul 27, 2026
9c457ef
feat(query): route --full's opinion body to the content store under t…
modelmirror Jul 27, 2026
c6332f3
Merge pull request #877 from ModelMirrorAI/fix/evaluate-backlog-memory2
modelmirror Jul 27, 2026
8a89d0d
Merge pull request #879 from ModelMirrorAI/feat/query-full-store-routing
modelmirror Jul 27, 2026
a23481a
feat(cost): add an ex-post spend backstop over the measured usage ledger
modelmirror Jul 27, 2026
673d902
chore(mcp): pin the CourtListener client to 1.1.0 and retire the stdi…
modelmirror Jul 27, 2026
dc21fac
Merge pull request #880 from ModelMirrorAI/feat/spend-backstop
modelmirror Jul 27, 2026
b6e3ac7
Merge pull request #882 from ModelMirrorAI/chore/mcp-pin-1-1-0
modelmirror Jul 27, 2026
f116021
ci: bound the assume-role step so a credential outage degrades, not h…
modelmirror Jul 27, 2026
6f83ab8
Merge pull request #884 from ModelMirrorAI/fix/bound-aws-credential-step
modelmirror Jul 27, 2026
939b571
docs(agents): let agents merge to staging, keep main maintainer-only
modelmirror Jul 27, 2026
899249e
ci: sync main into staging on a schedule, not at promotion time
modelmirror Jul 27, 2026
3eab37c
Merge pull request #885 from ModelMirrorAI/docs/agent-merge-authority
modelmirror Jul 27, 2026
4eb2068
Merge pull request #886 from ModelMirrorAI/ci/scheduled-staging-sync
modelmirror Jul 27, 2026
3028822
ops: report the promoted state, and stop reading promote's gates as b…
modelmirror Jul 27, 2026
3a8d75f
Merge pull request #887 from ModelMirrorAI/ops/promotion-regime-repor…
modelmirror Jul 27, 2026
f4f6f7e
test(salience): exercise the configured caps against a long-conferenc…
modelmirror Jul 27, 2026
50e307b
Merge pull request #888 from ModelMirrorAI/test/salience-cap-under-pr…
modelmirror Jul 27, 2026
b4e5a3d
docs(metrics): record why the generic back-test reports no baseline
modelmirror Jul 27, 2026
eefd711
docs(metrics): lead the no-baseline case with informativeness, not co…
modelmirror Jul 27, 2026
0a1912e
feat(backtest): report the always-deny floor and lift, per court and …
modelmirror Jul 27, 2026
e164d9a
Merge pull request #889 from ModelMirrorAI/docs/backtest-baseline-dec…
modelmirror Jul 27, 2026
2f3d655
feat(salience): state the segment base rate's lookback window, in cod…
modelmirror Jul 28, 2026
5d54233
docs(security): make the branch policy the staging environment's gate…
modelmirror Jul 28, 2026
91f84d4
Merge pull request #895 from ModelMirrorAI/feat/salience-lookback-window
modelmirror Jul 28, 2026
35f463f
Merge pull request #896 from ModelMirrorAI/docs/staging-env-branch-po…
modelmirror Jul 28, 2026
99c9917
Merge pull request #898 from ModelMirrorAI/main
modelmirror Jul 28, 2026
1a0bf44
docs(agents): scope the merge-commit rule to the two merges that need it
modelmirror Jul 28, 2026
2056952
Merge pull request #899 from ModelMirrorAI/docs/staging-merge-method-…
modelmirror Jul 28, 2026
73ae3a1
docs(salience): name the stratum the ops dashboard's baseline compari…
modelmirror Jul 28, 2026
af5c511
Merge pull request #900 from ModelMirrorAI/docs/ops-skill-stratum
modelmirror Jul 28, 2026
dad7da4
fix(devcontainer): restore the .claude chown a duplicate JSON key sil…
modelmirror Jul 28, 2026
faefa90
fix(backtest): let prior-vote read the population, not the twenty mos…
modelmirror Jul 28, 2026
30a01c6
docs(backtest): state what an uncapped prior-vote costs and what its …
modelmirror Jul 28, 2026
d770820
Merge pull request #901 from ModelMirrorAI/fix/devcontainer-postcreat…
modelmirror Jul 28, 2026
bb314a7
Merge pull request #906 from ModelMirrorAI/fix/prior-vote-sampling
modelmirror Jul 28, 2026
c0df037
docs(salience): record that the segment baseline stays SCOTUS-cert-only
modelmirror Jul 28, 2026
c9094b8
feat(analytics): report which configured MCP tools cells actually call
modelmirror Jul 28, 2026
afb44b0
Merge pull request #908 from ModelMirrorAI/docs/circuit-baseline-scope
modelmirror Jul 29, 2026
a8f7f64
Merge pull request #909 from ModelMirrorAI/feat/tool-usage-rollup
modelmirror Jul 29, 2026
b61a4bc
feat(analytics): flag cells that reached the open web instead of the MCP
modelmirror Jul 29, 2026
b3b727c
Merge pull request #912 from ModelMirrorAI/feat/web-search-signal
modelmirror Jul 29, 2026
97390bc
fix(devcontainer): let the container user own the Claude Code package
modelmirror Jul 29, 2026
569c070
Merge pull request #914 from ModelMirrorAI/fix/claude-npm-prefix-perm…
modelmirror Jul 29, 2026
9e33d79
feat(predict): give codex the live web search the other engines have
modelmirror Jul 29, 2026
eff1931
docs(process-version): name the retrieval surface as a digest input
modelmirror Jul 29, 2026
1c1a576
Merge pull request #915 from ModelMirrorAI/feat/codex-web-search
modelmirror Jul 29, 2026
b837607
docs(agents): add a stats-reviewer and state the interactive working …
modelmirror Jul 29, 2026
86e3bbf
Merge pull request #916 from ModelMirrorAI/docs/agent-workflow-and-st…
modelmirror Jul 29, 2026
8969d5b
docs: settle merge authority and correct verified doc staleness
modelmirror Jul 29, 2026
0e0363f
Merge pull request #917 from ModelMirrorAI/docs/clarify-merge-authority
modelmirror Jul 29, 2026
f0d12c7
fix(cli): render the stats --group-by help from its enum
modelmirror Jul 29, 2026
eefe26f
Merge pull request #920 from ModelMirrorAI/fix/stats-group-by-help
modelmirror Jul 29, 2026
8bcba9d
feat(gate): pre-flight a required status check against its producing job
modelmirror Jul 29, 2026
de91376
fix(cli): render option helps from the vocabularies they name
modelmirror Jul 29, 2026
b179481
Merge pull request #925 from ModelMirrorAI/feat/required-context-pref…
modelmirror Jul 29, 2026
aee238e
Merge pull request #926 from ModelMirrorAI/fix/cli-enum-help-drift
modelmirror Jul 29, 2026
dae44f1
refactor(cli): make the wide command signatures keyword-only
modelmirror Jul 29, 2026
a8fd3b8
Merge pull request #927 from ModelMirrorAI/fix/ruff-0.16-positional-args
modelmirror Jul 29, 2026
a43a5c1
feat(analytics): publish a court-facing docket pack
modelmirror Jul 29, 2026
333ee96
Merge pull request #930 from ModelMirrorAI/feat/docket-pack
modelmirror Jul 30, 2026
8da04fe
docs: pre-register the claim taxonomy and its scoring rule
modelmirror Jul 30, 2026
8dbf290
docs: correct the outcome-decomposition pre-registration
modelmirror Jul 30, 2026
c570232
docs: fix the scoring rule's shotgunning defence, which did not hold
modelmirror Jul 30, 2026
edffa50
predict: split the forecast of the court's reasoning from the predict…
modelmirror Jul 30, 2026
83e6fc8
Merge pull request #934 from ModelMirrorAI/docs/claim-taxonomy
modelmirror Jul 30, 2026
6bbdb5c
outcome: freeze the cert-stage docket signals at resolution
modelmirror Jul 30, 2026
84b7739
Merge pull request #935 from ModelMirrorAI/feat/outcome-cert-signals
modelmirror Jul 30, 2026
4a6c57a
evaluate: per-claim baselines, the score, and the floor
modelmirror Jul 30, 2026
706e47f
evaluate: keep the claim scoring rule, withdraw the claim set
modelmirror Jul 30, 2026
9b5eb82
Merge pull request #937 from ModelMirrorAI/docs/merits-stage-issue
modelmirror Jul 30, 2026
40364df
leaderboard: track evaluator agreement across the panel
modelmirror Jul 30, 2026
2f13852
fix(claims): retract a circular withdrawal and reweight its replacement
modelmirror Jul 30, 2026
e5be9d5
docs(model): pre-register one decision model for cert, interim, and m…
modelmirror Jul 30, 2026
e874f93
Merge pull request #938 from ModelMirrorAI/feat/evaluator-agreement
modelmirror Jul 31, 2026
e73e10a
Merge pull request #941 from ModelMirrorAI/fix/claim-withdrawal-reaso…
modelmirror Jul 31, 2026
315b20a
Merge pull request #944 from ModelMirrorAI/docs/decision-model
modelmirror Jul 31, 2026
25689b4
metrics: publish the rate a petition faced, beside the one it ended at
modelmirror Jul 31, 2026
787d283
Merge pull request #945 from ModelMirrorAI/fix/frozen-prediction-context
modelmirror Jul 31, 2026
8f773be
feat(predict): freeze the band a cell ran at, and score against the r…
modelmirror Jul 31, 2026
d2c76bd
feat(schema): give a vote its own vocabulary, and a merits judgment i…
modelmirror Jul 31, 2026
dc8b216
Merge pull request #946 from ModelMirrorAI/feat/frozen-prediction-con…
modelmirror Jul 31, 2026
5cc6c43
Merge pull request #948 from ModelMirrorAI/feat/vote-vocabulary
modelmirror Jul 31, 2026
8089f38
feat(backtest): truncate a replay snapshot by date instead of gutting it
modelmirror Jul 31, 2026
5556f1a
Merge pull request #947 from ModelMirrorAI/feat/replay-truncation
modelmirror Jul 31, 2026
bcfdcbe
metrics: give the predict prompt a cut of its own population
modelmirror Jul 31, 2026
33c674c
Merge pull request #949 from ModelMirrorAI/feat/paid-only-cert-cuts
modelmirror Aug 1, 2026
17b4e2c
metrics: warn that the granted/gvr split is ingestion history between…
modelmirror Jul 31, 2026
2d47d58
fix(corpus): strip a docket annotation so the identity join stops mis…
modelmirror Aug 1, 2026
076ea9a
metrics: refresh the scope manifest alongside the metrics artifacts
modelmirror Aug 1, 2026
b7782c2
Merge pull request #950 from ModelMirrorAI/fix/gvr-comparability
modelmirror Aug 1, 2026
da732b5
Merge pull request #954 from ModelMirrorAI/fix/docket-number-annotations
modelmirror Aug 1, 2026
0240b33
Merge pull request #955 from ModelMirrorAI/feat/scope-manifest-refresh
modelmirror Aug 1, 2026
c0c36f8
fix(corpus): catch the application spellings the non-cert scope rule …
modelmirror Aug 1, 2026
df11029
feat(live): address and identify interim-docket applications
modelmirror Aug 1, 2026
feeb4c9
feat(interim): read what an application asks, and how it ended
modelmirror Aug 1, 2026
b5081c4
Merge pull request #957 from ModelMirrorAI/fix/non-cert-scope-spellings
modelmirror Aug 1, 2026
d39af77
Merge pull request #958 from ModelMirrorAI/feat/interim-discovery
modelmirror Aug 1, 2026
95bbe1a
Merge pull request #959 from ModelMirrorAI/feat/interim-scope-events
modelmirror Aug 1, 2026
367d82c
feat(interim): the escalation ladder an interim forecast conditions on
modelmirror Aug 1, 2026
6e379b1
Merge pull request #960 from ModelMirrorAI/feat/interim-salience-signals
modelmirror Aug 1, 2026
137c02a
feat(live): poll the interim docket as a third numbering stream
modelmirror Aug 1, 2026
abd6fcb
Merge pull request #961 from ModelMirrorAI/feat/interim-polling
modelmirror Aug 1, 2026
f03f056
feat(corpus): keep the side counsel appears for, which ingestion disc…
modelmirror Aug 1, 2026
67524cc
Merge pull request #962 from ModelMirrorAI/refactor/arrival-time-counsel
modelmirror Aug 1, 2026
cfdd8fb
feat(historical): keep every decided petition, and make a full re-wal…
modelmirror Aug 1, 2026
ffd7b57
Merge pull request #964 from ModelMirrorAI/feat/historical-full-refresh
modelmirror Aug 1, 2026
7bafe97
feat(historical): let a refresh re-open one numbering stream
modelmirror Aug 1, 2026
8aa4b1b
ops(run-seed): dispatch input to re-open past Terms for a re-walk
modelmirror Aug 1, 2026
a605d2e
Merge pull request #965 from ModelMirrorAI/feat/refresh-historical-st…
modelmirror Aug 1, 2026
a6d16ab
Merge pull request #966 from ModelMirrorAI/ops/seed-refresh-dispatch
modelmirror Aug 1, 2026
397b859
docs(budget): measured ledger, an interim-docket reserve; fix the ops…
modelmirror Aug 1, 2026
aae3517
fix(statpack): grant-family series and the cross-Term GVR caveat (#978)
modelmirror Aug 1, 2026
2b6f2d6
ci(run-pull): a pipeline-runs dashboard and failure-only run logs
modelmirror Aug 1, 2026
192f86e
feat(live): re-poll unresolved applications and persist interim signa…
modelmirror Aug 1, 2026
2cfc7da
ci(run-pull): ambient token for logging writes; dashboard state harde…
modelmirror Aug 1, 2026
7e8d661
ci(run-backtest): salience-gate replay mode
modelmirror Aug 1, 2026
b283b0d
Merge origin/staging into ci/run-log-dashboard
modelmirror Aug 1, 2026
f24885f
ci(run-seed): dedupe live-minted duplicates in the maintenance tail
modelmirror Aug 1, 2026
a65a4fd
ci(run-backtest): route the salience-replay branch and tidy the PR arm
modelmirror Aug 1, 2026
84bc034
feat(salience): as-of projection and a gate replay over past Terms (#…
modelmirror Aug 1, 2026
cb512a2
fix(corpus): dedupe live-minted duplicate SCOTUS rows; true up the pe…
modelmirror Aug 1, 2026
5c4b644
Merge pull request #979 from ModelMirrorAI/ci/run-log-dashboard
modelmirror Aug 2, 2026
7a091af
Merge pull request #983 from ModelMirrorAI/ci/seed-dedupe-step
modelmirror Aug 2, 2026
a284a7d
Merge pull request #984 from ModelMirrorAI/ci/backtest-salience-gate-…
modelmirror Aug 2, 2026
1784e84
Merge remote-tracking branch 'origin/main' into staging
modelmirror Aug 2, 2026
df2fd0a
fix(corpus): read counsel through the optional-column contract (#985)
modelmirror Aug 2, 2026
3f036e1
ci(integration-test): an all scenario and branch-resolved deploy envi…
modelmirror Aug 2, 2026
db9aabd
ci(integration-test): close the freshness title-forgery channel
modelmirror Aug 2, 2026
4b93895
Merge pull request #987 from ModelMirrorAI/ci/integration-test-all
modelmirror Aug 2, 2026
17c33cd
chore(analytics): drop the unused weaker-or-equal band map
modelmirror Aug 2, 2026
802ea80
Merge pull request #988 from ModelMirrorAI/chore/drop-dead-band-map
modelmirror Aug 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
145 changes: 145 additions & 0 deletions .claude/agents/stats-reviewer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
---
name: stats-reviewer
description: Review statistical validity — reported numbers, metric definitions, leakage seams, pre-registration, and cross-engine comparability. Use whenever a change touches metrics/, scoring, the leaderboard, backtests, salience, analytics or ops reporting, process versioning, or the retrieval log, and before publishing any set of figures. Returns a verdict plus file:line findings; it reviews, it does not edit.
tools: Read, Grep, Glob, Bash
---

You review **statistical validity** for this repository: both the numbers it
reports and the code that produces them. You are a reviewer: you check the
claims against the data and the repo's own written standards, and report
findings with a clear verdict. You do **not** edit files — the calling agent
applies fixes.

This repo predicts a rare event against a deliberately non-representative
population: the whole-docket cert rate runs ~1–3%, but the gate selects into it,
so the baseline salience band runs 0.9%–2.6% while the high band runs
25.8%–48.0%. Nearly every way to be wrong here is a way to be *fooled by the
base rate* — including anchoring on the wrong one — and most of the rest is a
comparison between two things that were not measured the same way.

(`code-reviewer` owns whether a guard is implemented correctly; you own whether
the number that comes out of it is claimable.)

## First, gather context

You will be called in one of two modes. Establish which.

- **Reviewing a change**: `git diff --stat`, then the full diff (and
`git diff main...HEAD` on a branch). Read each touched file whole, plus the
module docstring — this codebase states its statistical reasoning in prose
beside the code, and a change that contradicts its own docstring is a finding
even when the tests pass.
- **Reviewing results**: get the numbers *and their provenance* — which
command produced them, over which population, at which process scope. A
figure whose denominator and stratum you cannot recover is itself the first
finding.

Then read the standards you are enforcing. They are the repo's, not yours:
`metrics/README.md` (what may be claimed and what may not),
`docs/salience.md` (selection bias, the lookback window, cross-court
incomparability), `docs/process-version.md` (the frozen partition), and
`AGENTS.md`'s leakage and artifact rules. Prefer citing one of these over
citing statistics in general — a finding lands when it names the standard the
repo already set for itself.

## Review checklist

- **Accuracy without a floor is arithmetic.** Under this class imbalance a
constant predictor scores its slice's base rate exactly. Any accuracy figure
must travel with the always-deny floor and the lift over it, and any claim of
*skill* must rest on Brier skill against a base rate, not on accuracy.
- **Every number carries its denominator.** Means here are taken over present
values only, so each metric on a row can have a different silent `n`. Check
that a reported figure states the sample it rests on, that a small `n` is
visible rather than rounded into a headline, and that an unknown denominator
prints as unknown rather than as zero.
- **Strata never blend.** Forward, retrospective, and procedural are separate
populations; no headline metric may mix them. Backtests — the retrospective
stratum, replay runs, `backtest.json`, `cert-backtest.json` — are iteration
instruments and are **never claimable performance**. Flag any prose that
presents a replay figure as a result.
- **Pooling that hides the failure.** A pooled cross-court or cross-band figure
is dominated by whichever slice supplies the most resolved events and can
average away a severe failure on the population actually being predicted, as
well as mixing outcome vocabularies. Check that the per-slice cut is
reported, and that anything incomparable across slices is not published as if
it were comparable.
- **Ranking is not a measurement.** The leaderboard's order rests on
N-unweighted point estimates, so a single lucky cell can outrank a large
honest sample. Flag a rank presented as a finding without the sample sizes
beside it, and flag any new rank key that inherits the same blindness.
- **The baseline is a stated choice.** The base-rate lookback window is config,
not a constant, and moving it re-bases every forward skill number at once —
per-Term high-band grant rates span roughly 26%–48%, so the window choice is
worth ~10 points of the number a Brier skill is scored against. A change to
it is a reviewable diff that must say why; a comparison across a window
change is not a comparison.
- **Selection effects.** The predicted population is a deliberately biased
subsample — salience gating selects high-relist and CVSG petitions, which
grant far above the whole-docket rate — so the docket-wide base rate is the
wrong anchor for it. Check that the anchor matches the population, and that
sampling or truncation (`limit`, head-first caps, recency ordering) samples
the population rather than its most recent tail.
- **Pre-registration holds on the digest, never the label.** The process digest
is the partition key; a label is sugar. The digest *moving* on a real process
change is the design — so the violation to hunt is a change that **suppresses**
the move: a capability added without a matching `ENGINE_RETRIEVAL` entry, a
field quietly excluded from the canonical config, or code that filters on
`label`. Equally, flag deliberately version-blind surfaces (the prediction
census, the leakage digest) being scoped to frozen-only — they exist to
surface shakedown contamination.
- **Leakage seams fail silently.** The high-risk ones: the snapshot redaction
blocklist is **key-name based**, so a new ingestion channel or
an upstream field rename un-redacts without any error — it is one flat list
that must enumerate every channel's keys, not a list per channel; the
base-rate guard
excludes the case's own Term by a single comparison operator; and the
retrieval parsers are deliberately tolerant, so a call type they stop
recognizing yields a clean-looking empty log and an unearned clean leakage
grade. Any new retrieval channel must be captured before it is enabled, or
the change moves reach from an audited channel to an unaudited one.
- **Cross-engine comparability.** A comparison between engines is valid only
while the prompt bytes, the kickoff, the tool and MCP surface, the scored
population, the stratum, and the process scope are all held constant — and
where a retrieval surface differs, the digest must differ so the cells
partition instead of pooling. Flag: an engine-conditional flag, sandbox, or
tool grant not mirrored into the declared retrieval surface; the sites that
build one engine's args drifting out of lockstep; differential parser
coverage; and differential cell-failure rates, which remove one engine's hard
cases from the scored set.
- **The garden of forking paths.** Predictors × evaluators × strata × salience
bands × calibration bins are all reported side by side. The repo runs no
significance testing and corrects for nothing, which is fine while the grid is
*described*; it stops being fine the moment one cell of it is lifted out as a
headline. Flag a claim selected from the grid after the fact, and say what
the pre-registered version of that claim would have been.
- **Caveats travel with the number.** Counts are denial-reweighted estimates,
filing censuses are upper bounds, tool-call zeros are not evidence of choice,
and an empty frozen headline is a shakedown state rather than a regression.
Where the code renders prose around a figure, the caveat belongs in the same
sentence the figure is in — a caveat one section away does not travel when
the line is quoted.
- **Tests pin the claim, not the arithmetic.** A scoring change needs a test
that fails on the wrong answer, over the repo's fixtures; a leakage guard
needs a test that fails when the guard is removed. Read the test against the
guard and satisfy yourself it would catch the removal — do not try it by
editing the tree. Where you cannot tell from reading, say so and recommend
the author delete the line, watch the test fail, and restore it.

Do **not** demand confidence intervals, p-values, power analysis, or
multiple-comparisons corrections as such: the repo has deliberately built none
of that machinery, and its stated instruments are `n=` counts, the always-deny
floor, Brier skill against a segment base rate, decile calibration bins, and
Kendall tau-b for rank agreement. Press for those. Where a question genuinely
cannot be answered without inferential machinery the repo lacks, say so plainly
and recommend the claim be weakened rather than the machinery be improvised.

## Report

A verdict first — **blockers** (a number that is wrong, unclaimable, or
incomparable as presented; a validity guard weakened), **recommended**, **nits**
— then findings as `severity · file:line — the claim → what the data or the
repo's standard actually supports`. For each finding on a reported figure,
give the corrected reading, not just the objection. Close with what you checked
and found sound, so the caller knows what not to re-derive. "No concerns" is a
complete answer when true.
5 changes: 2 additions & 3 deletions .devcontainer/devcontainer.json
Original file line number Diff line number Diff line change
Expand Up @@ -9,14 +9,13 @@
"ghcr.io/devcontainers/features/aws-cli:1": {}
},
"onCreateCommand": "uv sync",
"postCreateCommand": "sudo chown -R vscode:vscode /home/vscode/.claude && bash .devcontainer/setup-corpus-access.sh",
"postCreateCommand": "bash .devcontainer/setup-corpus-access.sh",
"postCreateCommand": "sudo chown -R vscode:vscode /home/vscode/.claude && bash .devcontainer/setup-claude-package-ownership.sh && bash .devcontainer/setup-corpus-access.sh",
"postStartCommand": "bash .devcontainer/check-corpus-session.sh",
"waitFor": "postCreateCommand",
"userEnvProbe": "loginInteractiveShell",
"secrets": {
"CLAUDE_CODE_OAUTH_TOKEN": {
"description": "Claude Code OAuth token from `claude setup-token`; skips interactive login (user-scoped, optional)"
"description": "Claude Code OAuth token from `claude setup-token`; skips interactive login (user-scoped, optional). Set this and do NOT run an interactive `claude` login in the container: a stored credential from one shadows this token once it expires, and `claude` then fails 401 at startup while `claude -p` still works. Recovery: move ~/.claude/.credentials.json aside."
},
"FEDCOURTS_COURTLISTENER_API_TOKEN": {
"description": "CourtListener REST API token, required by `fedcourts pull` (user-scoped, optional)"
Expand Down
29 changes: 29 additions & 0 deletions .devcontainer/setup-claude-package-ownership.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,29 @@
#!/usr/bin/env bash
# Give the container user ownership of the globally installed Claude Code
# package, so the CLI's in-place auto-update succeeds instead of failing with
# EACCES.
#
# The claude-code feature installs into the shared nvm prefix as root, which
# leaves the `@anthropic-ai` scope directory owned by root without group write.
# Only the owner is reset — the group comes from the prefix's setgid bit and
# must stay as the node feature set it.
#
# Deliberately advisory, like the corpus session check: no `set -e`, always
# exits 0. A container whose Claude Code lives outside this prefix, or which
# has none at all, is a valid state, and a failed create would be a far worse
# outcome than a package that cannot self-update.
set -uo pipefail

# A project-level `.npmrc` can steer this value: npm refuses to install against
# such a prefix but still prints it. Require an absolute path rather than
# trusting whatever comes back.
prefix="$(npm config get prefix)"
[[ -n "${prefix}" && "${prefix}" == /* ]] || exit 0

scope="${prefix}/lib/node_modules/@anthropic-ai"
[[ -d "${scope}" ]] || exit 0

if ! sudo chown -R "$(id -un)" "${scope}"; then
echo "Could not take ownership of ${scope} — Claude Code auto-update will fail."
fi
exit 0
7 changes: 7 additions & 0 deletions .github/actions/corpus-ranged/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -32,6 +32,13 @@ runs:
with:
role-to-assume: ${{ inputs.role-to-assume }}
aws-region: ${{ inputs.aws-region }}
# Bound the assume-role wall clock: unbounded, its 12 retries over a
# multi-minute connect timeout outlast every job budget here (see the
# workflow-traps list in docs/pipeline.md). Note the cost this buys —
# the action does not clear its timer on the error path, so ANY failing
# assume-role now burns the full 120s and ends on "Action timed out"
# rather than its own message; the real cause is earlier in the log.
action-timeout-s: 120
- name: Select the ranged corpus backend
shell: bash
env:
Expand Down
7 changes: 7 additions & 0 deletions .github/actions/corpus-readonly/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,13 @@ runs:
with:
role-to-assume: ${{ inputs.role-to-assume }}
aws-region: ${{ inputs.aws-region }}
# Bound the assume-role wall clock: unbounded, its 12 retries over a
# multi-minute connect timeout outlast every job budget here (see the
# workflow-traps list in docs/pipeline.md). Note the cost this buys —
# the action does not clear its timer on the error path, so ANY failing
# assume-role now burns the full 120s and ends on "Action timed out"
# rather than its own message; the real cause is earlier in the log.
action-timeout-s: 120
- name: Pull the corpus
shell: bash
env:
Expand Down
7 changes: 7 additions & 0 deletions .github/actions/corpus-sidecar/action.yml
Original file line number Diff line number Diff line change
Expand Up @@ -56,6 +56,13 @@ runs:
aws-region: ${{ inputs.aws-region }}
output-credentials: true
output-env-credentials: false
# Bound the assume-role wall clock: unbounded, its 12 retries over a
# multi-minute connect timeout outlast every job budget here (see the
# workflow-traps list in docs/pipeline.md). Note the cost this buys —
# the action does not clear its timer on the error path, so ANY failing
# assume-role now burns the full 120s and ends on "Action timed out"
# rather than its own message; the real cause is earlier in the log.
action-timeout-s: 120

- name: Launch the corpus query sidecar
shell: bash
Expand Down
Loading