Skip to content

Latest commit

 

History

History
260 lines (219 loc) · 13.1 KB

File metadata and controls

260 lines (219 loc) · 13.1 KB

Changelog

1.1.0 — 2026-09-30

Added

  • tianoshield-propagate case runs one case request and keeps every stage it reaches (preparation, analysis, generation, validation) with its outcome and diagnostics. Abstention, withheld evidence, and invalid input are recorded as outcomes, not errors.
  • tianoshield-propagate packet now builds a review packet from saved case evidence. The fixed research operator packet remains a separate replay.
  • tianoshield-propagate source-package prepares a base archive and ordered source patches, with upstream pre/post inputs bound to upstream Git objects.
  • tianoshield-propagate validate-build runs a pinned build and regression-test recipe in a bounded local process group. It is not a sandbox, and a passing build is evidence, not a safety verdict.
  • A cross-project dataset inventory and the September 2026 evaluation records in experiments/2026-09-propagation-evaluation/ (OpenSSL development and libxml2 held-out fix families). Their labels still need independent human review.
  • License files (MIT for original software, with third-party notices and the EDK II license inventory), CITATION.cff, SECURITY.md, CONTRIBUTING.md, and the guided example now live in the working repository, so every public release is exported from one source.
  • The guided example (scripts/try_example.py) builds a synthetic upstream Git repository, because preparation now binds upstream inputs to Git objects. tests/test_try_example.py keeps it working.

Changed

  • Classifier evidence records now record the analyzer release (classifier_evidence.ANALYZER_RELEASE, 1.0.0) as package.version instead of the installed package version. A package release therefore no longer changes the research records or the frozen benchmark anchors. The value changes only when analyzer output changes.

Removed

  • The automation/ multi-model feature orchestration harness, its feature contracts, the completed Phase A-H program queue (docs/PATCH_PROPAGATION_PROGRAM.md), and their tests. They supported only the finished Phase B-H implementation program and remain in Git history.
  • The Phase H fake-only external-action simulation (external_action.py), the external-action-authority-v0.1 schema, and their tests. TianoShield stops at human review, and real downstream actions are not planned. The external_action_authority fields in other records remain and are always false. The module was never exposed through the command line.

1.0.0 — 2026-09-14

This is an intentionally breaking architecture reset. It removes obsolete runtime contracts instead of preserving compatibility with them.

Changed

  • Manifests and reports use one strict version 1.0 contract. Unknown fields, duplicate target IDs, implicit directory selection, unsupported versions, and unsealed inputs are rejected.
  • Repository preparation publishes an atomic, hash-bound input set and uses a hardened Git object reader. The analysis engine verifies every input receipt before it reads source bytes.
  • Tree-sitter is the only authoritative syntax matcher. Structural and semantic analyzers consume one shared parse, patch-delta, and alignment context; semantic results remain evidence-only.
  • Maintained documentation is consolidated into the README, architecture, development, and research-status documents. The README now contains a reproducible prepare/analyze/generate/validate demo; generated experiment evidence remains unchanged.

Removed

  • The regular-expression text analyzer, semantic classify/projection mode, legacy CLI aliases, bare-manifest invocation, old manifest/report schemas, and deprecated benchmark metrics.
  • The Java SPIDER/SafePatch and Joern trees, QAS container/bootstrap path, Python 2 statistics scripts, generated Java reference output, and their separate dependency listing. The final snapshot remains available in Git history as documented in docs/ARCHITECTURE.md.

0.3.2 — 2026-07-15

Independent-review hardening release: packaging, provenance, metrics, and reproducibility corrections for the downstream candidate benchmark. No classifier behavior changed; all 37 per-target verdicts and all 20 scored outcomes are identical to 0.3.1.

Added

  • scripts/verify-prepared-target-blobs.py: every prepared target is re-fetched and byte-verified against its pinned remote Git blob (repository URL + 40-character ref + source path), producing the deterministic blob-verification.json ledger. All 20 targets were live-verified byte-for-byte with zero mismatches; target-lineage.json now carries a non-null, traceable git_blob_sha for every record, and the verification is a required release/review gate.
  • scripts/build-review-ledger.py and review-ledger.json: a machine-readable human-review ledger transcribed from the committed review documents, recording reviewer identity, review role, blinding status, visible information, rubric version, labels, evidence, confidence, and adjudication. Coverage tests fail if claims exceed the recorded reviews.
  • --deterministic harness mode: results and artifacts are byte-identical across reruns at one commit; timestamps and timing measurements move to the separate, documented-volatile timing.json (timing functionality is preserved). scripts/compare-normalized-results.py is the committed normalization comparator for non-deterministic runs.
  • Dependence/cluster-structure metrics: unique normalized function-content counts, duplicated-content clusters, repository+CVE clusters, and temporal lineage clusters, alongside the preserved raw target-row metrics.
  • Packaging tests that build the wheel and sdist with uv, verify their members, scan every archive member for private paths and secret patterns, and install/run the CLI from the built wheel.
  • Atomic evidence collection: full manifest validation (including candidate-id uniqueness) before any API call, staged output, atomic publish, and preservation of the previous valid evidence on failure.

Changed

  • The source distribution now ships only the installable package and metadata; it previously packaged the entire research repository, including internal notes with private absolute paths.
  • "Independent-backport recognition: 3/3" was corrected to its true units: one independent-backport temporal sequence correctly ordered (1/1 sequence; 3/3 correlated snapshot classifications; 1/1 fix-event recognition). The old independent_backport_recognition key is retained but deprecated, with an explicit unit and replacement pointer.
  • Exact/AST recognition rates are now framed as mechanical characterization (the lineage audit shares parsing/normalization code with the classifier), not independent validation of generalized accuracy.
  • Independent-review claims were narrowed to the recorded population: primary manual review covered all 20 target/CVE pairs (not blinded); the independent second review covered 11 selected nontrivial divergent and temporal-boundary cases. candidates.yaml reviewer lists were reconciled accordingly. No reviewer judgments were synthesized.
  • Reproducibility claims now distinguish deterministic classifier replay of committed prepared inputs, deterministic target-lineage regeneration, and acquisition/source verification; "byte-identical rerun" claims are limited to --deterministic runs at one commit.

0.3.1 — 2026-07-15

Downstream candidate benchmark evidence update.

Added

  • A candidate provenance schema, frozen repository/ref evidence, mandatory target-lineage review, independent review, and an advisor packet for the completed downstream benchmark.
  • A scored set of 20 authentic prepared targets spanning 8 downstream EDK II repositories, divided into 6 core candidates and 5 appendix candidates; one loongson/edk2 pure-upstream-ancestor control is rejected.
  • Stratified results: 10/10 exact/AST-pre, 4/4 exact/AST-post, 0/2 divergent-patched recognition with conservative abstention, 6/20 overall abstention, and zero confidently-wrong classifications. (This entry originally reported "3/3 independent backport"; 0.3.2 corrects the unit to one independent-backport temporal sequence observed as three correlated snapshots.)

Changed

  • Corrected fork-network candidate methodology: upstream commit-object resolution is not merge evidence, so repository eligibility uses compare status and still requires a separate target-level lineage audit.
  • Kept structurally divergent functions advisory-only and deferred them to human review instead of treating abstention as patch-state recognition.

0.3.0 — 2026-07-09

Repository-aware preparation and defensibility release.

Added

  • tianoshield-prepare-local, a read-only preparer for extracting an authentic target from an existing local Git repository pinned to a full commit SHA.
  • Bounded exact-path, filename, function, and history discovery with explicit ambiguity and negative-search outcomes.
  • preparation-lock.json provenance containing search attempts, resolved ref, source path and SHA-256, input kind, and extraction method.
  • Machine-readable analyzer reason codes and conservative handling of contradictory high-confidence results.
  • A reproducible eight-manifest experiment harness and committed public-safe rerun artifacts.

Changed

  • NOT_APPLICABLE is explicitly scoped to the provided prepared target file; it is not a repository-wide absence claim.
  • UNCERTAIN explicitly requires manual review and is never evidence of safety.
  • Confidence is described as matcher confidence, not exploitability, severity, patch-correctness, repository-wide applicability, or release confidence.

Safety

  • Local preparation never clones, fetches, checks out, resets, or modifies the source repository.
  • An explicit --source-path absent at the pinned commit is a hard SOURCE_PATH_NOT_FOUND failure; discovery never silently degrades from an explicit path to filename/function search.
  • Function search passes the pattern literally (git grep -F -e) so --prefixed strings cannot be parsed as options, and git grep errors are surfaced as SEARCH_ERROR instead of being recorded as absence.
  • Tree listing and search use NUL-terminated output, so non-ASCII/quoted paths are matched and extracted correctly.
  • The rerun harness exits non-zero on leak hits, public-safe nulling problems, expectation mismatches, or CLI exit-code mismatches, and verifies inputs.upstream_pre/inputs.upstream_post are nulled in public-safe JSON.
  • Output paths inside the source repository, non-empty collisions, unsafe target IDs, missing upstream files, non-C targets, and symlink targets are rejected before extraction.
  • External posting, vendor notification, and patch submission remain outside the tool's authority.

0.2.0 — 2026-06-09

First hardened release of the patch-propagation engine.

Added

  • Packaged tianoshield-propagate CLI (pyproject.toml, src/ layout, uv workflow); root propagate.py retained as a back-compat shim, and both run manifest.yaml and bare manifest.yaml invocation shapes work.
  • Stable v0.2 JSON report schema (schema/report-v0.2.schema.json, schema/report-list-v0.2.schema.json) with UTC Z-suffixed timestamps, all analyzer observations recorded per target, recommended human actions, and target/report-level limitations. The current report contract is documented in docs/ARCHITECTURE.md.
  • Typed ManifestReader with early validation and helpful errors; manifest contract now superseded by the strict v1 contract documented in docs/ARCHITECTURE.md.
  • VerdictClassifier with conservative absence-agreement rule: NOT_APPLICABLE requires all detectors to agree the function is absent.
  • Advisory artifact renderer (--write-advisory) following the TianoShield advisory comment contract (summary, evidence, confidence, recommended human action, limitations; advisory_only / human_review_required labels).
  • --public-safe mode across JSON/Markdown/advisory outputs: local path fields are nulled (not omitted, so the schema still validates) and fixture-only fields are suppressed; targets are labeled repo@ref.
  • pytest suite (85 tests): Stage 1/1.5 characterization, Stage 2 kraxel/edk2 clean replay, mu_tiano_plus UNCERTAIN boundary (overclassification guard), external-fork absence, CLI exit-code contract (0/1/2), schema validation, renderer and path-leakage tests.
  • Stage 2 fixtures committed with provenance/licensing documentation (mock-supply-chain/stage2-real-candidates/README.md).
  • Deferred Z3/ConditionDiff audit with an explicit revisit trigger; the current evidence-only semantic boundary is documented in docs/ARCHITECTURE.md and docs/RESEARCH_STATUS.md.

Changed

  • Exit-code contract formalized: 0 success (UNCERTAIN is a valid conservative verdict), 1 declared-expectation mismatch, 2 manifest/IO error.
  • Generated artifacts now go to git-ignored out/ (legacy reports/ contains historical Java-SPIDER text reports only).

Explicit non-goals (unchanged)

  • No GitHub posting, patch submission, vendor notification, or private/embargoed workflows. The tool analyzes prepared local files and writes local artifacts only.