Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 10 additions & 3 deletions .tdd-swarm/final-submission-manifest.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,19 @@
# Final-submission execution manifest (session-lease scope repair round 3)

> **Historical planning artifact — not the final release manifest.** Current release facts and
> pending bindable fields live in `SUBMISSION.md` and
> `docs/submission-artifacts/RELEASE_BINDING.md`. Nothing here overrides retained manifests.

[locked-decision] Canonical requirements remain `Week_3_AgentForge.pdf`; this is not a replacement PRD/roadmap.

## Evidence boundary

- Baseline `23490ea`: 1001 Python passed/3 skipped, 75 console, 4 browser, dual CI green.
- Baseline `23490ea`: 1001 Python passed/3 skipped, 75 console, 4 browser, GitHub CI green and the
exact commit mirrored to GitLab. GitLab pipeline status is not a release gate.
- Planning catalog baseline `efd5ce3` adds only the tracked secret-free target catalog; this repair does not amend it. Later shared-worktree commits are not attributed to this planning repair.
- Run `aceddc495808427992efbd2b73b3598d`: 9 HTTP 200, 9 evidence, 9 `INDETERMINATE`, $0.09 outbound.
- Legacy summary `aceddc495808427992efbd2b73b3598d` claims 9 HTTP 200 / 9 evidence /
9 `INDETERMINATE`. Its `$0.09` figure has no retained billing or immutable request-cost manifest
and is quarantined: actual spend is unavailable, not `$0.09`.
- Judge baseline 60% agreement/33.3% false negatives/60% abstention is failed.
- [locked-decision] None proves calibrated safe/unsafe outcomes, findings, current-SHA live trace, production isolation, performance, or report completeness.

Expand Down Expand Up @@ -53,7 +60,7 @@
- T-F05g consumes and reauthenticates every separate chain artifact. No caller-combined document, local queue count, process boolean, overwrite, reload, rolling overlap, or in-place session/patient swap is authority.
- T-F05f permits only atomic queue claim before immediate J/K/M validation. All other mutation, resolution, adapter/client construction, network, and spend follow success. Every physical attempt obtains a new L→H/I refresh through M.
- T-F05e timestamp/reference languages remain Python 3.12 ASCII `re.fullmatch` plus portable Draft 2020-12 patterns, exact lengths, semantic validation, and explicit control rejection including terminal CR/LF.
- No PHI/secrets/sessions/raw hostile evidence in artifacts; staging/test is never called production. Swarm never merges main, publishes critical findings, remediates, load-tests, or posts socially autonomously.
- No PHI/secrets/sessions/raw hostile evidence in artifacts; staging/test is never called production. Swarm never merges main, publishes any finding/report, remediates, load-tests, or posts socially autonomously.

## Deterministic order

Expand Down
10 changes: 6 additions & 4 deletions .tdd-swarm/progress.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,17 +162,19 @@ NEXT: report to owner + STOP for explicit bounded live-campaign authorization.

## 2026-07-24 — Phase 0 / final-submission plan created
- [locked-decision] Production-grade posture retained; base `23490ea` recorded as 1001 Python passed/3
skipped, 75 console tests, 4 browser tests, dual CI green.
skipped, 75 console tests, 4 browser tests, GitHub CI green, and exact GitLab mirroring. GitLab
pipeline status is not a release gate.
- [locked-decision] Created the ten-ticket minimum manifest T-F01..T-F10 with exclusive same-wave scopes,
strict RED→test-review→freeze→GREEN→code/security review sequencing, four-slot accounting, traceability,
gate mapping, and role prompts. Planning only; no agent dispatch, application/test edit, live traffic,
credential action, commit, push, or main merge.
- [open-question] Execution remains gated by fresh staging SMART lease, exact authorization, distinct
Approver, paid provider/model activation, passing Judge calibration, separate load approval, and any
true-production isolation/credential choice.
- [locked-decision] Existing 9×HTTP-200/9×evidence/9×INDETERMINATE/$0.09 run is transport evidence only;
no finding/report is fabricated. PRD-32 remains conditional on three genuine independently reproduced
findings.
- [locked-decision] The legacy 9×HTTP-200/9×evidence/9×INDETERMINATE summary is transport evidence
only. Its `$0.09` prose is unsupported by a retained billing/request-cost manifest and is not
measured spend. No finding/report is fabricated. PRD-32 remains conditional on three genuine,
independently reproduced findings.

## 2026-07-24 — Phase 0 plan-review repairs complete
- [locked-decision] Adversarial review findings C-1..C-5 and I-1..I-4 repaired before dispatch. Superseded
Expand Down
5 changes: 3 additions & 2 deletions .tdd-swarm/reports/RTG-orchestrator.md
Original file line number Diff line number Diff line change
Expand Up @@ -155,8 +155,9 @@ The owner supplied a **live target authorization** in answer to §5. It is recor
the out-of-repo bundle; never committed here.
- Target is **already partially wired** in-repo (`README.md`, `docs/evidence/zap/*`,
`evals/results/live-campaign-20260724/*`, three `AF-VULN-2026-0724-*` reports,
`scripts/live_probe.py`, `console/vite.config.ts`); a prior live run (`aceddc4…`) already hit
it (9×200 / 9 INDETERMINATE / $0.09).
`scripts/live_probe.py`, `console/vite.config.ts`); a legacy `aceddc4…` summary claims
9×200 / 9 `INDETERMINATE`. The associated `$0.09` prose has no retained billing/request-cost
manifest and is quarantined rather than reported as spend.

**What this changes:** the *owner-authorized deployed live target URL + surfaces + provisioned
synthetic principals* requirement is now satisfied — a real advance for RT-07 (Week2 exposes the
Expand Down
6 changes: 3 additions & 3 deletions .tdd-swarm/reports/RTG-target-authorization.md
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,6 @@ WP-21B–E live executors consume. This file is authorization **metadata only**.
`README.md`, `docs/evidence/zap/{AUTHORIZATION.md,README.md,zap-target.json}`,
`evals/results/live-campaign-20260724/{summary.json,responses.jsonl}`,
`docs/vulnerabilities/AF-VULN-2026-0724-00{1,2,3}-*.md`, `scripts/live_probe.py`,
`console/vite.config.ts`. A prior live campaign run (`aceddc4…`, per the final-submission
manifest) already produced 9×HTTP-200 / 9 evidence / 9 INDETERMINATE / $0.09 outbound against
this exact target.
`console/vite.config.ts`. A legacy summary for `aceddc4…` claims 9×HTTP-200 / 9 evidence /
9 `INDETERMINATE` against this target. Its `$0.09` statement is unsupported by a retained billing
or immutable request-cost manifest and must not be used as measured cost.
6 changes: 4 additions & 2 deletions .tdd-swarm/reports/final-gap-planner.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,8 +45,10 @@ authorizations or immutable prerequisites miss noon. The plan does not claim ful

[locked-decision] Oracle precedence, fail-closed calibration, synthetic-only data, distinct approval,
critical-publication/remediation gates, package-authoritative contracts, staging-not-production, and
no-fabricated-findings remain binding. Run `aceddc495808427992efbd2b73b3598d` remains exactly 9 HTTP 200,
9 evidence, 9 `INDETERMINATE`, $0.09 outbound; the 60%/33.3%/60% calibration remains failed.
no-fabricated-findings remain binding. A legacy summary for
`aceddc495808427992efbd2b73b3598d` claims 9 HTTP 200, 9 evidence, and 9 `INDETERMINATE`;
its `$0.09` prose lacks a retained billing/request-cost manifest and is quarantined. The
60%/33.3%/60% calibration remains failed.

## Review-finding closure table

Expand Down
3 changes: 2 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,8 @@ deployed URL. No target code here.
## Non-negotiables
- Deployed target URL submitted every checkpoint; test a live system, not a mock.
- Multi-agent, not a pipeline. The Judge is independent of attack generation.
- Human approval gate before publishing critical findings or remediation.
- Human approval gate before publishing any finding/report, regardless of severity, or performing
remediation.
- Every eval = boundary | invariant | regression, mapped to OWASP Web + LLM Top 10.
- The Judge must never approve a confirmed exploit. Cost is never tokens × N.
- "Optional Engineering Deliverables" are mandatory (the PRD grades them).
Expand Down
Loading
Loading