Files
set50-system/docs/test-evidence/2026-08-23-replay-report-integrity-remediation.md
Kunthawat Greethong ead9aeb25c chore: pre-existing in-tree work (event-study/research/vintages/prices + migration script + integrity docs)
Committing the prior uncommitted working-tree state that predates this session's
data-source work (was already modified/untracked at session start) so the tree
is clean before push. Includes: event-study + research report integrity/forward
observation work, prices tests, research hash migration script, and the
2026-08-23/24 engineering-log + test-evidence notes. Verified green as part of
the full 362-test suite.
2026-08-29 09:19:24 +07:00

3.5 KiB

Test evidence — 2026-08-23 replay and research-report integrity remediation

Automated

PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v
Ran 79 tests ... OK

PYTHONPATH=backend .venv/bin/python -m compileall -q backend/app backend/scripts backend/tests

npm run build
Vite build completed successfully.

npm audit --omit=dev --audit-level=high
found 0 vulnerabilities

git diff --check
static scan of added lines: no findings

Ruff, mypy, and ESLint were not installed/configured in this checkout and were skipped without masking an exit code.

Regression coverage

  • Replay reads through VintageStore.load_snapshot().
  • Tampered normalized vintage snapshots return HTTP 422 from the replay endpoint.
  • Research report content tampering is rejected by the report hash.
  • Research manifest report-hash tampering is rejected.
  • Research manifest metadata tampering is rejected by the manifest-entry hash.
  • Non-object report JSON is rejected as invalid rather than raising an uncaught attribute error.
  • Non-object manifests and malformed manifest entries are rejected by load_manifest(), list_runs(), and latest().
  • list_runs() verifies each report and manifest-entry hash before returning metadata.
  • persist() validates every existing manifest entry before returning or writing a new report, and refuses malformed entries without leaving an orphan report.
  • Source replayability uses the same verified load-and-compute helper as the replay endpoint, so semantically invalid snapshots are marked unreplayable.
  • Tourism replay computation rejects non-list and non-object observation/exposure shapes as controlled semantic validation failures instead of leaking AttributeError.
  • Replay storage I/O failures return HTTP 422 instead of leaking an uncontrolled 500.
  • Explicit legacy report migration validates identity/metadata, is resumable after an interrupted report write, and is not automatic.
  • Direct report loads bind the manifest entry identity to the requested run_id.
  • Invalid UTF-8 in persisted manifest/report files is converted to ResearchRunError instead of escaping as a decode exception.
  • VintageStore rejects non-object snapshots with VintageStoreError.
  • Invalid UTF-8 in vintage manifests/snapshots is converted to VintageStoreError.
  • Raw vintage payload read failures are converted to VintageStoreError.
  • Research lineage includes normalized snapshot hashes and algorithms.

Runtime smoke

GET /api/v1/research/tourism/latest → HTTP 200; report_hash present
GET /api/v1/backtest/tourism?min_events=12 → HTTP 409; reason=insufficient_vintages

BOT collector rerun → same published release, `seen_count=14`, independent release count remains 1, gate remains blocked.

Independent review

No valid approval for the current tree. deleg_c227c439 found a persist-wide manifest-validation gap in the prior staged tree; the gap is now covered by a regression test and fixed before writing new reports. deleg_cf1da27d timed out after 600 seconds with no schema-valid JSON verdict, so it is inconclusive and does not approve the tree. deleg_feb40653 was dispatched against the pre-semantic-fix tree, reported truncated cached diffs, and is stale/non-approving. deleg_87a0f686 is the current bounded post-fix exact-diff review; no schema-complete current-tree approval has been surfaced yet.

Scope boundary

This remediation improves local artifact integrity only. It does not unlock the backtest: the current evidence remains below 12 independent releases and price data remains point_in_time=false.