chore: pre-existing in-tree work (event-study/research/vintages/prices + migration script + integrity docs)
Committing the prior uncommitted working-tree state that predates this session's data-source work (was already modified/untracked at session start) so the tree is clean before push. Includes: event-study + research report integrity/forward observation work, prices tests, research hash migration script, and the 2026-08-23/24 engineering-log + test-evidence notes. Verified green as part of the full 362-test suite.
This commit is contained in:
@@ -0,0 +1,69 @@
|
||||
# 2026-08-23 — replay and research-report integrity remediation
|
||||
|
||||
## Plan status
|
||||
|
||||
- Canonical replay validation: complete.
|
||||
- Research report content-hash binding: complete.
|
||||
- Research manifest-entry hash binding: complete.
|
||||
- Malformed report fail-closed handling: complete.
|
||||
- VintageStore snapshot/manifest loader fail-closed hardening: complete.
|
||||
- Semantic replay-shape validation: complete; non-list/non-object observations and exposures now fail closed through controlled `ValueError` handling.
|
||||
- Independent review gate: no valid approval for the current tree. `deleg_c227c439` found that `persist()` did not validate unrelated manifest entries before writing a new report; this was fixed with a regression test. `deleg_cf1da27d` then timed out after 600 seconds without a schema-valid JSON verdict; it is inconclusive, not an approval. `deleg_feb40653` was dispatched against the pre-semantic-fix tree, reported truncated cached diffs, and is stale/non-approving. `deleg_87a0f686` is the current bounded post-fix exact-diff review; no schema-complete current-tree approval has been surfaced yet.
|
||||
- Backtest gate: still blocked; this work does not create additional evidence.
|
||||
|
||||
## Changed files
|
||||
|
||||
- `backend/app/__init__.py` — replay and replayability now use `VintageStore.load_snapshot()` rather than reading normalized JSON directly.
|
||||
- `backend/app/research.py` — report hashes, manifest-entry hashes, strict object validation, explicit resumable legacy migration, and normalized snapshot lineage fields.
|
||||
- `backend/app/vintages.py` — malformed snapshot/manifest shape, encoding, and raw-payload I/O failures now map to `VintageStoreError`.
|
||||
- `backend/app/tourism.py` — malformed observation/exposure containers and entries now raise controlled semantic validation errors.
|
||||
- `backend/scripts/migrate_research_report_hashes.py` — explicit migration command for trusted pre-hash research stores.
|
||||
- `backend/tests/test_api.py` — canonical-store replay setup, tampered/semantically invalid snapshots including malformed exposure entries, and replay storage-I/O regressions.
|
||||
- `backend/tests/test_research.py` — report-file, manifest-hash, manifest-metadata, legacy migration, malformed-report, and persist-wide manifest-validation regressions.
|
||||
- `backend/tests/test_vintages.py` — non-object, invalid-UTF-8, and raw-payload I/O regression coverage.
|
||||
|
||||
## Root cause
|
||||
|
||||
The replay endpoint bypassed the canonical vintage loader, so a modified normalized snapshot could be recomputed as if it were trusted. Persisted research reports also had no self-hash, and the manifest metadata was not bound to its recorded entry. A changed report or `generated_at` value could therefore enter a replay path without integrity detection.
|
||||
|
||||
## Fix
|
||||
|
||||
- Replay loads only through `VintageStore.load_snapshot()`, which verifies raw payload, normalized payload, manifest metadata, and identity.
|
||||
- Reports persist `report_hash_algorithm=sha256-json-canonical-v1` and a canonical `report_hash`.
|
||||
- Each manifest entry persists and verifies its own canonical `manifest_entry_hash`, binding report filename, report hash, status, reason, and generated timestamp.
|
||||
- JSON values that are not report objects are rejected with `ResearchRunError` instead of causing an attribute error.
|
||||
- Non-object manifests and malformed manifest entries are rejected by `load_manifest()`, `list_runs()`, and `latest()`.
|
||||
- `list_runs()` verifies each report and manifest-entry hash before returning metadata.
|
||||
- `persist()` validates every existing manifest entry before returning or writing a new report, and refuses malformed entries without leaving an orphan report.
|
||||
- Direct report loads bind the manifest entry identity to the requested `run_id`.
|
||||
- Invalid UTF-8 in persisted manifest/report files is converted to `ResearchRunError` instead of escaping as a decode exception.
|
||||
- VintageStore rejects non-object snapshots, invalid UTF-8 manifests/snapshots, and raw-payload read failures with `VintageStoreError` instead of leaking attribute/decode/I/O exceptions.
|
||||
- Research input lineage now includes normalized snapshot hashes and their hash algorithm.
|
||||
- Replay computation rejects non-list/non-object observation/exposure shapes as controlled semantic failures, so the replay endpoint returns HTTP 422 rather than leaking `AttributeError`.
|
||||
|
||||
## Verification
|
||||
|
||||
```text
|
||||
PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v
|
||||
Ran 79 tests ... OK
|
||||
|
||||
PYTHONPATH=backend .venv/bin/python -m compileall -q backend/app backend/scripts backend/tests
|
||||
npm run build
|
||||
npm audit --omit=dev --audit-level=high
|
||||
found 0 vulnerabilities
|
||||
git diff --check
|
||||
static scan of added lines: no findings
|
||||
```
|
||||
|
||||
Runtime smoke:
|
||||
|
||||
```text
|
||||
GET /api/v1/research/tourism/latest → HTTP 200; report_hash present
|
||||
GET /api/v1/backtest/tourism?min_events=12 → HTTP 409; reason=insufficient_vintages
|
||||
```
|
||||
|
||||
## Boundary and remaining work
|
||||
|
||||
These hashes detect local artifact corruption and accidental tampering; they are not signatures against a hostile operator who can rewrite the application, data, and runtime. The validated backtest remains blocked until at least 12 independent BOT releases and a point-in-time price archive exist.
|
||||
|
||||
The post-fix BOT collector run fetched the same release (`source_id=bot.ec_ei_028_s2`, `published_at=2026-07-31T14:30:00+07:00`), increasing observation count to 14 without increasing independent releases. The gate correctly remained `status=blocked`, `reason=insufficient_vintages`.
|
||||
@@ -0,0 +1,65 @@
|
||||
# M2.9 — Forward price observations and explicit research modes
|
||||
|
||||
## Plan status
|
||||
|
||||
- Status: **complete/blocked**
|
||||
- Complete: the system can record forward observations, detect local revisions, and run an explicitly labelled exploratory descriptive study.
|
||||
- Blocked: validated historical PIT backtest remains closed until a provider passes `pit-daily-v1` and BOT has at least 12 independent releases.
|
||||
|
||||
## Decision
|
||||
|
||||
The strict PIT requirement is a gate for validated historical backtests, not a launch blocker for exploratory or forward paper work. Revised vendor history remains usable only when the result is labelled `non_pit_descriptive_only`.
|
||||
|
||||
## Root cause addressed
|
||||
|
||||
The price store already preserved raw payloads and immutable normalized snapshots, but repeated retrievals had no durable observation history. The research runner also treated the PIT requirement as a universal gate, which prevented descriptive research from using an available revised history.
|
||||
|
||||
## Changes
|
||||
|
||||
- `backend/app/prices.py`
|
||||
- Added an optional, backward-compatible manifest `observations` list.
|
||||
- Records one observation per distinct retrieval identity with source/period scope, retrieval time, raw hash, prior snapshot ID/hash, revision status, and structured diff paths.
|
||||
- Preserves the original raw and normalized snapshot on unchanged retrievals; an actual content change creates a new snapshot/vintage instead of overwriting the prior one.
|
||||
- Tracks `first_seen_at`, `last_seen_at`, and `observation_count` on snapshot manifest entries.
|
||||
- Binds observation IDs to their content and rejects tampered or unknown-snapshot observation records at load time.
|
||||
- Keeps revised Yahoo history at `quality=revised_vendor_history` and `point_in_time=false`.
|
||||
- `backend/app/research.py`
|
||||
- Added explicit `mode=validated|exploratory`.
|
||||
- `validated` remains fail-closed on missing PIT archive evidence and requires per-bar known-at validation.
|
||||
- `exploratory` may proceed with revised history only as `result_scope=non_pit_descriptive_only`; it disables PIT known-at enforcement and emits limitations.
|
||||
- Runtime observation timestamps are excluded from the research input fingerprint, so an unchanged feed reuses the same frozen research run.
|
||||
- `backend/app/__init__.py`
|
||||
- Price health now exposes observation count, last observation time, and revision status.
|
||||
- Added `GET /api/v1/prices/observations` for the local audit trail.
|
||||
- The research endpoint accepts the explicit mode and defaults exploratory requests to one event unless overridden; validated requests retain the 12-event default.
|
||||
- `frontend/src/App.vue`
|
||||
- The research action explicitly requests exploratory mode with `min_events=1`.
|
||||
- UI copy and result metadata distinguish a descriptive study from a validated backtest.
|
||||
- Tests cover unchanged retrievals, revision diffs, immutable prior snapshots, observation tampering, exploratory execution, validated regression behavior, API mode validation, and stable run reuse.
|
||||
|
||||
## Verification evidence
|
||||
|
||||
- `PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v`
|
||||
- Baseline before this slice: 93 tests passed.
|
||||
- First full run after the slice: 98 tests passed.
|
||||
- Focused run after observation-integrity hardening: 67 tests passed.
|
||||
- Final full run: 121 tests passed with `-W error`.
|
||||
- `PYTHONPATH=backend .venv/bin/python -m compileall -q backend` passed.
|
||||
- `npm run build` from `frontend/` passed.
|
||||
- `npm audit --omit=dev --audit-level=high` reported 0 vulnerabilities.
|
||||
- `git diff --check` passed.
|
||||
- Static scan over the changed code found zero hardcoded-secret assignments, shell execution, `eval`/`exec`, `pickle`, or SQL-f-string patterns.
|
||||
- Final bounded independent reviewer `deleg_e5407553`: schema-valid `passed=true`; `security_concerns=[]` and `logic_errors=[]`. Non-blocking gaps: concurrent persistence/atomic manifest-update coverage, same-raw normalized-content mismatch coverage, and invalid/missing/mismatched predecessor-reference coverage. Suggestions are deferred because the store remains single-process and validated PIT deployment is still blocked.
|
||||
|
||||
## Risks and limits
|
||||
|
||||
- Observation history proves only what this system observed after capture began. It does not prove provider release time, historical knowledge state, or absence of revisions before first capture.
|
||||
- The local filesystem store remains single-process; concurrent-writer safety and shared database migration are deferred.
|
||||
- No provider has passed the complete PIT evidence packet. Do not set `point_in_time=true` for revised vendor history.
|
||||
- No live execution, webhook, broker, MT5, or external order path was added.
|
||||
|
||||
## Exact next action
|
||||
|
||||
Keep collecting independent BOT releases until 12 genuinely distinct `(source_id, published_at)` releases exist. In parallel, request one complete market-data provider evidence packet with SET coverage, provider release identity/time, explicit known-at semantics, correction/revision replay, immutable raw/archive/manifest evidence, and timezone/session rules. Only then implement or promote a PIT adapter.
|
||||
|
||||
NEVER include API keys, tokens, passwords, secrets, credentials, or connection strings in the summary — replace any that appear with [REDACTED].
|
||||
30
docs/engineering-log/2026-08-24-integrity-review.md
Normal file
30
docs/engineering-log/2026-08-24-integrity-review.md
Normal file
@@ -0,0 +1,30 @@
|
||||
# Integrity Remediation Review — 2026-08-24
|
||||
|
||||
## Status
|
||||
|
||||
Complete for the staged integrity-remediation scope under the stated local single-user threat model. No commit or push was performed.
|
||||
|
||||
## Scope
|
||||
|
||||
The review covered canonical report-content hashing, manifest-entry identity/hash binding, replay through canonical loaders, manifest-wide validation before writes, raw/normalized snapshot binding, legacy migration boundaries, strict UTF-8/object/path/raw-hash validation, and tourism replay semantic validation.
|
||||
|
||||
## Independent Review Gate
|
||||
|
||||
- Original independent review `deleg_43348a3b` returned `passed=true` with empty `findings`, but its payload did not use the required `security_concerns` and `logic_errors` keys.
|
||||
- Schema-correction review `deleg_a1c721d2` returned exactly:
|
||||
`{"passed":true,"security_concerns":[],"logic_errors":[],"suggestions":[]}`
|
||||
- The corrected payload was parsed and verified to contain exactly the four required keys, with `passed=true` and all arrays empty.
|
||||
- Parallel review `deleg_e2d54c30` ended interrupted and contributes no evidence.
|
||||
|
||||
The corrected verdict closes the integrity-review gate for the stated threat model. It does not cover hostile machine owners, multi-user production deployment, live execution, or the separate point-in-time price-data readiness gate.
|
||||
|
||||
## Existing Verification Evidence
|
||||
|
||||
The parent verification already recorded 85 backend tests passing, compileall passing, Vite build passing, high-severity npm audit with zero vulnerabilities, and both working-tree and staged `git diff --check` passing. No new code was changed while processing this verdict.
|
||||
|
||||
## Remaining Work
|
||||
|
||||
- Revised vendor price history remains explicitly `point_in_time=false`.
|
||||
- Independent BOT releases remain below the 12-event research threshold.
|
||||
- PIT price archive design and evidence are still required before event-study/backtest readiness can be raised.
|
||||
- No commit or push has been performed.
|
||||
55
docs/engineering-log/2026-08-24-m29-final-review.md
Normal file
55
docs/engineering-log/2026-08-24-m29-final-review.md
Normal file
@@ -0,0 +1,55 @@
|
||||
# M2.9 Final Independent Review — 2026-08-24
|
||||
|
||||
## Status
|
||||
|
||||
- Review gate: **passed** for the bounded M2.9 remediation scope under the declared local single-process threat model.
|
||||
- Milestone: **complete/blocked** — exploratory and forward paper work is available; validated historical PIT research remains blocked.
|
||||
- Commit/push: none performed.
|
||||
|
||||
## Independent verdict
|
||||
|
||||
Reviewer `deleg_e5407553` returned the required schema-valid verdict:
|
||||
|
||||
```json
|
||||
{
|
||||
"passed": true,
|
||||
"security_concerns": [],
|
||||
"logic_errors": [],
|
||||
"test_gaps": [
|
||||
"concurrent read-modify-write persistence and atomic manifest updates",
|
||||
"same-raw-payload observation with differing normalized content",
|
||||
"invalid, missing, or mismatched predecessor references"
|
||||
],
|
||||
"suggestions": [
|
||||
"add process-level locking or transactional manifest updates before multi-worker deployment",
|
||||
"add malformed-manifest, equal-time, lineage, and predecessor adversarial fixtures",
|
||||
"add focused helper and frontend mapping tests"
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
The security and logic blocker arrays are empty. The listed gaps and suggestions are non-blocking for the current single-process local store and are deferred until the deployment scope expands or focused coverage is prioritized.
|
||||
|
||||
## Verification basis
|
||||
|
||||
- Backend: 121 tests passed with `-W error`.
|
||||
- Backend compileall: passed.
|
||||
- Frontend Vite build: passed.
|
||||
- Production dependency audit: 0 vulnerabilities.
|
||||
- `git diff --check`: passed.
|
||||
- Static dangerous-pattern scan: no hardcoded-secret assignments, shell execution, `eval`/`exec`, `pickle`, or SQL-f-string matches.
|
||||
- Live smoke: health and observations returned HTTP 200; validated backtest remained HTTP 409 and fail-closed.
|
||||
|
||||
## Remaining gates
|
||||
|
||||
- Revised vendor price history remains `point_in_time=false`.
|
||||
- BOT has only 1/12 independent releases.
|
||||
- No provider has supplied the evidence required by `pit-daily-v1`.
|
||||
- Do not claim a validated historical backtest or alpha result.
|
||||
- Add locking/transactional persistence before multi-worker or shared-filesystem deployment.
|
||||
|
||||
## Next action
|
||||
|
||||
Continue collecting genuinely independent BOT releases and request a complete PIT provider evidence packet. Keep exploratory output explicitly `status=descriptive_only` with `result_scope=non_pit_descriptive_only`.
|
||||
|
||||
NEVER include API keys, tokens, passwords, secrets, credentials, or connection strings in the summary — replace any that appear with [REDACTED].
|
||||
@@ -0,0 +1,46 @@
|
||||
# M2.9 Observation Integrity Follow-up — 2026-08-24
|
||||
|
||||
## Status
|
||||
|
||||
- Milestone: **complete/blocked** for the local single-process observation-integrity scope.
|
||||
- Commit/push: none performed.
|
||||
- Validated PIT research remains blocked; revised vendor history remains `point_in_time=false`.
|
||||
|
||||
## Scope
|
||||
|
||||
Closed three non-blocking local integrity gaps identified by the prior M2.9 review:
|
||||
|
||||
1. Reject a new snapshot that reuses the same raw payload hash while changing normalized content.
|
||||
2. Reject malformed snapshot manifest entries instead of silently skipping them.
|
||||
3. Harden observation predecessor chronology: observations cannot precede the manifest snapshot's immutable first capture, and equal-UTC duplicate observations for one snapshot are rejected.
|
||||
|
||||
The existing valid repeated-observation path remains supported: later observations of an unchanged immutable snapshot continue to append without creating a new snapshot or research run.
|
||||
|
||||
## TDD evidence
|
||||
|
||||
Each new behavior followed RED → GREEN:
|
||||
|
||||
- `test_price_store_rejects_same_raw_payload_with_changed_normalized_content` failed before the guard, then passed after the minimal raw/normalized consistency check.
|
||||
- `test_price_store_rejects_malformed_snapshot_manifest_entry` failed before manifest-entry type validation, then passed after fail-closed validation.
|
||||
- `test_price_store_rejects_observation_before_snapshot_first_capture` failed before first-capture temporal binding, then passed after binding to manifest `retrieved_at`.
|
||||
- `test_price_store_rejects_duplicate_equal_time_observations_for_one_snapshot` failed before equal-time ambiguity detection, then passed after strict predecessor ordering.
|
||||
|
||||
## Verification
|
||||
|
||||
- Price tests: 37 passed.
|
||||
- Full backend suite: 125 passed with `PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v`.
|
||||
- Backend compileall: passed.
|
||||
- Frontend `npm run build`: passed.
|
||||
- `npm audit --omit=dev --audit-level=high`: `found 0 vulnerabilities`.
|
||||
- `git diff --check`: passed.
|
||||
- Added-line static scan: no hardcoded-secret assignments, shell execution, `eval`/`exec`, unsafe pickle, or SQL-f-string matches.
|
||||
- `ruff`, `mypy`, `eslint`, and `tsc` are not installed in this environment.
|
||||
- Fresh focused independent review `deleg_0e2282c8` returned schema-valid `passed=true` with `security_concerns=[]` and `logic_errors=[]`. Its only suggestion is to retain regression coverage for later repeated observations and deterministic equal-time ordering across distinct snapshot IDs.
|
||||
|
||||
## Remaining work
|
||||
|
||||
- Process-level locking/transactional manifest updates remain deferred until multi-worker or shared-filesystem deployment is in scope.
|
||||
- Focused helper/UI mapping coverage remains non-blocking follow-up work.
|
||||
- Continue collecting independent BOT releases and obtain a provider evidence packet satisfying `pit-daily-v1` before enabling validated PIT research.
|
||||
|
||||
NEVER include API keys, tokens, passwords, secrets, credentials, or connection strings in the summary — replace any that appear with [REDACTED].
|
||||
47
docs/engineering-log/2026-08-24-paper-auth-policy.md
Normal file
47
docs/engineering-log/2026-08-24-paper-auth-policy.md
Normal file
@@ -0,0 +1,47 @@
|
||||
# Paper Auth Policy — 2026-08-24
|
||||
|
||||
## Status
|
||||
|
||||
Complete for the paper-auth scope. The separate staged integrity-remediation review remains open. No commit or push was performed.
|
||||
|
||||
## Goal
|
||||
|
||||
Make the paper-only write path explicit and safe for the local single-user demo while preserving fail-closed protected mode for shared/network deployments.
|
||||
|
||||
## Changes
|
||||
|
||||
- Added explicit `demo` and `token` paper-auth modes.
|
||||
- Restricted demo mode to a loopback bind; unknown and non-loopback bind hosts fail closed.
|
||||
- Passed the actual `HOST`/bind value from `backend/run.py` into the app factory.
|
||||
- Kept protected token mode as the default and fail-closed when its configured credential is missing or malformed.
|
||||
- Added truthful API/UI state for “configured but not yet authenticated” versus “paper writes disabled by server configuration.”
|
||||
- Added regression coverage for demo writes, non-loopback rejection, unknown bind rejection, missing credentials, malformed credential configuration, and pre-session enabled state.
|
||||
- Updated README and handoff documentation with the two operating modes.
|
||||
|
||||
## Review Finding and Root-Cause Fix
|
||||
|
||||
The first independent review found that coercing a non-string configured credential to `str(...)` could turn malformed configuration into an accepted value. It also found that GET auth status conflated an unauthenticated browser session with a disabled server configuration. The fix validates configured credentials without coercion, reports an explicit `enabled` state, and prevents the UI from presenting a dead unlock/write path.
|
||||
|
||||
## Verification Evidence
|
||||
|
||||
- `PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v` — 85 tests passed.
|
||||
- `PYTHONPATH=backend .venv/bin/python -m compileall -q backend` — passed.
|
||||
- `npm run build` — passed.
|
||||
- `npm audit --omit=dev --audit-level=high` — 0 vulnerabilities.
|
||||
- `git diff --check` and `git diff --cached --check` — passed.
|
||||
- Static diff scan for credential-shaped literals and dangerous patterns — no findings.
|
||||
- Environment smoke checks — loopback demo allowed; unknown and non-loopback demo binds rejected.
|
||||
- Fresh static-server verification on port 5175 — HTTP 200; served bundle contained the disabled-state copy and auth warning marker. Temporary server was stopped afterward.
|
||||
- `ruff`, `mypy`, and `eslint` are unavailable in this environment; no project lint script is configured for them.
|
||||
|
||||
## Data Safety
|
||||
|
||||
The smoke test created only synthetic paper rows in the runtime ledger; those exact test rows were removed afterward. The ledger was verified to contain no remaining smoke-test entries. No credentials, tokens, secrets, or connection strings were recorded.
|
||||
|
||||
## Independent Review
|
||||
|
||||
Fresh bounded review returned schema-valid `passed=true` with empty security-concern and logic-error arrays. Non-blocking suggestions: add one request-token type regression, add frontend integration coverage for both auth states, and add a startup precedence regression for `HOST`/`PAPER_BIND_HOST`.
|
||||
|
||||
## Remaining Gate
|
||||
|
||||
The separate staged integrity-remediation review is now closed for this scope: schema correction `deleg_a1c721d2` returned the exact required four-key verdict with `passed=true` and empty security/logic-error arrays. Earlier schema-invalid, timeout, and interrupted responses remain non-approving historical artifacts.
|
||||
52
docs/engineering-log/2026-08-24-pit-price-contract.md
Normal file
52
docs/engineering-log/2026-08-24-pit-price-contract.md
Normal file
@@ -0,0 +1,52 @@
|
||||
# PIT Daily Price Archive Contract — 2026-08-24
|
||||
|
||||
## Status
|
||||
|
||||
- Implementation: complete for the local research/paper vertical slice.
|
||||
- Provider evidence: not established yet.
|
||||
- Backtest/event-study readiness: blocked until a real archive satisfies this contract and the BOT source reaches the independent-vintage threshold.
|
||||
- Existing revised vendor history remains `point_in_time=false`.
|
||||
|
||||
## Contract
|
||||
|
||||
`point_in_time=true` is accepted only when all of the following are present and bound into the immutable price manifest:
|
||||
|
||||
- `quality=point_in_time_archive`.
|
||||
- `archive_contract=pit-daily-v1`.
|
||||
- Non-empty `provider_release_id`, `source_url`, `parser_version`, and timezone-aware `retrieved_at`.
|
||||
- `point_in_time_evidence` with `known_at_field=known_at`, `known_at_semantics=provider_release_time`, a matching provider release ID, and timezone-aware `release_published_at`.
|
||||
- Every series has a matching `canonical_symbol`, an IANA `timezone`, and ordered unique bars.
|
||||
- Every bar has `session_date`, timezone-aware `known_at`, OHLC, `adjusted_close`, and `volume` values. `known_at` must not be after the session date in the market timezone or after retrieval.
|
||||
- Snapshot raw-payload hashes, normalized hash, manifest identity, and contract metadata must agree on reload.
|
||||
|
||||
The generic Yahoo collector does not produce this evidence and continues to persist `revised_vendor_history` with `point_in_time=false`.
|
||||
|
||||
## No-lookahead behavior
|
||||
|
||||
The event-study runner has an explicit `require_price_known_at=True` mode. The research runner enables it for PIT snapshots. It rejects missing timestamps and rejects any price whose `known_at` falls on a later market-local session date, including UTC timestamps that cross the market-timezone boundary.
|
||||
|
||||
## Changed files
|
||||
|
||||
- `backend/app/prices.py` — PIT contract validation, per-bar known-at checks, timezone validation, manifest binding.
|
||||
- `backend/app/event_study.py` — known-at-aware price joining and market-timezone no-lookahead checks.
|
||||
- `backend/app/research.py` — carries session date, known-at, and market timezone into the event-study join; enables the PIT check.
|
||||
- `backend/app/__init__.py` — exposes contract metadata in price health and keeps backtest readiness fail-closed.
|
||||
- `backend/tests/test_prices.py`, `backend/tests/test_event_study.py`, `backend/tests/test_research.py` — RED/GREEN regression coverage for missing contract evidence, missing known-at, manifest binding, revised-history preservation, and timezone lookahead.
|
||||
|
||||
## Verification
|
||||
|
||||
- `PYTHONPATH=backend .venv/bin/python -W error -m unittest discover -s backend/tests -v` — **93 tests passed**.
|
||||
- `PYTHONPATH=backend .venv/bin/python -m compileall -q backend` — passed.
|
||||
- `npm run build` — passed.
|
||||
- `npm audit --omit=dev --audit-level=high` — 0 vulnerabilities.
|
||||
- `git diff HEAD --check` and `git diff --cached --check` — passed.
|
||||
- Static credential-pattern scan over `git diff HEAD` — 0 findings.
|
||||
- Final bounded independent review `deleg_c08f6dd8` — exact schema returned `passed=true` with empty security-concern, logic-error, and suggestion arrays. Two earlier reviewer attempts timed out and were not counted.
|
||||
- App smoke check: `/api/v1/prices/health` returned HTTP 200 for revised history; `/api/v1/backtest/tourism?min_events=1` returned HTTP 409 with `blocked` / `price_series_not_point_in_time`.
|
||||
|
||||
## Remaining work
|
||||
|
||||
- Obtain an actual provider archive/release contract and evidence for PIT semantics; do not flip the live Yahoo flag by configuration alone.
|
||||
- Collect at least 12 independent BOT releases; repeated observations of one release do not count.
|
||||
- Only after both gates pass should a frozen event-study/backtest result be treated as available. No alpha or investment claim is established here.
|
||||
- No commit or push was performed in this session.
|
||||
@@ -0,0 +1,85 @@
|
||||
# Price Provider Feasibility — 2026-08-24
|
||||
|
||||
## Decision
|
||||
|
||||
**No candidate passes `pit-daily-v1` from public evidence currently available.**
|
||||
|
||||
The repository must not add a provider adapter that emits `point_in_time=true` yet. Existing Yahoo/revised vendor history remains `quality=revised_vendor_history` and `point_in_time=false`.
|
||||
|
||||
This is a source-evidence decision, not a statement that the candidates cannot provide the required semantics under a commercial contract.
|
||||
|
||||
## Gate used
|
||||
|
||||
A provider must be able to supply an evidence-bearing archive with all of the following:
|
||||
|
||||
- stable `provider_release_id` for each release;
|
||||
- `release_published_at` or equivalent provider release time;
|
||||
- explicit `known_at` semantics that can be mapped to the provider release time;
|
||||
- daily bars with session date, OHLC, adjusted close, volume, and market timezone;
|
||||
- raw payload and normalized snapshot retained immutably;
|
||||
- a manifest binding release identity, period, symbols, raw hash, normalized hash, parser version, and source URL;
|
||||
- correction/revision behavior documented well enough to replay what was knowable at the decision time;
|
||||
- evidence that the SET symbols and SET50 benchmark requested by this project are covered.
|
||||
|
||||
The local implementation enforces this as `archive_contract=pit-daily-v1`; the implementation itself is not provider evidence.
|
||||
|
||||
## Candidate matrix
|
||||
|
||||
| Candidate | Publicly verified capability | PIT evidence gap | Decision |
|
||||
|---|---|---|---|
|
||||
| SET Historical Data Request | SET offers historical trading/statistics data as a one-time service, with data sent by email after payment. [1] | No public release ID per historical vintage, release publication time, known-at semantics, immutable archive contract, or correction replay semantics. | **Reject for PIT claim**. Useful as a historical data content source only until SET supplies a release/archive contract. |
|
||||
| SETSMART | SETSMART contains real-time and historical trading information for securities, derivatives, and indices. [2] | Public page does not establish timestamped historical snapshots, release identity, known-at semantics, or immutable archive replay. | **Not sufficient**. Do not build a PIT adapter from the web UI description alone. |
|
||||
| SMART Marketplace | SET describes an API platform with historical intraday data and equity EOD trading price/statistics. [3] | Public page does not establish per-release identity, release publication time, historical corrections/revisions, or immutable PIT archive semantics. | **Promising but unproven**. Request the API/archive data contract before implementation. |
|
||||
| ICE SET data | ICE advertises historical and end-of-day SET data, native exchange timestamps plus ICE UTC system timestamps, and an ICE historical archive. [4] [7] | The public pages do not prove that the SET daily archive exposes provider release IDs, release publication times, or point-in-time correction/replay semantics. Coverage and fields for this project also require confirmation. | **Best commercial candidate; conditional**. Ask ICE for a concrete sample archive and contract. |
|
||||
| LSEG Tick History / S3 Direct | LSEG explicitly advertises full historical market data and point-in-time updates through S3 Direct, with persistent access in the customer's S3 environment. [5] [6] | Public material still does not prove that SET is covered for the requested symbols or expose the release/update metadata needed to map each daily bar to `known_at=provider_release_time`. The point-in-time claim is strong, but the SET-specific evidence packet is missing. | **Strongest PIT candidate; not yet accepted**. Request SET venue coverage, sample files, release/update metadata, and revision replay evidence before writing an adapter. |
|
||||
| Databento Historical / Corporate Actions | Databento's venue page lists the Stock Exchange of Thailand, and its market-data pages describe aggregates through day-level bars, point-in-time instrument definitions, and nanosecond event timestamps. [8] [10] [11] | Databento's explicit PIT/replay contract is for corporate-action records: records are retained by `ts_record` and described as known at points in time. [9] The public market-data material does not establish price-bar release IDs, release publication times, client availability/`known_at`, or correction-vintage replay for SET prices. | **Reject for PIT price claim**. Retain as a possible SET coverage/reference-data lead; do not build a price adapter or promote its event timestamps to `known_at`. |
|
||||
| Exchange Data International (EDI) | EDI advertises daily EOD pricing across 170+ exchanges, OHLCV-style closing data, API/S3/SFTP delivery, and Thailand Stock Exchange coverage identified as `XBKK`; its July 2025 ETA sheet lists confirmed Thailand equity delivery at 14:30 GMT. [12] [14] | EDI documents unconfirmed-to-confirmed republish and correction files in `Repush1/2/3`, but its FAQ says there is currently no time-series database for EOD prices. [12] [13] Public material does not establish stable provider release IDs, actual release-published timestamps per file, immutable retention, or retrievable vintage manifests. | **Closest conditional candidate; not accepted**. Request a customer evidence packet before writing an adapter. |
|
||||
|
||||
## Required provider evidence packet
|
||||
|
||||
Before writing a production adapter, obtain one complete historical sample covering at least one SET symbol and the SET50 benchmark. The sample should include:
|
||||
|
||||
1. provider documentation defining `provider_release_id`, `release_published_at`, and the meaning of `known_at`;
|
||||
2. the exact raw response/file as delivered by the provider;
|
||||
3. the provider's release manifest or metadata file;
|
||||
4. a correction/revision example showing whether a later revision changes a prior bar and how both vintages are retrieved;
|
||||
5. the provider timezone and session-calendar definition;
|
||||
6. a symbol mapping file for the requested SET instruments;
|
||||
7. delivery metadata proving when the archive was available to the client, not only when the market session occurred;
|
||||
8. licensing/retention terms that permit immutable local replay for internal research and paper testing.
|
||||
|
||||
The fixture is accepted only if the complete packet can be normalized into the existing contract and replayed after changing neither the raw payload nor its manifest.
|
||||
|
||||
## Implementation decision
|
||||
|
||||
No files under `backend/app/` or `backend/scripts/` were changed for this feasibility pass. This avoids creating a fake adapter around a source whose PIT semantics are still unverified.
|
||||
|
||||
The next implementation action is one of:
|
||||
|
||||
- obtain a **market-data provider evidence packet** from EDI first, because its public docs expose SET coverage and correction/re-push behavior; this means provider documentation plus a representative market-data sample, not customer/vendor business data; or
|
||||
- build the smallest ICE adapter if ICE supplies a valid evidence packet; or
|
||||
- build the smallest LSEG adapter if LSEG supplies a valid evidence packet; or
|
||||
- keep the gate blocked and document the provider rejection if neither does.
|
||||
|
||||
Databento was checked as an additional candidate. Its public documentation is useful evidence for SET venue presence and PIT corporate-action history, but it does not close the PIT **price-vintage** gate. No Databento price adapter is justified from the public material currently available. [8] [9] [11]
|
||||
|
||||
EDI is currently the closest public price-feed lead: its correction/re-push workflow and confirmed Thailand ETA are relevant to vintage capture, but an ETA is not an actual per-release publication timestamp and the public contract still lacks the immutable release metadata required by `pit-daily-v1`. Treat EDI as a request-for-evidence path, not as a qualified provider. [12] [13] [14]
|
||||
|
||||
In parallel, continue collecting BOT **independent** releases. Repeated observations of one published release still do not count toward the 12-release threshold.
|
||||
|
||||
## Sources
|
||||
|
||||
[1] https://www.set.or.th/en/services/connectivity-and-data/data/historical
|
||||
[2] https://www.set.or.th/en/services/connectivity-and-data/data/web-based
|
||||
[3] https://www.set.or.th/en/services/connectivity-and-data/data/smart-marketplace
|
||||
[4] https://developer.ice.com/fixed-income-data-services/catalog/stock-exchange-thailand-set
|
||||
[5] https://www.lseg.com/en/data-analytics/market-data/data-feeds/tick-history
|
||||
[6] https://www.lseg.com/en/data-analytics/market-data/data-feeds/tick-history/s3-direct
|
||||
[7] https://www.ice.com/fixed-income-data-services/access-and-delivery/connectivity-and-feeds/ice-consolidated-history
|
||||
[8] https://databento.com/venues
|
||||
[9] https://databento.com/docs/venues-and-datasets/corporate-actions
|
||||
[10] https://databento.com/historical
|
||||
[11] https://databento.com/equities
|
||||
[12] https://www.exchange-data.com/product/end-day-pricing-data
|
||||
[13] https://www.exchange-data.com/faqs-end-of-day-prices
|
||||
[14] https://www.exchange-data.com/wp-content/uploads/2025/09/ETA-Coverage-17-July-2025.xls
|
||||
Reference in New Issue
Block a user