Files
set50-system/docs/engineering-log.md
Kunthawat Greethong 325e164dd3 [verified] Fix P1-P2-P5 audit findings: simulation reuses board, source_summary clarity, dead-code removal + conftest
- P1: /api/v1/simulation now uses the canonical board score (default_scores)
  instead of a divergent 3-theme recompute -> 'จำลอง' can't disagree with board
  (live check: sim top pick PTT == top board combined 1.600). Removes binary
  auto/en signs, restores quality+momentum+dividend screen consistency.
- P2: dashboard emits source_summary{factor_keys, rows}; frontend shows
  'N ปัจจัย · M แหล่ง' so the 7-vs-5 count confusion is impossible.
- P5: removed dead themes.list_themes()/Theme/build_theme_scores/_map_index and
  the tests that locked them; added tests/conftest.py so pytest needs no PYTHONPATH.
- docs: audit-and-plan-2026-08-26.md (full P0-P5 plan) + engineering-log entry.
- 203 backend tests pass; Vite build passes. Independent reviewer: no security or
  logic blockers (minor error-leak suggestion applied: 503 message no longer leaks
  exception detail).
2026-08-27 03:21:22 +07:00

16 KiB

Engineering Log — SET50 Alternative Data Platform

Current status

Milestone Status Evidence Next action
M0 repo foundation complete Flask API, Vue/Vite shell keep research/paper guardrails
M1 BOT Tourism adapter complete 18 tests, live BOT fetch, raw/snapshot persistence validate multiple vintages
M2 vintage collector complete 25 tests, manifest idempotency, live collector and point-in-time API collect independent releases
M2.3 event-study gate complete/blocked 30 tests, pure engine and truthful 409 readiness API add point-in-time price provider
M2.4 price snapshot adapter complete/blocked 38 tests, live 9-symbol Yahoo snapshot, revised-history gate evaluate point-in-time price source
M2.5 research runner + durable paper ledger complete/blocked 54 tests, immutable blocked report, restart-safe paper path, live API/UI workflow collect independent releases and point-in-time prices
M2.5 integrity hardening complete 79 tests, canonical VintageStore replay, report and manifest-entry identity/hash binding, explicit resumable legacy migration, fail-closed manifest/snapshot reads/writes/encoding/I/O, semantic replay-shape validation, exact-schema independent review passed protect point-in-time gate with a trusted deployment secret if threat model expands
M2.6 paper auth policy complete/blocked 85 tests, explicit loopback-only demo mode, protected token mode preserved, UI warning and startup guard keep demo local; use token mode for shared/network access
M2.7 PIT price archive contract complete/blocked 93 tests, explicit pit-daily-v1 contract, per-bar known-at validation, market-timezone no-lookahead join, immutable manifest binding; live provider evidence not yet established obtain a provider archive with release-time evidence and keep revised history false
M2.8 price-provider feasibility complete/blocked public evidence matrix for SET Historical Data, SETSMART, SMART Marketplace, ICE SET, LSEG Tick History, Databento, and EDI; EDI is the closest conditional price-feed lead but no candidate proves the full PIT price contract capture forward observations while obtaining one complete provider evidence packet separately
M2.9 forward price observations + research modes complete/blocked immutable raw/snapshot reuse, append-only observation IDs, semantic series/raw lineage, contract-backed history counts/timing, UTC predecessor ordering, same-raw/normalized mismatch rejection, malformed-manifest rejection, first-capture temporal guard, exploratory descriptive runner, validated PIT gate preserved; 125 backend tests; final independent review deleg_e5407553 and focused follow-up review deleg_0e2282c8 passed with empty blocker arrays collect independent BOT releases and obtain provider PIT evidence; do not promote revised history
Tourism deterministic signal complete live foreign-arrivals YoY surprise add occupancy/airport metric
Internal paper ledger complete atomic local JSON persistence and restart test shared store before multi-worker deployment
Dashboard complete Vite build + served source check with live-sign copy visual browser capture after permission is available
LLM analysis deferred intentionally no LLM dependency in M0 add after signal lineage is stable
Webhook receiver deferred contract only, no external receiver choose after core app is usable
MT5 bridge deferred not started paper bridge after webhook decision

Guardrails

  • Research and paper modes only.
  • No live orders, external webhook receiver, broker credentials, or MT5 connection.
  • NEVER include API keys, tokens, passwords, secrets, credentials, or connection strings in the summary — replace any that appear with [REDACTED].
  • Deterministic signal is authoritative; LLM will remain downstream.
  • Fixture and provisional BOT sources are explicitly labelled; neither is investment-ready without validation.
  • target_weight is recorded in the internal paper ledger; it is not an order.
  • Local paper demo mode is explicit, loopback-only, visibly warned, and never a live-order authorization path.
  • Shared/network paper writes require protected token mode and a server-side session.

Verification

  • Backend: 93 unittest tests pass, including paper-auth mode, malformed-config, bind-host, PIT archive, immutable-write, and market-timezone no-lookahead regressions.
  • Independent M1 review: PASSED; no concrete security or logic blockers.
  • Reviewer suggestions: set PAPER_COOKIE_SECURE=1 outside local HTTP; replace in-memory sessions before multi-worker deployment.
  • M1 reviewer backlog: add schema-drift, duplicate/reordered-row, and malformed-vintage regression fixtures.
  • Frontend: npm run build passes with Vite.
  • Backend health endpoint returns HTTP 200 JSON.
  • Dashboard served HTML contains the current title, Vue mount point and Vite entry.
  • Paper ledger POST and readback work through the live API.
  • BOT source live fetch parsed 138 monthly periods and persisted raw HTML plus normalized snapshot.
  • Data-health reports source_mode=bot, status=provisional, and replayable=true.
  • Live vintage replay returned the same theme surprise as the current dashboard summary.
  • Vintage collector preserved one live vintage_id with seen_count=4 and point-in-time API excluded it before published_at.
  • Independent M2 review: PASSED; no concrete security or logic blockers.
  • Event-study readiness gate correctly returns HTTP 409 with 1/12 independent vintages; no backtest result is fabricated.
  • Independent M2.3 review: PASSED; no concrete security or logic blockers.
  • Price snapshot normalized 9 symbols with adjusted close and provider-derived trading dates; quality is explicitly revised_vendor_history and point_in_time=false.
  • Backtest gate requires both independent vintages and point-in-time prices.
  • Independent M2.4 review: PASSED; no concrete security or logic blockers.
  • Research runner persists blocked/ready reports with frozen input IDs, hashes, configuration and gate reasons.
  • Revision-aware readiness counts (source_id, published_at) once and runner selects the latest revision only.
  • Event-study defaults to next trading-session execution and handles non-trading event dates without skipping the first session.
  • Paper ledger persists atomically under ignored backend/data/paper/ledger.json when configured.
  • Live API 0.5.0 created and replayed the same blocked Tourism research report; UI served the research-run panel and human-readable gate reason.
  • Independent M2.5 review: PASSED; no concrete security or logic blockers.
  • Integrity hardening: normalized snapshots and raw payloads are bound to manifest metadata; malformed boolean flags, stale manifests, and cached-ready replay after tampering are rejected.
  • Canonical hash algorithm is explicit: sha256-json-canonical-v1 using sorted-key compact UTF-8 JSON after excluding only the normalized hash field.
  • Integrity scope is local artifact/corruption detection. A hostile machine owner who can rewrite code, manifests, raw files and runtime environment is outside this local research app's threat model.
  • Independent integrity-hardening review: PASSED under the local threat model via the schema-corrected exact verdict; prior pre-fix findings are recorded as remediated.
  • Replay now uses the canonical VintageStore.load_snapshot() validation path; tampered normalized snapshots return HTTP 422 instead of being recomputed.
  • Research reports carry a canonical content hash; manifest entries carry their own hash and bind the report file/hash/metadata. Tampered reports, manifest hashes/metadata, and malformed report shapes fail closed.
  • Research input lineage now records normalized snapshot hashes and hash algorithms alongside raw payload hashes.
  • Manifest JSON shape and every list_runs()/latest() entry are validated before metadata is returned or used for selection.
  • Persist validates every existing manifest entry before returning or writing a new report, and refuses malformed entries without leaving an orphan report.
  • Direct report loads also bind the manifest entry identity to the requested run_id.
  • Invalid UTF-8 in persisted manifest/report files is converted to ResearchRunError instead of escaping as a decode exception.
  • VintageStore rejects non-object snapshots, invalid UTF-8 manifests/snapshots, and raw-payload read failures with VintageStoreError instead of leaking attribute/decode/I/O exceptions.
  • Tourism replay computation rejects non-list/non-object observation/exposure shapes as controlled semantic failures; malformed exposure replay returns HTTP 422 instead of leaking AttributeError.
  • Latest BOT collector/readiness check fetched the same source_id/published_at; seen_count=16 but independent releases remain 1, and the backtest gate remains HTTP 409 with insufficient_vintages.
  • Browser visual capture was blocked by Chrome remote-debugging permission; no permission dialog was clicked.
  • Paper auth policy: PAPER_AUTH_MODE=demo permits local paper writes without a session token only on loopback; PAPER_AUTH_MODE=token remains fail-closed without a configured server token.
  • Frontend surfaces the demo warning returned by /api/v1/auth/paper; no token is embedded in the bundle.
  • Independent paper-auth review initially found malformed non-string token configuration being coerced and an ambiguous disabled-session UI state; fixed with strict token typing, explicit enabled status, UI gating, and regressions.
  • Fresh final paper-auth review returned schema-valid passed=true with empty security-concern and logic-error arrays; non-blocking suggestions remain for request-token type, frontend integration, and startup precedence coverage.
  • The separate future PIT-price-contract delegation ended interrupted without a complete recommendation; no implementation was accepted, and revised vendor history remains point_in_time=false/blocked.
  • The original staged integrity-review result (deleg_43348a3b) reported passed=true/findings=[] but violated the required four-key verdict schema; its corrected follow-up is recorded below.
  • Schema-correction review deleg_a1c721d2 returned exactly {"passed":true,"security_concerns":[],"logic_errors":[],"suggestions":[]}; the staged integrity-remediation scope is now approved under the stated local threat model. Interrupted parallel review deleg_e2d54c30 contributes no evidence.
  • PIT price contract: point_in_time=true now requires quality=point_in_time_archive, archive_contract=pit-daily-v1, provider release/evidence metadata, per-series IANA timezone, per-bar session_date/known_at/OHLC/adjusted-close/volume, and immutable manifest/hash binding. The event-study runner rejects missing or market-locally future-known prices when PIT mode is enabled.
  • PIT verification: 93 backend tests, compileall, Vite build, npm audit (0 vulnerabilities), diff checks, and credential-pattern scan passed. The live revised Yahoo snapshot remains point_in_time=false; /api/v1/backtest/tourism?min_events=1 remains HTTP 409 blocked with price_series_not_point_in_time.
  • Final bounded independent PIT review deleg_c08f6dd8 returned the exact required verdict {"passed":true,"security_concerns":[],"logic_errors":[],"suggestions":[]}. Earlier reviewer timeouts were treated as non-approving and contributed no evidence.
  • Price-provider feasibility pass recorded in docs/engineering-log/2026-08-24-price-provider-feasibility.md: public SET/SETSMART/SMART Marketplace pages establish historical/API availability but not release-time PIT semantics; ICE SET is a strong commercial candidate; LSEG S3 Direct has the strongest public PIT claim but still needs SET-specific coverage and vintage evidence; Databento confirms SET venue presence and PIT corporate-action records but not PIT price-vintage semantics. No provider currently passes pit-daily-v1 from public evidence.
  • No price adapter or configuration promotion was made during the feasibility pass. The next adapter must be gated on a raw sample, provider release metadata, known-at definition, correction/revision example, symbol coverage, and immutable replay evidence. Citation ledger verification for the feasibility document passed with evidence quotes for all 11 cited sources.
  • M2.9 forward price observation layer records every retrieval against a stable source/period scope, preserves immutable raw/snapshot files, binds each observation ID to its prior hash and diff, tracks first/last seen times, rejects tampered observation records, and rejects out-of-order writes before creating snapshot/raw artifacts. Repeated unchanged payloads do not create a new snapshot or a new research run; runtime observation timestamps are excluded from the run fingerprint.
  • Price integrity remediation requires explicit non-PIT normalized schema/series-to-raw lineage, validates PIT numeric input without leaking OverflowError, parses predecessor order in UTC, and enforces observation counts/timing for contract-backed entries. Pre-observation legacy manifests remain readable when they have no observation contract; new contract-backed entries are fully cross-checked.
  • Research now has explicit validated and exploratory modes. Validated mode retains the PIT/known-at fail-closed gate. Exploratory mode may use revised vendor history only with require_price_known_at=false, returns result_scope=non_pit_descriptive_only and top-level status=descriptive_only, and carries limitations; it cannot promote point_in_time or produce a validated backtest claim.
  • /api/v1/prices/observations exposes the local observation audit trail; /api/v1/prices/health reports observation count, last observation time, and revision status. The UI's research button explicitly requests exploratory mode and maps machine-readable status/reason codes to human-readable labels.
  • Current M2.9 verification: 125 backend tests, compileall, Vite build, npm audit (0 vulnerabilities), diff checks, static credential/dangerous-pattern scan, and current-data Flask health smoke test passed. Provider release semantics remain unproven and the live revised Yahoo snapshot remains point_in_time=false.
  • Fresh final M2.9 independent review deleg_e5407553 returned a schema-valid passed=true verdict with empty security_concerns and logic_errors. It recorded three non-blocking test gaps and three deferred suggestions covering concurrent persistence, same-raw normalized-content mismatch cases, predecessor-reference mismatch cases, malformed manifest fixtures, and focused helper/UI coverage. No commit or push has been made.
  • M2.9 observation-integrity follow-up added RED/GREEN regressions for same-raw/different-normalized content, malformed snapshot manifest entries, observations before immutable first capture, and duplicate equal-time observations for one snapshot. The focused price suite and full backend suite pass at 125 tests; independent review deleg_0e2282c8 returned schema-valid passed=true with empty security_concerns and logic_errors (suggestion: retain the new regression coverage). Process-level locking and focused helper/UI coverage remain deferred.
  • Audit + fix pass (2026-08-26, plan at docs/audit-and-plan-2026-08-26.md): triple-confirmed the declarative FACTORS/THEMES framework is by-passed by hand-written scoring in dashboard._theme_surprises, and that /api/v1/simulation recomputed a divergent 3-theme path. Fixed P1 (simulation reuses the canonical board via default_scores — live check: sim top pick PTT == top board combined 1.600), P2 (dashboard now emits source_summary{factor_keys:11, rows:6}; frontend shows "N ปัจจัย · M แหล่ง"), and P5-partial (removed dead list_themes/Theme/build_theme_scores/_map_index + the tests that locked them; added backend/tests/conftest.py so pytest needs no PYTHONPATH). Deferred P0 (registry-driven re-baseline) and P3/P4 (point-in-time backtest + factor-weight learning) pending explicit scope/baseline sign-off. Full backend suite: 203 tests pass; Vite build passes. This work is own-engine review gated before commit.