fix(evolution): persist safe proposal diagnostics and validate input readiness - #183
Conversation
|
Review base: 2c15302 🤖 DeepSeek V4 Pro PR ReviewReviewed the diff end-to-end (CI, docs, migration 0059, evolver proposer + checkpoint persistence, research extractor, and the two new root scripts + tests). The change is careful and internally consistent (diagnostics are whitelisted and leak-tested, the DB column is bounded, the replay tool is pinned to loopback/audited targets). A few things are worth resolving before merge. Findings[medium] services/research/src/inalpha_research/event_extractor.py:99 — dropping [medium] scripts/check-e2-readiness.py:23 ( [medium] .github/workflows/ci.yml:247 — the new step runs repo-root Notes (below threshold, no action required)
|
Deploying inalpha-web with
|
| Latest commit: |
134498b
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://87d898d2.inalpha-web.pages.dev |
| Branch Preview URL: | https://codex-event-evidence-chain.inalpha-web.pages.dev |
Summary
Real crypto news names Bitcoin, Ethereum or Solana without tickers. Recognize explicit full names, require uppercase textual tickers, preserve structured symbols, and exclude Bitcoin Cash/SV/Gold and Ethereum Classic, including hyphenated names and URL slugs. New facts use
deterministic-event-extractor-v2; existing facts and snapshots remain immutable.Record real input validation and a paid-provider five-generation execution with restart recovery.
docs/04-current-state.mdowns the evidence and remaining acceptance work; the sanitized structured report isdocs/validation/e2-real-acceptance-2026-09-30.json. Event taxonomy, statistical thresholds, public APIs and credential grant flow are unchanged.Validation
Real E2 acceptance result
Real execution, checkpoint restart recovery and insufficient-evidence rejection passed. The complete positive E2 chain has not passed acceptance.
In isolated local services, replayed 345 distinct CoinDesk/Kraken Blog records through normal APIs, preserving historical visibility. Started from a normal Paper SMA strategy backtest, froze 631 closed one-hour BTC perpetual bars, and used owner-scoped orchestration grants and Dashboard credential redemption with the real
deepseek/deepseek-flashAPI. Automatic collectors, extraction workers and scheduler were disabled.CAMPAIGN_NOT_ADOPTABLE; old-phase credential renewal after termination was rejected. That terminal-state rejection does not independently prove active stale-worker fencing.The three BTC facts in automatic event types split into discovery 2, selection 0, sealed holdout 1. No candidate qualified; the loop ended with
INSUFFICIENT_FDR_EVIDENCE. No champion was locked, no Forward sandbox was created, and holdout was not consumed. Thresholds and event timestamps were not relaxed.6 of 10 proposer batches fell back to deterministic scaffolds; only four batches adopted model output. Current records do not persist specific output rejection reasons, so their root cause remains unverified. The 40 slots must not be described as 40 successfully model-authored hypotheses.
Remaining acceptance
runner_eligible=false.This is local development service-level evidence, not production deployment verification. The isolated audit database is retained locally; owned test services were stopped and temporary credential material was removed.
Review
Independent extraction testing, maintainability, security, performance and adversarial reviews completed earlier. The hyphenated fork/URL attribution issue was fixed with regression tests. An independent read-only review of the new acceptance report found no discrepancies or exposed credentials.
Proposer reliability follow-up
Real reproduction identified four JSON-array parsing failures and two confirmation-field DSL failures. Persist bounded sanitized rejection categories atomically with immutable proposal receipts (migration 0059); never save provider content, exception messages or credentials. Supply the actual nested DSL schema and model-owned event-type enum. Ownership, call topology, fee reservation and lease fencing are unchanged.
A subsequent real five-generation run adopted 8 of 10 model batches, with zero JSON/DSL rejections. Two request-stage failures retain unknown usage: $0.0577704 settled plus $0.0340608 reserved, not zero-cost failures. Their exact causes were not captured by the earlier generic category. Future diagnostics distinguish timeout/status/connection and local input limits. The final enum addition has automated coverage and was not followed by another paid rerun solely for that addition.
Selection evidence is still zero; both follow-up runs correctly ended with insufficient evidence. No positive champion, Forward, holdout or adoption claim is made. Detailed real receipts are recorded in
docs/validation/e2-proposer-acceptance-2026-09-30.json.Add a read-only BTC real-input preflight with natural 60/20/20 bar windows, source provenance, freshness, gaps and the evaluator's 24-hour independence rule. Input readiness is a necessary condition only; matched controls, FDR and profitability still decide qualification. The preflight makes no model calls. Forward must accumulate real time; final adoption remains the user's decision.
Final validation: 192 passed / 49 skipped across the Evolver offline suite and two readiness regression tests; 40 passed in related checks including a migrated real PostgreSQL database. These are automated checks, separate from real provider execution. Consistency remains 19 passed, zero failures, 12 existing warnings. Independent review found and fixed event-type enum omission and article-versus-independent-event counting; follow-up review found no further issues.
Real input refresh and extraction recovery
Replay the latest genuine archive through normal Data service endpoints into the existing isolated audit database. Add a reusable loopback-only replay CLI, source read-only SQL, destination routing sentinel, immutable provenance checks and safe resumable receipts. The existing Research durable worker performs extraction; the CLI does not insert facts or call a model.
Measured refresh: 348 independent original records, 349 raw versions and 349 fact versions. Three new originals and one content revision were added relative to the previous 345-record sample. All 349 extraction jobs completed. Source identity, content hashes, first-seen and publication timestamps were retained; normal Data revision semantics conservatively advanced one revision's local fetched/accepted timestamps. This is explicitly not an exact historical copy for that revision, and no timestamp was backdated.
Repeated replay created zero versions and preserved the source metadata hash. Restarting the owned Research service preserved all completed jobs and raw/fact counts. Creating a snapshot twice at the same cutoff reused its ID and events hash. The official FrozenBarsLoader froze 632 closed one-hour and 158 closed four-hour bars via normal Binance/Data APIs. Both natural window splits still contain 2 discovery / 0 selection / 1 sealed-holdout events. The formal event-study evaluator found zero selection events and zero matched controls for all eight initial scaffold arms; no sealed returns were inspected.
There were zero new model calls or fees; persistent model reservation counts were unchanged. Existing unknown-usage reservations remain retained. No new search, champion, Forward, holdout consumption or adoption occurred. The input stage is verified; the full positive E2 chain remains pending real evidence.
Updated the Research README, central current-state document and sanitized
docs/validation/e2-input-refresh-2026-09-30.json. Owned input services were stopped and temporary credential material removed; audit data is retained. Thirteen new replay regressions plus two readiness regressions passed (15 total); Ruff, diff checks and consistency (19 passed / 0 failed / 12 existing warnings) passed. Independent read-only replay review found no additional issues. Code and documentation are separate logical commits.Merge-readiness fixes
Integrate current main while preserving Event Fact v2 and bounded research CLI documentation. Sequence diagnostics as migration
0059after wallet revision0058; CI exercises all three migration regressions and the input tools. Remove private chat follow-up instructions from public documentation.Fix archive replay routing: reject database target overrides, pin effective libpq host/address/port/database, reject loopback aliases of the same service and require a destination sentinel absent from the source database before any API writes. Independent read-only review verified the fixes.
Current automated validation: 241 Evolver/readiness tests passed with zero skips against a disposable PostgreSQL database; 3 migration tests passed, preserving wallet balances and historical receipts and refusing destructive diagnostic downgrade; 23 input-tool tests passed; 166 Research tests passed, 4 integration tests deselected. These checks use fake providers and make no paid model calls.
Final review disposition
The latest review's lowercase-ticker recall observation matches the documented precision policy: textual tickers must be uppercase, while explicit asset names and structured symbols remain supported. Do not restore global uppercasing, which incorrectly identifies ordinary
link/dotwords as assets. The readiness script's duplicated rules are a maintenance concern: the current split, independence and filters were independently verified against the evaluator, and its output explicitly reports necessary input coverage rather than formal qualification. Centralizing these rules remains a follow-up. The missing-JWT concern is disproved by the locked Evolver → Paper → PyJWT dependency chain and successful collection/execution.Final SHA
134498bc92afe53f7dd15ac885c46d68e5573654: all 15 checks passed, including 258 Evolver tests, 23 input-tool tests, 3 migration regressions, dashboard production build, container build and health smoke. One previous dashboard Turbopack/font build failed transiently; the latest-SHA build passed. Independent reviews found no remaining correctness/security blockers.