Skip to content

fix(evolution): persist safe proposal diagnostics and validate input readiness - #183

Merged
mirror29 merged 17 commits into
mainfrom
codex/event-evidence-chain
Sep 30, 2026
Merged

mirror29 merged 17 commits into
mainfrom
codex/event-evidence-chain

Conversation

@mirror29

@mirror29 mirror29 commented Sep 30, 2026 •

Copy link
Copy Markdown
Owner

Summary

Real crypto news names Bitcoin, Ethereum or Solana without tickers. Recognize explicit full names, require uppercase textual tickers, preserve structured symbols, and exclude Bitcoin Cash/SV/Gold and Ethereum Classic, including hyphenated names and URL slugs. New facts use deterministic-event-extractor-v2; existing facts and snapshots remain immutable.

Record real input validation and a paid-provider five-generation execution with restart recovery. docs/04-current-state.md owns the evidence and remaining acceptance work; the sanitized structured report is docs/validation/e2-real-acceptance-2026-09-30.json. Event taxonomy, statistical thresholds, public APIs and credential grant flow are unchanged.

Validation

  • Research offline suite: 163 passed, 4 integration tests deselected; Ruff and diff checks passed.
  • Earlier input validation extracted 100 archive versions through normal Research/Data APIs, replayed 94 distinct records preserving hashes and acceptance timestamps, verified retry/snapshot reuse, and froze 1,633 closed four-hour bars. The disposable database was removed afterward.
  • Earlier durable-loop automated validation: 30 passed using test(evolver): validate durable loops in isolated databases #181, with fake LLMs. This is separate from the real run below.
  • Current consistency: 19 passed, zero failures, 12 existing warnings, matching current main. The historical 120-warning baseline belongs to an earlier checkout.

Real E2 acceptance result

Real execution, checkpoint restart recovery and insufficient-evidence rejection passed. The complete positive E2 chain has not passed acceptance.

In isolated local services, replayed 345 distinct CoinDesk/Kraken Blog records through normal APIs, preserving historical visibility. Started from a normal Paper SMA strategy backtest, froze 631 closed one-hour BTC perpetual bars, and used owner-scoped orchestration grants and Dashboard credential redemption with the real deepseek/deepseek-flash API. Automatic collectors, extraction workers and scheduler were disabled.

  • Baseline completed and handed off to five generations: 40 hypothesis slots, 120 successful implementation evaluations, zero evaluation failures.
  • Duplicate trigger reused the same loop and campaign.
  • Restarted Evolver after the generation-four proposal checkpoint, with no provider request in flight. Waited for natural lease expiration; recovered the same loop/campaign/dataset, preserved checkpoints and costs, and completed generation five without repeating completed proposal calls.
  • 11 settled real model calls. Ledger cost $0.0907749, within the explicit $0.50 cap, with zero reserved balance. This uses provider usage tokens and frozen estimated unit prices; the provider invoice was not independently verified.
  • Owner read succeeded; unauthenticated and other-owner access were rejected. Premature adoption returned CAMPAIGN_NOT_ADOPTABLE; old-phase credential renewal after termination was rejected. That terminal-state rejection does not independently prove active stale-worker fencing.

The three BTC facts in automatic event types split into discovery 2, selection 0, sealed holdout 1. No candidate qualified; the loop ended with INSUFFICIENT_FDR_EVIDENCE. No champion was locked, no Forward sandbox was created, and holdout was not consumed. Thresholds and event timestamps were not relaxed.

6 of 10 proposer batches fell back to deterministic scaffolds; only four batches adopted model output. Current records do not persist specific output rejection reasons, so their root cause remains unverified. The 40 slots must not be described as 40 successfully model-authored hypotheses.

Remaining acceptance

  1. Continue checking provider request failures and retained unknown-usage reservations; the DSL diagnostic/contract repair is now implemented and exercised with real calls.
  2. Obtain legitimate discovery/selection event coverage before rerunning the positive champion/Forward handoff.
  3. Accumulate real-time Forward evidence, verify one-time sealed holdout and explicit experimental adoption with runner_eligible=false.
  4. Exercise in-flight provider crash semantics, active stale-worker fencing, and the remaining stage-handoff recovery cases.

This is local development service-level evidence, not production deployment verification. The isolated audit database is retained locally; owned test services were stopped and temporary credential material was removed.

Review

Independent extraction testing, maintainability, security, performance and adversarial reviews completed earlier. The hyphenated fork/URL attribution issue was fixed with regression tests. An independent read-only review of the new acceptance report found no discrepancies or exposed credentials.

Proposer reliability follow-up

Real reproduction identified four JSON-array parsing failures and two confirmation-field DSL failures. Persist bounded sanitized rejection categories atomically with immutable proposal receipts (migration 0059); never save provider content, exception messages or credentials. Supply the actual nested DSL schema and model-owned event-type enum. Ownership, call topology, fee reservation and lease fencing are unchanged.

A subsequent real five-generation run adopted 8 of 10 model batches, with zero JSON/DSL rejections. Two request-stage failures retain unknown usage: $0.0577704 settled plus $0.0340608 reserved, not zero-cost failures. Their exact causes were not captured by the earlier generic category. Future diagnostics distinguish timeout/status/connection and local input limits. The final enum addition has automated coverage and was not followed by another paid rerun solely for that addition.

Selection evidence is still zero; both follow-up runs correctly ended with insufficient evidence. No positive champion, Forward, holdout or adoption claim is made. Detailed real receipts are recorded in docs/validation/e2-proposer-acceptance-2026-09-30.json.

Add a read-only BTC real-input preflight with natural 60/20/20 bar windows, source provenance, freshness, gaps and the evaluator's 24-hour independence rule. Input readiness is a necessary condition only; matched controls, FDR and profitability still decide qualification. The preflight makes no model calls. Forward must accumulate real time; final adoption remains the user's decision.

Final validation: 192 passed / 49 skipped across the Evolver offline suite and two readiness regression tests; 40 passed in related checks including a migrated real PostgreSQL database. These are automated checks, separate from real provider execution. Consistency remains 19 passed, zero failures, 12 existing warnings. Independent review found and fixed event-type enum omission and article-versus-independent-event counting; follow-up review found no further issues.

Real input refresh and extraction recovery

Replay the latest genuine archive through normal Data service endpoints into the existing isolated audit database. Add a reusable loopback-only replay CLI, source read-only SQL, destination routing sentinel, immutable provenance checks and safe resumable receipts. The existing Research durable worker performs extraction; the CLI does not insert facts or call a model.

Measured refresh: 348 independent original records, 349 raw versions and 349 fact versions. Three new originals and one content revision were added relative to the previous 345-record sample. All 349 extraction jobs completed. Source identity, content hashes, first-seen and publication timestamps were retained; normal Data revision semantics conservatively advanced one revision's local fetched/accepted timestamps. This is explicitly not an exact historical copy for that revision, and no timestamp was backdated.

Repeated replay created zero versions and preserved the source metadata hash. Restarting the owned Research service preserved all completed jobs and raw/fact counts. Creating a snapshot twice at the same cutoff reused its ID and events hash. The official FrozenBarsLoader froze 632 closed one-hour and 158 closed four-hour bars via normal Binance/Data APIs. Both natural window splits still contain 2 discovery / 0 selection / 1 sealed-holdout events. The formal event-study evaluator found zero selection events and zero matched controls for all eight initial scaffold arms; no sealed returns were inspected.

There were zero new model calls or fees; persistent model reservation counts were unchanged. Existing unknown-usage reservations remain retained. No new search, champion, Forward, holdout consumption or adoption occurred. The input stage is verified; the full positive E2 chain remains pending real evidence.

Updated the Research README, central current-state document and sanitized docs/validation/e2-input-refresh-2026-09-30.json. Owned input services were stopped and temporary credential material removed; audit data is retained. Thirteen new replay regressions plus two readiness regressions passed (15 total); Ruff, diff checks and consistency (19 passed / 0 failed / 12 existing warnings) passed. Independent read-only replay review found no additional issues. Code and documentation are separate logical commits.

Merge-readiness fixes

Integrate current main while preserving Event Fact v2 and bounded research CLI documentation. Sequence diagnostics as migration 0059 after wallet revision 0058; CI exercises all three migration regressions and the input tools. Remove private chat follow-up instructions from public documentation.

Fix archive replay routing: reject database target overrides, pin effective libpq host/address/port/database, reject loopback aliases of the same service and require a destination sentinel absent from the source database before any API writes. Independent read-only review verified the fixes.

Current automated validation: 241 Evolver/readiness tests passed with zero skips against a disposable PostgreSQL database; 3 migration tests passed, preserving wallet balances and historical receipts and refusing destructive diagnostic downgrade; 23 input-tool tests passed; 166 Research tests passed, 4 integration tests deselected. These checks use fake providers and make no paid model calls.

Final review disposition

The latest review's lowercase-ticker recall observation matches the documented precision policy: textual tickers must be uppercase, while explicit asset names and structured symbols remain supported. Do not restore global uppercasing, which incorrectly identifies ordinary link/dot words as assets. The readiness script's duplicated rules are a maintenance concern: the current split, independence and filters were independently verified against the evaluator, and its output explicitly reports necessary input coverage rather than formal qualification. Centralizing these rules remains a follow-up. The missing-JWT concern is disproved by the locked Evolver → Paper → PyJWT dependency chain and successful collection/execution.

Final SHA 134498bc92afe53f7dd15ac885c46d68e5573654: all 15 checks passed, including 258 Evolver tests, 23 input-tool tests, 3 migration regressions, dashboard production build, container build and health smoke. One previous dashboard Turbopack/font build failed transiently; the latest-SHA build passed. Independent reviews found no remaining correctness/security blockers.

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review base: 2c15302

🤖 DeepSeek V4 Pro PR Review

Reviewed the diff end-to-end (CI, docs, migration 0059, evolver proposer + checkpoint persistence, research extractor, and the two new root scripts + tests). The change is careful and internally consistent (diagnostics are whitelisted and leak-tested, the DB column is bounded, the replay tool is pinned to loopback/audited targets). A few things are worth resolving before merge.

Findings

[medium] services/research/src/inalpha_research/event_extractor.py:99 — dropping .upper() in re.findall(r"�(?:BTC|ETH|SOL|XRP|ADA|DOGE|BNB|AVAX|DOT|LINK)�", evidence) makes all ticker matching case-sensitive, not just the ambiguous link/dot cases you wanted to fix; tickers without a full-name fallback (XRP, AVAX) lose recall entirely for lowercase/mixed-case mentions. Concrete: an article body "xrp and avax lead the bounce" previously yielded ["XRP","AVAX"] and now yields [], while "link"/"dot" correctly stop matching. The full-name map added in this PR covers only BTC/ETH/SOL/ADA/DOGE/DOT/LINK/BNB, so XRP/AVAX have no fallback. — evidence: general principle (correctness / regression); CLAUDE.md §3.1 expects conservative-but-not-lossy extraction. Consider case-insensitive matching for the unambiguous tickers and case-sensitive only for LINK/DOT, or add XRP/AVAX names.

[medium] scripts/check-e2-readiness.py:23 (independent_counts) and the selection-window SQL — the readiness gate re-implements the evaluator's independence window, 60/20/20 split, and the type/source/severity/confidence filter set instead of reusing evaluator/event_study.py (the docstring even says “Mirror evaluator/event_study.py's 24-hour BTC/type independence window”). These are now two independent sources of truth for the exact gate the loop enforces. Concrete: if the evaluator's independence rule or qualifying-type set changes, the preflight can report ready_to_attempt_search: true while the real campaign still returns INSUFFICIENT_FDR_EVIDENCE (or vice-versa), i.e. operators get a green/red signal that disagrees with the actual search gate. — evidence: general principle (design: duplicated logic / single source of truth).

[medium] .github/workflows/ci.yml:247 — the new step runs repo-root scripts/tests/test_e2_*.py inside services/evolver's locked environment, but those tests runpy the scripts, which import httpx, PyJWT (jwt) and python-dotenv at module load. The Research README itself says the replay tool must run in “an environment with psycopg/httpx/PyJWT/python-dotenv installed”, which signals these are not guaranteed in every service env; if PyJWT (not an obvious evolver dependency) is absent from evolver's lock, the evolver test job errors at collection and the CI red line (§4) fails for an unrelated service. Concrete: uv run pytest ../../scripts/tests/test_e2_archive_replay.py → ModuleNotFoundError: No module named 'jwt'. — evidence: CLAUDE.md §4 (CI red line; missing-dependency CI breaks). Prefer running these under data (owns ingest/httpx/psycopg) or a dedicated job, or add the deps explicitly.

Notes (below threshold, no action required)

  • The invalid_json_array kind: "syntax" branch looks unreachable given _parse_json_array appears to surface non-JSON as ValueError; harmless.
  • Migration 0059's CHECK(jsonb_array_length(diagnostics)<=2) is hard-coupled to exactly two proposer batches; currently safe because commit_proposal requires exactly 8 hypotheses (2 batches), but a future change to batch count would turn the insert into a hard CheckViolation that fails the whole generation checkpoint.
  • docs/04-current-state.md and the validation JSONs are internally consistent and reference only public paths (no docs/miro/ leakage).

@cloudflare-workers-and-pages

cloudflare-workers-and-pages Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Deploying inalpha-web with  Cloudflare Pages  Cloudflare Pages

Latest commit: 134498b
Status: ✅  Deploy successful!
Preview URL: https://87d898d2.inalpha-web.pages.dev
Branch Preview URL: https://codex-event-evidence-chain.inalpha-web.pages.dev

View logs

@mirror29 mirror29 changed the title fix(research): recognize crypto names and verify real event inputs fix(research): recognize crypto names and validate real E2 execution Sep 30, 2026
@mirror29 mirror29 changed the title fix(research): recognize crypto names and validate real E2 execution fix(evolution): improve real proposal reliability and acceptance evidence Sep 30, 2026
@mirror29 mirror29 changed the title fix(evolution): improve real proposal reliability and acceptance evidence fix(evolution): persist safe proposal diagnostics and validate input readiness Sep 30, 2026
@mirror29
mirror29 merged commit 1d8494d into main Sep 30, 2026
19 of 28 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant