diff --git a/.github/workflows/data-origin-contract.yml b/.github/workflows/data-origin-contract.yml index 2a38759..b7fdc02 100644 --- a/.github/workflows/data-origin-contract.yml +++ b/.github/workflows/data-origin-contract.yml @@ -10,7 +10,7 @@ on: paths: - 'workers/data-isamples-org/**' - 'tests/test_data_origin_contract.py' - - 'isamples_202608_release_manifest.json' + - 'isamples_202609_release_manifest.json' - '.github/workflows/data-origin-contract.yml' schedule: - cron: '17 6 * * *' # daily 06:17 UTC diff --git a/CANONICAL.md b/CANONICAL.md index 60e4eba..cb4b337 100644 --- a/CANONICAL.md +++ b/CANONICAL.md @@ -6,33 +6,42 @@ accumulation is superseded versions and orphan snapshots around it. This file names the canonical set and documents every non-canonical survivor so the accident is neutralized even where files remain on the bucket.* +*Updated 2026-09 for the 202608 → 202609 generation flip: §1 now names the 202609 +set; §2's rows are unchanged history (all of them are 202608 files); §3 records +202608 as superseded.* + **Companion:** the machine-readable release manifest -(`tools/build_release_manifest.py` → `isamples_202608_release_manifest.json`) +(`tools/build_release_manifest.py` → `isamples_202609_release_manifest.json`; the +superseded `isamples_202608_release_manifest.json` stays deployed for old links) enumerates the same canonical set with sizes/checksums; the Explorer cross-checks it at boot (#334 v0). A file absent from the manifest is non-canonical **by definition**. -## 1. The canonical (live) set — 202608 +## 1. The canonical (live) set — 202609 Exactly what production `explorer.qmd` loads, all under `https://data.isamples.org/`: | File | Role (plain English) | |---|---| -| `isamples_202608_wide.parquet` | Full sample detail, one row per entity — everything derives from this | -| `isamples_202608_samples_map_lite_v3.parquet` | Slim map/table columns (coords, label, place, date, h3 cells) | -| `isamples_202608_sample_facets_v4.parquet` | Per-sample facet URIs + search-text blob (the `?fts=off` scan target) | -| `isamples_202608_h3_summary_res{4,6,8}.parquet` | Pre-counted globe clusters at 3 zoom tiers | -| `isamples_202608_facet_summaries.parquet` | Facet checkbox counts, no filters active | -| `isamples_202608_facet_cross_filter.parquet` | Facet counts under single cross-filter | -| `isamples_202608_facet_tree_summaries.parquet` | Hierarchical (tree) facet counts | -| `isamples_202608_facet_tree_cross_filter.parquet` | Tree facet counts under cross-filter | -| `isamples_202608_sample_facet_membership.parquet` | Sample↔facet-tree membership | -| `isamples_202608_sample_facet_masks.parquet` | Bitmask substrate for multi-filter counts | -| `isamples_202608_facet_node_bits.parquet` | Facet-node bit assignments for the masks | -| `isamples_202608_sample_facet_index.parquet` | Per-sample facet index (multi-filter fast path) | -| `isamples_202608_sample_facet_index_meta.parquet` | ~1 KB trusted manifest for the above (#313/#317 boot-race fix) | -| `vocab_labels_202608.parquet` | URI → human label (539 entries; verified complete 2026-07-22) | -| `isamples_202608_search_index_v1/` | Sharded FTS index — the default search since 2026-07-17. Runtime loads: `build_stats.json`, `hot_tokens.json`, `shard_sizes.json`, `hot_topk.parquet`, + 256 base token shards and ~591 hot-token sub-files (852 objects total incl. sidecars; base inventory in `shard_sizes.json`, totals in `build_stats.json`); `df.parquet` is offline-only | +| `isamples_202609_wide.parquet` | Full sample detail, one row per entity — everything derives from this | +| `isamples_202609_samples_map_lite_v3.parquet` | Slim map/table columns (coords, label, place, date, h3 cells) | +| `isamples_202609_sample_facets_v4.parquet` | Per-sample facet URIs + search-text blob (the `?fts=off` scan target) | +| `isamples_202609_h3_summary_res{4,6,8}.parquet` | Pre-counted globe clusters at 3 zoom tiers | +| `isamples_202609_facet_summaries.parquet` | Facet checkbox counts, no filters active | +| `isamples_202609_facet_cross_filter.parquet` | Facet counts under single cross-filter | +| `isamples_202609_facet_tree_summaries.parquet` | Hierarchical (tree) facet counts | +| `isamples_202609_facet_tree_cross_filter.parquet` | Tree facet counts under cross-filter | +| `isamples_202609_sample_facet_membership.parquet` | Sample↔facet-tree membership | +| `isamples_202609_sample_facet_masks.parquet` | Bitmask substrate for multi-filter counts | +| `isamples_202609_facet_node_bits.parquet` | Facet-node bit assignments for the masks | +| `isamples_202609_sample_facet_index.parquet` | Per-sample facet index (multi-filter fast path) | +| `isamples_202609_sample_facet_index_meta.parquet` | ~1 KB trusted manifest for the above (#313/#317 boot-race fix) | +| `vocab_labels_202609.parquet` | URI → human label (539 entries, same count as 202608) | +| `isamples_202609_search_index_v1/` | Sharded FTS index — the default search since 2026-07-17. Runtime loads: `build_stats.json`, `hot_tokens.json`, `shard_sizes.json`, `hot_topk.parquet`, + 256 base token shards and 769 hot-token sub-files (1,030 objects total incl. sidecars; base inventory in `shard_sizes.json`, totals in `build_stats.json`); `df.parquet` is offline-only | + +The `_v3`/`_v4` suffixes carry over from 202608 so the Explorer's file names stay +recognizable; 202609 was built with those fixes natively and has **no** earlier +`_v2`/`_v3` siblings on the bucket. The ~9-file facet family looks baroque but is load-bearing: it is the price of fast multi-filter counts with no server. See `EXPLORER_QUERIES.md` for how each @@ -64,7 +73,8 @@ boot; this table makes them legible to humans.) | `isamples_202601_*` | **Archival** — the coherent snapshot in Zenodo draft 21288719 | | `isamples_202604_*` | **NOT orphan — backs the stable `/current/wide.parquet` alias** used by data.qmd and multiple tutorials (see SERIALIZATIONS.md). Must NOT be moved until the alias is repointed and consumers audited | | `isamples_202606_*` | Pipeline intermediate (the #272 OC-concept-enriched wide that fed 202608 — see DATA_PROVENANCE.md step 3b). Not loaded by anything at runtime; safe to attic post-grant, lineage documented | -| `isamples_202608_*` | **Live** (production Explorer) | +| `isamples_202608_*` | **Superseded by 202609** (live from the #284 cutover, merged 2026-06-14, until the 202609 flip). Still served, bytes unchanged, so old links and notebooks keep working. Not loaded by the Explorer. Its §2 suffix history stays valid | +| `isamples_202609_*` | **Live** (production Explorer). A deterministic rebuild with the same sample identities and counts as 202608 — but **not identical content**: (a) 144 fewer OpenContext location rows, because site de-duplication now picks its winner deterministically (pqg#26); (b) the flat single-value `material` differs for 41,372 SESAR samples, because the flat value is the first non-root entry of `p__has_material_category` (see DATA_PROVENANCE.md, "Material selection") and pqg#26 made that array input-ordered instead of arbitrary. So flat `facet_summaries` material counts moved (e.g. soil 32,149 → 13,806) and ~110 `facet_cross_filter` rows changed; (c) `place_name` differs for 11,114 located samples in capitalisation only (e.g. `UTAH` → `Utah`; all 11,114 are equal case-insensitively). Unchanged: tree facet summaries and tree cross-filter (row for row); sample coordinates (0 of 6,026,242 moved); H3 cells with the same per-cell sample count, dominant source and source count (centres within 0.000001°, rounding); located-sample and per-source counts | ## 4. Conveniences (not load-bearing; fine to keep, labeled) @@ -72,6 +82,8 @@ boot; this table makes them legible to humans.) |---|---| | `isamples_202601_wide_h3.parquet` | Analyst convenience (wide + precomputed H3 cells). In the Zenodo deposit; NOT loaded by the Explorer | | `*.csv` twins of lite/H3 (~640 MB) | Convenience exports; parquet is authoritative; excluded from Zenodo by design | +| `staging/isamples_202609/*` | Pre-cutover rehearsal copy of the §1 set (hash-verified against the build). NOT loaded by the Explorer; safe to delete once the real keys are verified | +| `staging/isamples-0909-smoketest/*` | Upload-tool smoke test (4 small files). Not referenced anywhere; safe to delete | ## 5. Post-grant simplification options (deliberately NOT done in-grant) @@ -81,5 +93,6 @@ boot; this table makes them legible to humans.) 9 files trade storage (cheap) for browser CPU (scarce); collapsing them is an engineering project, not cleanup. 3. If a fresher export ever lands (#320), cut ONE coherent generation and retire - 202608 wholesale via the manifest. + the then-live generation wholesale via the manifest. (202609 is not that + fresher cut — it rebuilt the same sample records reproducibly.) diff --git a/DATA_PROVENANCE.md b/DATA_PROVENANCE.md index 0613833..054402e 100644 --- a/DATA_PROVENANCE.md +++ b/DATA_PROVENANCE.md @@ -4,19 +4,20 @@ How the explorer's derived parquet files are generated, from root to publish. **Not exhaustive as of 2026-08-05 — see the coverage caveat below.** *Reviewed 2026-06-02 (CC, via codebase audit). Complements `SERIALIZATIONS.md` (format/schema reference); this file is the end-to-end build chain + the automation gaps.* -> ⚠️ **Coverage caveat (2026-08-05).** The DAG below documents the **seven-file +> ⚠️ **Coverage caveat (2026-08-05; live generation updated 2026-09).** The DAG below documents the **seven-file > derived substrate as of the 2026-06-02 review**. It does *not* cover the whole -> live `202608` family. Known omissions: +> live `202609` family. Known omissions: > > - `sample_facet_masks`, `facet_node_bits`, `sample_facet_index`, > `sample_facet_index_meta` (the bitmask count path, #299/#304/#305/#313) > - `sample_facet_membership` > - `facet_tree_summaries`, `facet_tree_cross_filter` (the tree facet path, #290) -> - the sharded search index `isamples_202608_search_index_v1/` (#171) +> - the sharded search index `isamples_202609_search_index_v1/` (#171) > -> The *build chain and the automation gaps* it describes are still accurate for -> the files it does cover; treat it as incomplete rather than wrong. Authoritative -> current inventory: `isamples_202608_release_manifest.json` / `CANONICAL.md`. +> The *build chain and the automation gaps* it describes were accurate, as of that +> review, for the files it covers. They have not been re-audited for the 202609 +> re-run, so treat this file as incomplete and dated rather than wrong. Authoritative +> current inventory: `isamples_202609_release_manifest.json` / `CANONICAL.md`. > **Load-bearing constraint:** the **root export cannot be regenerated.** It was produced from the iSamples Central Solr API (`central.isample.xyz`), **offline since Aug 2025**. The Zenodo-archived export is a **frozen root**. Any *new* data (e.g. concept URIs, thumbnails) therefore must come from a **per-source supplementary file merged into the base by `pid`** — the "sidecar" pattern (see Stage 3) — not from re-exporting. @@ -87,11 +88,15 @@ Eric Kansa maintains OpenContext PQG **independently** on GCS (`storage.googleap > ⚠️ **Snapshot note (2026-08-05).** The version-skew bullets below were written > when the deployed derived files were `202601` and the wide was `202604`. That is -> no longer the live state: the Explorer now serves the **`202608`** family -> (`sample_facets_v4`, `samples_map_lite_v3`, `wide` at 300,303,095 B). The -> *reproducibility* gap the bullets describe is still real and still unresolved — -> only the version numbers in them are historical. Authoritative current inventory: -> [`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json), +> no longer the live state: the Explorer now serves the **`202609`** family +> (`sample_facets_v4`, `samples_map_lite_v3`, `wide` at 236,995,193 B). The +> reproducibility and version-skew bullets below describe **those historical builds** +> (the `202601` derived files and the `202604` wide) and are kept as a record. The +> 202609 generation came from a full re-run of the chain with per-file hashes recorded +> (`tools/verify_release.py`; toolchain pin in #358). These bullets have not been +> re-audited against that re-run, so don't read them as statements about 202609. +> Authoritative current inventory: +> [`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json), > human twin `CANONICAL.md`. - ⚠️ **The deployed `202601` derived files are NOT reproducible** from any available wide. A rebuild yields **528,983** root-material rows (pre-#271); the deployed `sample_facets_v2` has **346,768** — so the live files came from a different/unrecorded Stage-4 process, *and* the data has since rolled (wide is now `202604`). Treat a fresh `build_frontend_derived.py` run as the new source of truth, not as a bit-for-bit reproduction of the deployed files. diff --git a/EXPLORER_QUERIES.md b/EXPLORER_QUERIES.md index 93e7c2a..aab57d3 100644 --- a/EXPLORER_QUERIES.md +++ b/EXPLORER_QUERIES.md @@ -23,22 +23,22 @@ server. You can open any of these URLs directly, or point DuckDB at them one place, `explorer.qmd` around **line 800-864**, e.g.: ```js -lite_url = `${R2_BASE}/isamples_202608_samples_map_lite_v3.parquet` // map points + table -wide_url = `${R2_BASE}/isamples_202608_wide.parquet` // full sample detail -facets_url = `${R2_BASE}/isamples_202608_sample_facets_v4.parquet` // material/context/object_type + search text -h3_res4_url = `${R2_BASE}/isamples_202608_h3_summary_res4.parquet` // pre-counted globe dots (world zoom) +lite_url = `${R2_BASE}/isamples_202609_samples_map_lite_v3.parquet` // map points + table +wide_url = `${R2_BASE}/isamples_202609_wide.parquet` // full sample detail +facets_url = `${R2_BASE}/isamples_202609_sample_facets_v4.parquet` // material/context/object_type + search text +h3_res4_url = `${R2_BASE}/isamples_202609_h3_summary_res4.parquet` // pre-counted globe dots (world zoom) ``` | File | Plain-English role | Roughly how big | |---|---|---| -| `..._wide.parquet` | Full detail for every sample (one row each) — everything else is derived from this | ~300 MB | +| `..._wide.parquet` | Full detail for every sample (one row each) — everything else is derived from this | ~237 MB | | `..._samples_map_lite_v3.parquet` | Slim version with just what the map/table need: coords, label, place, date | ~63 MB | | `..._sample_facets_v4.parquet` | One row per sample: material/context(sampled feature)/object_type as plain URIs, plus a search-text blob | ~69 MB | | `..._h3_summary_res{4,6,8}.parquet` | Pre-counted dots for the globe at 3 zoom tiers (continent / region / neighborhood), so zooming out never counts 6M rows live | tiny–few MB | | `..._facet_summaries.parquet`, `..._facet_cross_filter.parquet`, `..._facet_tree_*.parquet` | Pre-computed facet-checkbox counts at various levels of "how many filters are active" — the whole point of these is to avoid a live COUNT over millions of rows | KB–tens of MB | | `..._sample_facet_masks.parquet`, `..._sample_facet_index.parquet` | Bitmask tricks so 2+ facet filters at once are still fast (see `SERIALIZATIONS.md` §4.12 if you want the gory detail) | ~10 MB each | | `vocab_labels_*.parquet` | URI → human-readable label lookup (e.g. `.../material/1.0/rock` → "Rock") | ~60 KB | -| `..._search_index_v1/` (852 files) | Pre-built search index (like a book's back-of-book index, sharded): token shards + tiny sidecars (`hot_tokens.json`, `df.parquet`, `build_stats.json`). The default search path since 2026-07-17 | few KB–few MB per shard | +| `..._search_index_v1/` (1,030 files) | Pre-built search index (like a book's back-of-book index, sharded): token shards + tiny sidecars (`hot_tokens.json`, `df.parquet`, `build_stats.json`). The default search path since 2026-07-17 | few KB–few MB per shard | *Full list with exact schemas: `SERIALIZATIONS.md`. This table is the subset that matters for "what happens when I click around the Explorer."* @@ -87,7 +87,7 @@ of filters" doesn't pre-aggregate cleanly. **Default path (since 2026-07-17): a pre-built search index.** Your query is split into tokens, and each token maps (by hash) to a small parquet shard of -a pre-built inverted index (`isamples_202608_search_index_v1/`, 852 files — +a pre-built inverted index (`isamples_202609_search_index_v1/`, 1,030 files — think "the index at the back of a book, one file per drawer"). The browser fetches only the few KB-to-MB shards for *your* tokens, intersects the matching sample ids, and ranks them by relevance (BM25 — the standard @@ -151,7 +151,7 @@ LEFT JOIN read_parquet('vocab_labels.parquet') mat_lbl ON mat_lbl.uri = mat.pid WHERE s.pid = '' ``` This is the one query that reads from `wide.parquet` on click (everything -above deliberately avoids touching the 300 MB wide file until you actually +above deliberately avoids touching the 237 MB wide file until you actually need full detail on one sample). ## Try it yourself @@ -162,7 +162,7 @@ You don't need the browser — any of this works from the DuckDB CLI or ```sql -- how many samples per source, right now, live off the public URL SELECT n AS source, COUNT(*) -FROM read_parquet('https://data.isamples.org/isamples_202608_wide.parquet') +FROM read_parquet('https://data.isamples.org/isamples_202609_wide.parquet') WHERE otype = 'MaterialSampleRecord' GROUP BY n ORDER BY 2 DESC; @@ -170,7 +170,7 @@ GROUP BY n ORDER BY 2 DESC; -- the default path since 2026-07-17 probes the sharded search index instead -- (JS, not a single SQL statement — see SEARCH_INDEX_V1.md) SELECT pid, label, source -FROM read_parquet('https://data.isamples.org/isamples_202608_sample_facets_v4.parquet') +FROM read_parquet('https://data.isamples.org/isamples_202609_sample_facets_v4.parquet') WHERE description ILIKE '%pottery%' LIMIT 20; ``` diff --git a/README.md b/README.md index adbda4e..55ba71a 100644 --- a/README.md +++ b/README.md @@ -3,14 +3,14 @@ title: isamples.github.io subtitle: README for the isamples.github.io source --- -> ⚠️ **Snapshot note (2026-08-05).** The **`202601`** file examples below pin a +> ⚠️ **Snapshot note (2026-08-05; live generation updated 2026-09).** The **`202601`** file examples below pin a > stable, versioned snapshot. Those objects still exist and their byte counts are > correct, but they are **not what the Interactive Explorer serves today** — the -> live app reads the **`202608`** family, including `sample_facets_v4`, +> live app reads the **`202609`** family, including `sample_facets_v4`, > `samples_map_lite_v3`, `wide`, and the sharded search index. (Other generations > may also appear here, such as the `current/` alias or the `202512` narrow file.) > Use `202601` for a stable citable snapshot; for what the Explorer actually -> loads, see `isamples_202608_release_manifest.json` (human twin: `CANONICAL.md`). +> loads, see `isamples_202609_release_manifest.json` (human twin: `CANONICAL.md`). # isamplesorg.github.io diff --git a/SERIALIZATIONS.md b/SERIALIZATIONS.md index f2afdc6..4ee4dc2 100644 --- a/SERIALIZATIONS.md +++ b/SERIALIZATIONS.md @@ -7,14 +7,14 @@ toc: true categories: [data, architecture, parquet] --- -> ⚠️ **Snapshot note (2026-08-05).** The **`202601`** file examples below pin a +> ⚠️ **Snapshot note (2026-08-05; live generation updated 2026-09).** The **`202601`** file examples below pin a > stable, versioned snapshot. Those objects still exist and their byte counts are > correct, but they are **not what the Interactive Explorer serves today** — the -> live app reads the **`202608`** family, including `sample_facets_v4`, +> live app reads the **`202609`** family, including `sample_facets_v4`, > `samples_map_lite_v3`, `wide`, and the sharded search index. (Other generations > may also appear here, such as the `current/` alias or the `202512` narrow file.) > Use `202601` for a stable citable snapshot; for what the Explorer actually -> loads, see `isamples_202608_release_manifest.json` (human twin: `CANONICAL.md`). +> loads, see `isamples_202609_release_manifest.json` (human twin: `CANONICAL.md`). ## 1. Purpose and scope diff --git a/_quarto.yml b/_quarto.yml index 9b2e782..bd5e501 100644 --- a/_quarto.yml +++ b/_quarto.yml @@ -2,6 +2,9 @@ project: type: website output-dir: docs resources: + - isamples_202609_release_manifest.json + # Superseded release's manifest stays deployed so already-published links and + # cached older Explorer renders keep resolving (CANONICAL.md §3). - isamples_202608_release_manifest.json - assets/js/source-palette.js - assets/js/sql-builders.js diff --git a/data.qmd b/data.qmd index d238026..66b762c 100644 --- a/data.qmd +++ b/data.qmd @@ -6,18 +6,18 @@ categories: [data, parquet, download] --- ::: {.callout-important} -## Snapshot note (2026-08-05) +## Snapshot note (2026-08-05; live generation updated 2026-09) The **`202601`** file examples on this page pin a stable, versioned snapshot. Those objects still exist and their byte counts are correct, but they are **not -what the Interactive Explorer serves today** — the live app reads the **`202608`** +what the Interactive Explorer serves today** — the live app reads the **`202609`** family, including `sample_facets_v4`, `samples_map_lite_v3`, `wide`, and the sharded search index. (This page may also reference other generations, such as the `current/` alias or the `202512` narrow file.) Use `202601` when you want a stable, citable reference. For "what the Explorer is actually loading right now", the authoritative inventory is -[`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json) +[`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json) (human twin: `CANONICAL.md`), which the Explorer cross-checks at boot. ::: diff --git a/explorer.qmd b/explorer.qmd index 34ad324..1fd08c1 100644 --- a/explorer.qmd +++ b/explorer.qmd @@ -12,9 +12,33 @@ format: include-in-header: text: | - - - + --- ```{=html} @@ -828,15 +852,16 @@ R2_BASE = (() => { // default and absolute overrides (http://localhost:8099/data) pass through. return raw.startsWith('/') ? new URL(raw, location.origin).href : raw; })() -h3_res4_url = `${R2_BASE}/isamples_202608_h3_summary_res4.parquet` -h3_res6_url = `${R2_BASE}/isamples_202608_h3_summary_res6.parquet` -h3_res8_url = `${R2_BASE}/isamples_202608_h3_summary_res8.parquet` +h3_res4_url = `${R2_BASE}/isamples_202609_h3_summary_res4.parquet` +h3_res6_url = `${R2_BASE}/isamples_202609_h3_summary_res6.parquet` +h3_res8_url = `${R2_BASE}/isamples_202609_h3_summary_res8.parquet` // #300: the _v2 lite carries h3_res4/h3_res6 (for in-browser filtered-cluster // aggregation). A new filename rather than overwriting the original keeps the // immutable-cache contract intact (isamples_YYYYMM_*.parquet is served // immutable/1-yr) — every visitor fetches fresh data, no cache purge. The next // generation builds res4/res6 into the canonical name natively (build change), -// so this _v2 suffix is a one-off retrofit for 202608. +// so this _v2 suffix is a one-off retrofit for 202608. (202609 kept the _v3 +// filename as-is rather than renaming at the generation boundary.) // // #311: _v3 rebuilds place_name/result_time via the SamplingEvent/ // SamplingSite graph traversal (the fix in build_frontend_derived.py) — _v2's @@ -848,20 +873,20 @@ h3_res8_url = `${R2_BASE}/isamples_202608_h3_summary_res8.parquet` // immutable-cache reasoning as _v2: new filename, never overwrite. // #351: queries name this URL, but the db cell's query wrapper serves them from // an in-memory copy fetched once by the liteFile cell (see both). -lite_url = `${R2_BASE}/isamples_202608_samples_map_lite_v3.parquet` +lite_url = `${R2_BASE}/isamples_202609_samples_map_lite_v3.parquet` // Explicit versioned wide (#272: OC concept-enriched — popups read material/ // object-type from this file). The stable alias `current/wide.parquet` still // points at the previous wide until the production cutover flips the manifest; // pinning the version here keeps staging and prod each self-consistent. -wide_url = `${R2_BASE}/isamples_202608_wide.parquet` +wide_url = `${R2_BASE}/isamples_202609_wide.parquet` // v2 carries object_type alongside material and context (URI-string columns). -facets_url = `${R2_BASE}/isamples_202608_sample_facets_v4.parquet` -facet_summaries_url = `${R2_BASE}/isamples_202608_facet_summaries.parquet` +facets_url = `${R2_BASE}/isamples_202609_sample_facets_v4.parquet` +facet_summaries_url = `${R2_BASE}/isamples_202609_facet_summaries.parquet` // Pre-aggregated single-filter cache for fast cross-filtered facet counts. -cross_filter_url = `${R2_BASE}/isamples_202608_facet_cross_filter.parquet` +cross_filter_url = `${R2_BASE}/isamples_202609_facet_cross_filter.parquet` // SKOS prefLabels for Material / Sampled Feature / Specimen Type URIs. // ~60 KB lookup; falls back to URI tail if a URI isn't covered. -vocab_labels_url = `${R2_BASE}/vocab_labels_202608.parquet` +vocab_labels_url = `${R2_BASE}/vocab_labels_202609.parquet` // #281/#282/#291 facet hierarchy — SHIPPED, default ON. Material, Sampled Feature // (context) and Specimen Type (object_type) render as expandable trees backed by the @@ -880,21 +905,21 @@ FACET_TREE = (() => { })() // Hierarchical counts (one row per concept node: facet_type, concept_uri, // parent_uri, depth, count) and per-sample membership (pid ↔ every ancestor). -facet_tree_url = `${R2_BASE}/isamples_202608_facet_tree_summaries.parquet` -membership_url = `${R2_BASE}/isamples_202608_sample_facet_membership.parquet` +facet_tree_url = `${R2_BASE}/isamples_202609_facet_tree_summaries.parquet` +membership_url = `${R2_BASE}/isamples_202609_sample_facet_membership.parquet` // #290/#293 single-active-filter cross-filter cube over the trees + source. // Precomputed COUNT(DISTINCT pid) for each target node given ONE active filter, // so global-view cross-filtered tree counts are instant instead of a live // membership near-full-scan. ~1k rows; same schema as facet_cross_filter. -tree_cross_filter_url = `${R2_BASE}/isamples_202608_facet_tree_cross_filter.parquet` +tree_cross_filter_url = `${R2_BASE}/isamples_202609_facet_tree_cross_filter.parquet` // #293 bitmask filter: per-pid tree-membership masks + the concept_uri→bit map. // facetFilterSQL filters broad multi-tree selections with a columnar bitwise // predicate over sample_facet_masks (~10 MB, one row/pid) instead of the // 39M-row membership GROUP BY that stalls DuckDB-WASM. node_bits (~56 rows) // maps each selected node to its bit. Best-effort: if either fails to load, // facetFilterSQL falls back to the membership scan (no regression). -masks_url = `${R2_BASE}/isamples_202608_sample_facet_masks.parquet` -node_bits_url = `${R2_BASE}/isamples_202608_facet_node_bits.parquet` +masks_url = `${R2_BASE}/isamples_202609_sample_facet_masks.parquet` +node_bits_url = `${R2_BASE}/isamples_202609_facet_node_bits.parquet` // #304/#305 complete per-pid facet index: one row per LOCATED sample (incl. the // ~29,917 with no tree membership that sample_facet_masks omits, #306) + source. // The multi-filter global-view count path scans this with the node_bits bitmask @@ -904,7 +929,7 @@ node_bits_url = `${R2_BASE}/isamples_202608_facet_node_bits.parquet` // multi-filter path shows "Loading…"/"unavailable" (#313) rather than a // misleading baseline (the honesty rule — never baseline under active // filters). -index_url = `${R2_BASE}/isamples_202608_sample_facet_index.parquet` +index_url = `${R2_BASE}/isamples_202609_sample_facet_index.parquet` // #313 P1: tiny trusted manifest (source, count, build_id, schema_version, // total_rows) built DIRECTLY from samp_geo at build time — NOT read back from // index_url — and independently validated against the real on-disk index by @@ -912,15 +937,15 @@ index_url = `${R2_BASE}/isamples_202608_sample_facet_index.parquet` // of scanning the 9.68 MB index_url on every page load; index_url itself is // now touched only lazily, when a user's multi-filter count query actually // runs. Always deployed paired with index_url (same build_id). -index_meta_url = `${R2_BASE}/isamples_202608_sample_facet_index_meta.parquet` +index_meta_url = `${R2_BASE}/isamples_202609_sample_facet_index_meta.parquet` // Single source of truth for the search-index location (#334 round-1 review): // the runtime substrate path and the manifest check MUST share this cell, so a // version bump can't diverge them silently. -search_index_base = `${R2_BASE}/isamples_202608_search_index_v1` +search_index_base = `${R2_BASE}/isamples_202609_search_index_v1` // === #334 v0-detect: release-manifest boot cross-check === -// The release manifest (isamples_202608_release_manifest.json, committed in +// The release manifest (isamples_202609_release_manifest.json, committed in // this repo and deployed with the site; generated by // tools/build_release_manifest.py; human twin: CANONICAL.md) enumerates the // canonical data objects for this release. This cell cross-checks every URL @@ -957,7 +982,7 @@ manifestCheck = { }))]; const result = { status: 'unchecked', missing: [], extra: [], release_id: null }; try { - const resp = await fetch(new URL('isamples_202608_release_manifest.json', document.baseURI)); + const resp = await fetch(new URL('isamples_202609_release_manifest.json', document.baseURI)); if (!resp.ok) { console.warn(`#334 manifest check: manifest fetch HTTP ${resp.status} — skipping`); result.status = 'manifest-unavailable'; @@ -6312,7 +6337,7 @@ zoomWatcher = { // ===== #171: substrate search path (default; escape hatch ?fts=off) ===== // Replaces buildSearchFilter's 69 MB facets ILIKE scan with the - // sharded inverted index at data.isamples.org/isamples_202608_search_index_v1/ + // sharded inverted index at data.isamples.org/isamples_202609_search_index_v1/ // (built by tools/build_search_index.py, contract SEARCH_INDEX_V1.md). // Downstream shares one contract: both producers emit the singleton // `search_pids` table (pid, label, source, place_name, relevance_score); @@ -7797,17 +7822,17 @@ Pre-aggregated H3 hexagonal indices achieve near-instant globe rendering, with s | **Zoom in** | H3 res6 | 1.3 MB | 112K clusters (city) | | **Zoom more** | H3 res8 | 2.0 MB | 176K clusters (neighborhood) | | **Zoom deep** | Map lite | 63 MB (range req.) | Up to 5K individual samples | -| **Click sample** | Full dataset | 300 MB (range req.) | Full metadata for 1 sample | +| **Click sample** | Full dataset | 237 MB (range req.) | Full metadata for 1 sample | *Sizes are decimal MB (10⁶ bytes), matching the manifest's byte counts.* **Static files, zero backend.** All queries run in your browser via DuckDB-WASM with HTTP range requests — only the bytes you need are transferred. The manifest directly inventories 21 objects — including five search-index sidecars — and -summarizes the index's 847 shard files (256 base plus 591 hot); the index -directory holds 852 objects in all. Any one view touches only a handful of them. +summarizes the index's 1,025 shard files (256 base plus 769 hot); the index +directory holds 1,030 objects in all. Any one view touches only a handful of them. The authoritative inventory is -[`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json) +[`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json) (human twin: `CANONICAL.md`), which the Explorer cross-checks at boot. + If you edit these, re-derive from the manifest — do not round from memory. + + Generation flip 202608 -> 202609 (2026-09). Re-derived from the 202609 release + manifest: res4 505,630 B; res6 1,322,935 B; res8 2,009,551 B; map_lite_v3 + 62,923,910 B; wide 236,995,193 B. Only the wide row's rounded figure moved + (300 -> 237 MB): same 170 row groups and ZSTD codec as 202608, 144 fewer + location rows; the size drop itself was not investigated further. Cluster + counts unchanged — 38,462 / 112,019 / 176,669 — checked against the parquet + row counts of BOTH generations. Search index: 256 base + 769 hot = 1,025 + shard files, 1,030 objects with the five sidecars (was 847 / 852). The + manifest is deployed with the site (_quarto.yml resources), so its link is + isamples.org, not data.isamples.org (the old link 404'd). --> ::: {.callout-tip collapse="true"} ## Check these numbers against the source @@ -7849,7 +7885,7 @@ budget are properties of the app, not the data.) To re-derive the sizes without taking this page's word for it: ```bash -curl -s https://data.isamples.org/isamples_202608_release_manifest.json \ +curl -s https://isamples.org/isamples_202609_release_manifest.json \ | jq '.files | to_entries[] | select(.key|test("h3_summary|_wide\\.|map_lite")) | {file: .key, bytes: .value.size_bytes}' ``` @@ -7858,7 +7894,7 @@ Cluster counts come from the parquet row counts, not the manifest: ```sql SELECT count(*) FROM read_parquet( - 'https://data.isamples.org/isamples_202608_h3_summary_res4.parquet'); + 'https://data.isamples.org/isamples_202609_h3_summary_res4.parquet'); ``` ::: diff --git a/how-to-use.qmd b/how-to-use.qmd index 5860e30..1b6baab 100644 --- a/how-to-use.qmd +++ b/how-to-use.qmd @@ -5,18 +5,18 @@ number-sections: false --- ::: {.callout-important} -## Snapshot note (2026-08-05) +## Snapshot note (2026-08-05; live generation updated 2026-09) The **`202601`** file examples on this page pin a stable, versioned snapshot. Those objects still exist and their byte counts are correct, but they are **not -what the Interactive Explorer serves today** — the live app reads the **`202608`** +what the Interactive Explorer serves today** — the live app reads the **`202609`** family, including `sample_facets_v4`, `samples_map_lite_v3`, `wide`, and the sharded search index. (This page may also reference other generations, such as the `current/` alias or the `202512` narrow file.) Use `202601` when you want a stable, citable reference. For "what the Explorer is actually loading right now", the authoritative inventory is -[`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json) +[`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json) (human twin: `CANONICAL.md`), which the Explorer cross-checks at boot. ::: diff --git a/isamples_202609_release_manifest.json b/isamples_202609_release_manifest.json new file mode 100644 index 0000000..036178a --- /dev/null +++ b/isamples_202609_release_manifest.json @@ -0,0 +1,132 @@ +{ + "schema_version": 1, + "release_id": "isamples_202609", + "generated_at_utc": "2026-09-10T19:59:55+00:00", + "base": "https://data.isamples.org", + "files": { + "isamples_202609_wide.parquet": { + "size_bytes": 236995193, + "etag": "91f0d23bdc9807cd32710794433712f5-29", + "last_modified": null + }, + "isamples_202609_samples_map_lite_v3.parquet": { + "size_bytes": 62923910, + "etag": "7add8109033752038634f3e15417f5dc-8", + "last_modified": null + }, + "isamples_202609_sample_facets_v4.parquet": { + "size_bytes": 69382360, + "etag": "0c29b5168c24b29a20304ef955e410dc-9", + "last_modified": null + }, + "isamples_202609_h3_summary_res4.parquet": { + "size_bytes": 505630, + "etag": "0b9222b83a6e635bce94e59637efd8c6", + "last_modified": null + }, + "isamples_202609_h3_summary_res6.parquet": { + "size_bytes": 1322935, + "etag": "a3fb750d67eea0eddc48a106d279bb6e", + "last_modified": null + }, + "isamples_202609_h3_summary_res8.parquet": { + "size_bytes": 2009551, + "etag": "e4d30c67ab1f1a7f7c509ffdeba94682", + "last_modified": null + }, + "isamples_202609_facet_summaries.parquet": { + "size_bytes": 1688, + "etag": "7b994ffabf940ee18821f5af662b8211", + "last_modified": null + }, + "isamples_202609_facet_cross_filter.parquet": { + "size_bytes": 5835, + "etag": "cc5c834eb4fe92bf5da55f9f327884eb", + "last_modified": null + }, + "isamples_202609_facet_tree_summaries.parquet": { + "size_bytes": 2773, + "etag": "8631b499654eb19b356f30bd1cbc10f7", + "last_modified": null + }, + "isamples_202609_facet_tree_cross_filter.parquet": { + "size_bytes": 7102, + "etag": "fe9155ec83548e873c29657ae569ae7e", + "last_modified": null + }, + "isamples_202609_sample_facet_membership.parquet": { + "size_bytes": 30908543, + "etag": "24d5021eda3c1b75060e265c18da36bd-4", + "last_modified": null + }, + "isamples_202609_sample_facet_masks.parquet": { + "size_bytes": 10138648, + "etag": "ae02298f33d10b0f09f029b2619d9118-2", + "last_modified": null + }, + "isamples_202609_facet_node_bits.parquet": { + "size_bytes": 1839, + "etag": "687e2ee9e72184c1affb293aec861850", + "last_modified": null + }, + "isamples_202609_sample_facet_index.parquet": { + "size_bytes": 10152968, + "etag": "b0dab48dde932b0cfcf691ff8a0d9dfc-2", + "last_modified": null + }, + "isamples_202609_sample_facet_index_meta.parquet": { + "size_bytes": 1154, + "etag": "2e7686fd6f008666cbde31dc047f344c", + "last_modified": null + }, + "vocab_labels_202609.parquet": { + "size_bytes": 62440, + "etag": "437c546dfe17f282bfb82124c09452f0", + "last_modified": null + }, + "isamples_202609_search_index_v1/build_stats.json": { + "size_bytes": 1982, + "etag": "2617eae0618cec33998e9e40e98e8255", + "last_modified": null + }, + "isamples_202609_search_index_v1/hot_tokens.json": { + "size_bytes": 20359, + "etag": "5932f186375a2baf9f744607dc5d324a", + "last_modified": null + }, + "isamples_202609_search_index_v1/shard_sizes.json": { + "size_bytes": 7934, + "etag": "5e05788f0d4d5593e840173a0ccaf86d", + "last_modified": null + }, + "isamples_202609_search_index_v1/hot_topk.parquet": { + "size_bytes": 762875, + "etag": "21a1a24bbd95aa3ba7b012ddd5886e03", + "last_modified": null + }, + "isamples_202609_search_index_v1/df.parquet": { + "size_bytes": 7990281, + "etag": "a42de35cafeeb8bfae58c91ef09023e9", + "last_modified": null + } + }, + "search_index": { + "path": "isamples_202609_search_index_v1/", + "base_shard_count": 256, + "base_shard_bytes": 528220676, + "total_shard_files": 1025, + "hot_shard_files": 769, + "hot_token_count": 149, + "runtime_sidecars": [ + "build_stats.json", + "hot_tokens.json", + "shard_sizes.json", + "hot_topk.parquet" + ], + "offline_only": [ + "df.parquet" + ], + "note": "base shards from shard_sizes.json; hot sub-files enumerated from hot_tokens.json; totals cross-checked vs build_stats shard_files" + }, + "docs": "CANONICAL.md (human twin); #334 v0-detect" +} diff --git a/query-spec.qmd b/query-spec.qmd index a224879..095f358 100644 --- a/query-spec.qmd +++ b/query-spec.qmd @@ -22,9 +22,9 @@ Known instances (illustrative, not exhaustive): - **Filenames are a generation stale.** The bindings below name `sample_facets_v2.parquet`; the Explorer now reads - `isamples_202608_sample_facets_v4.parquet` (and `samples_map_lite_v3`). The + `isamples_202609_sample_facets_v4.parquet` (and `samples_map_lite_v3`). The authoritative inventory is - [`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json), + [`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json), human twin `CANONICAL.md`. - **Text search is described as the old full-scan path.** The `text MATCHES` binding presents the `ILIKE` scan as what happens "currently" and the diff --git a/tests/playwright/explorer-characterization.spec.js b/tests/playwright/explorer-characterization.spec.js index 62c67ee..aabf1e4 100644 --- a/tests/playwright/explorer-characterization.spec.js +++ b/tests/playwright/explorer-characterization.spec.js @@ -3,7 +3,7 @@ * * These tests (all tagged [data]) pin the behaviors that Codex review of the * PR1 smoke gate named as missing characterization coverage. They depend on - * remote parquet loads from data.isamples.org (202608 dataset) and are + * remote parquet loads from data.isamples.org (202609 dataset) and are * intentionally NOT in the CI smoke gate (explorer-e2e.yml stays unchanged). * Run manually or via workflow_dispatch with spec_filter=explorer-characterization. * diff --git a/tests/playwright/facet-tree.spec.js b/tests/playwright/facet-tree.spec.js index adbb5b3..0eb4034 100644 --- a/tests/playwright/facet-tree.spec.js +++ b/tests/playwright/facet-tree.spec.js @@ -78,7 +78,9 @@ test.describe('Material facet tree (#281/#282 preview)', () => { expect(total).toBeGreaterThan(0); }); - // Known 202608 subtree/union totals — deterministic for this dataset, so polling + // Known subtree/union totals — identical in 202608 and 202609 (tree summaries + // compared row-for-row; the mineral ∪ soil union recounted from both + // generations' membership files, 2026-09-10) — deterministic, so polling // to the exact value also guarantees the filter has applied (no stale-pager read). const EARTHMATERIAL_TOTAL = 4091133; // earthmaterial subtree const MINERAL_OR_SOIL_TOTAL = 333253; // mineral ∪ soil (peers) diff --git a/tests/playwright/fts-v1.spec.js b/tests/playwright/fts-v1.spec.js index 44e071f..70d7a67 100644 --- a/tests/playwright/fts-v1.spec.js +++ b/tests/playwright/fts-v1.spec.js @@ -1,12 +1,13 @@ // #171/#172: substrate search path (DEFAULT since 2026-07-17; ?fts=off // escape hatch, ?fts=v1 still accepted) — end-to-end against the published -// index at data.isamples.org/isamples_202608_search_index_v1/. +// index at data.isamples.org/isamples_202609_search_index_v1/. // // Run: BASE_URL=https://rdhyee.github.io/isamplesorg.github.io \ // npx playwright test tests/playwright/fts-v1.spec.js // -// Ground truths are properties of the 202608 index (see PR #329's -// committed build stats): 'axial seamount summit caldera' → 284; +// Ground truths were first recorded on the 202608 index (see PR #329's +// committed build stats) and re-observed unchanged on the 202609 index +// (2026-09-10, in-browser): 'axial seamount summit caldera' → 284; // 'pottery Cyprus' → 1,305; 'basalt' alone → 785. const { test, expect } = require('@playwright/test'); @@ -113,7 +114,7 @@ test.describe('default-flip routing (#172)', () => { window.__searchFilter && window.__searchFilter.substrate === true, null, { timeout: 120_000 }); const total = await page.evaluate(() => window.__searchFilter.total); - expect(total).toBe(785); // == 'basalt' ground truth for the 202608 index + expect(total).toBe(785); // == 'basalt' ground truth (202608 and 202609 indexes) }); test('escape hatch: ?fts=off → interim search', async ({ page }) => { diff --git a/tests/playwright/pin-overlay.spec.js b/tests/playwright/pin-overlay.spec.js index 21922a9..87a9a5f 100644 --- a/tests/playwright/pin-overlay.spec.js +++ b/tests/playwright/pin-overlay.spec.js @@ -15,7 +15,7 @@ // window.__clickSearchPin(i) → replays the on-globe result-pin click ceremony // (shared helper + pid hash push) by index, since // Cesium canvas picking isn't feasible in Playwright. -// Ground truths reuse the FTS suite's 202608 index facts: 'pottery Cyprus' → +// Ground truths reuse the FTS suite's index facts (same in 202608/202609): 'pottery Cyprus' → // 1,305 matches; 'basalt' → 785 — both exceed the LIMIT 50 display cap, so the // displayed (and pinned) set is capped at 50. 'ark:/28722/k2000hz7r' is the FTS // suite's known-present single pid. diff --git a/tests/test_data_origin_contract.py b/tests/test_data_origin_contract.py index e038d92..56815e0 100644 --- a/tests/test_data_origin_contract.py +++ b/tests/test_data_origin_contract.py @@ -33,7 +33,7 @@ ORIGIN = os.environ.get("ISAMPLES_DATA_ORIGIN", "https://data.isamples.org") PAGE_ORIGIN = "https://isamples.org" -MANIFEST = pathlib.Path(__file__).resolve().parent.parent / "isamples_202608_release_manifest.json" +MANIFEST = pathlib.Path(__file__).resolve().parent.parent / "isamples_202609_release_manifest.json" # The Worker 403s some default user agents; identify honestly. UA = {"User-Agent": "isamples-ci-contract/1.0 (+https://isamples.org)"} @@ -56,8 +56,8 @@ def _boot_critical_large_file(): """ files = _manifest_files() name = None - for candidate in ("isamples_202608_samples_map_lite_v3.parquet", - "isamples_202608_sample_facet_masks.parquet"): + for candidate in ("isamples_202609_samples_map_lite_v3.parquet", + "isamples_202609_sample_facet_masks.parquet"): if candidate in files: name = candidate break diff --git a/tools/build_release_manifest.py b/tools/build_release_manifest.py index de9f38e..320884b 100644 --- a/tools/build_release_manifest.py +++ b/tools/build_release_manifest.py @@ -10,7 +10,7 @@ Usage: python3 tools/build_release_manifest.py # print JSON - python3 tools/build_release_manifest.py --out isamples_202608_release_manifest.json + python3 tools/build_release_manifest.py --out isamples_202609_release_manifest.json python3 tools/build_release_manifest.py --base https://data.isamples.org Fail-closed: any missing file or non-strict probe result aborts with exit 1 — a @@ -23,39 +23,39 @@ import sys import urllib.request -RELEASE_ID = "isamples_202608" +RELEASE_ID = "isamples_202609" # The canonical set. Keep in lockstep with CANONICAL.md §1 — a suffix bump edits # BOTH in the same change (policy in CANONICAL.md §2). CANONICAL_FILES = [ - "isamples_202608_wide.parquet", - "isamples_202608_samples_map_lite_v3.parquet", - "isamples_202608_sample_facets_v4.parquet", - "isamples_202608_h3_summary_res4.parquet", - "isamples_202608_h3_summary_res6.parquet", - "isamples_202608_h3_summary_res8.parquet", - "isamples_202608_facet_summaries.parquet", - "isamples_202608_facet_cross_filter.parquet", - "isamples_202608_facet_tree_summaries.parquet", - "isamples_202608_facet_tree_cross_filter.parquet", - "isamples_202608_sample_facet_membership.parquet", - "isamples_202608_sample_facet_masks.parquet", - "isamples_202608_facet_node_bits.parquet", - "isamples_202608_sample_facet_index.parquet", - "isamples_202608_sample_facet_index_meta.parquet", - "vocab_labels_202608.parquet", + "isamples_202609_wide.parquet", + "isamples_202609_samples_map_lite_v3.parquet", + "isamples_202609_sample_facets_v4.parquet", + "isamples_202609_h3_summary_res4.parquet", + "isamples_202609_h3_summary_res6.parquet", + "isamples_202609_h3_summary_res8.parquet", + "isamples_202609_facet_summaries.parquet", + "isamples_202609_facet_cross_filter.parquet", + "isamples_202609_facet_tree_summaries.parquet", + "isamples_202609_facet_tree_cross_filter.parquet", + "isamples_202609_sample_facet_membership.parquet", + "isamples_202609_sample_facet_masks.parquet", + "isamples_202609_facet_node_bits.parquet", + "isamples_202609_sample_facet_index.parquet", + "isamples_202609_sample_facet_index_meta.parquet", + "vocab_labels_202609.parquet", ] -# The search index is a directory of ~852 objects; the manifest pins its +# The search index is a directory of ~1,030 objects (202609); the manifest pins its # self-describing sidecars (the index validates itself via build_stats). SEARCH_INDEX_SIDECARS = [ - "isamples_202608_search_index_v1/build_stats.json", - "isamples_202608_search_index_v1/hot_tokens.json", - "isamples_202608_search_index_v1/shard_sizes.json", - "isamples_202608_search_index_v1/hot_topk.parquet", # loaded by the topk query path + "isamples_202609_search_index_v1/build_stats.json", + "isamples_202609_search_index_v1/hot_tokens.json", + "isamples_202609_search_index_v1/shard_sizes.json", + "isamples_202609_search_index_v1/hot_topk.parquet", # loaded by the topk query path # df.parquet is OFFLINE-ONLY per SEARCH_INDEX_V1.md — archived, not runtime; # it is probed for existence but excluded from the Explorer's runtime check. - "isamples_202608_search_index_v1/df.parquet", + "isamples_202609_search_index_v1/df.parquet", ] @@ -132,7 +132,7 @@ def main(): inv = {"base_shard_count": None, "base_shard_bytes": None, "total_shard_files": None, "hot_shard_files": None} try: - sizes = fetch_json(args.base, "isamples_202608_search_index_v1/shard_sizes.json") + sizes = fetch_json(args.base, "isamples_202609_search_index_v1/shard_sizes.json") entries = sizes if isinstance(sizes, dict) else {} if not entries: errors.append("shard_sizes.json: empty/unexpected shape") @@ -142,7 +142,7 @@ def main(): except Exception as e: # noqa: BLE001 errors.append(f"shard_sizes.json inventory: {e}") try: - stats = fetch_json(args.base, "isamples_202608_search_index_v1/build_stats.json") + stats = fetch_json(args.base, "isamples_202609_search_index_v1/build_stats.json") # build_stats: shard_count = LOGICAL shards (256, matches shard_sizes); # shard_files = PHYSICAL files (base + hot-token sub-files). tot = stats.get("shard_files") @@ -162,7 +162,7 @@ def main(): try: # Hot sub-files enumerated from their authoritative source: each # hot_tokens.json entry declares its physical sub_files count. - ht = fetch_json(args.base, "isamples_202608_search_index_v1/hot_tokens.json") + ht = fetch_json(args.base, "isamples_202609_search_index_v1/hot_tokens.json") toks = ht.get("tokens") if not isinstance(toks, dict) or not toks: errors.append("hot_tokens.json: missing/empty tokens map") @@ -199,7 +199,7 @@ def main(): "base": args.base, "files": files, "search_index": { - "path": "isamples_202608_search_index_v1/", + "path": "isamples_202609_search_index_v1/", **inv, "runtime_sidecars": ["build_stats.json", "hot_tokens.json", "shard_sizes.json", "hot_topk.parquet"], diff --git a/tutorials/explorer_guided_tour.qmd b/tutorials/explorer_guided_tour.qmd index 72a1adf..9e5e393 100644 --- a/tutorials/explorer_guided_tour.qmd +++ b/tutorials/explorer_guided_tour.qmd @@ -75,9 +75,9 @@ sql = async (q) => { // The canonical data files — the SAME ones the Explorer loads (see CANONICAL.md). // All three used on this page are kilobyte-scale, so these cells load instantly. -FACET_SUMMARIES = 'https://data.isamples.org/isamples_202608_facet_summaries.parquet' -VOCAB_LABELS = 'https://data.isamples.org/vocab_labels_202608.parquet' -H3_RES4 = 'https://data.isamples.org/isamples_202608_h3_summary_res4.parquet' +FACET_SUMMARIES = 'https://data.isamples.org/isamples_202609_facet_summaries.parquet' +VOCAB_LABELS = 'https://data.isamples.org/vocab_labels_202609.parquet' +H3_RES4 = 'https://data.isamples.org/isamples_202609_h3_summary_res4.parquet' ``` ## Step 1 — The world at a glance diff --git a/tutorials/index.qmd b/tutorials/index.qmd index fc02fb3..27ce9ad 100644 --- a/tutorials/index.qmd +++ b/tutorials/index.qmd @@ -5,18 +5,18 @@ number-sections: false --- ::: {.callout-important} -## Snapshot note (2026-08-05) +## Snapshot note (2026-08-05; live generation updated 2026-09) The **`202601`** file examples on this page pin a stable, versioned snapshot. Those objects still exist and their byte counts are correct, but they are **not -what the Interactive Explorer serves today** — the live app reads the **`202608`** +what the Interactive Explorer serves today** — the live app reads the **`202609`** family, including `sample_facets_v4`, `samples_map_lite_v3`, `wide`, and the sharded search index. (This page may also reference other generations, such as the `current/` alias or the `202512` narrow file.) Use `202601` when you want a stable, citable reference. For "what the Explorer is actually loading right now", the authoritative inventory is -[`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json) +[`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json) (human twin: `CANONICAL.md`), which the Explorer cross-checks at boot. ::: diff --git a/tutorials/why_h3.qmd b/tutorials/why_h3.qmd index cb05ab7..2b4df41 100644 --- a/tutorials/why_h3.qmd +++ b/tutorials/why_h3.qmd @@ -10,18 +10,18 @@ format: --- ::: {.callout-important} -## Snapshot note (2026-08-05) +## Snapshot note (2026-08-05; live generation updated 2026-09) The **`202601`** file examples on this page pin a stable, versioned snapshot. Those objects still exist and their byte counts are correct, but they are **not -what the Interactive Explorer serves today** — the live app reads the **`202608`** +what the Interactive Explorer serves today** — the live app reads the **`202609`** family, including `sample_facets_v4`, `samples_map_lite_v3`, `wide`, and the sharded search index. (This page may also reference other generations, such as the `current/` alias or the `202512` narrow file.) Use `202601` when you want a stable, citable reference. For "what the Explorer is actually loading right now", the authoritative inventory is -[`isamples_202608_release_manifest.json`](https://data.isamples.org/isamples_202608_release_manifest.json) +[`isamples_202609_release_manifest.json`](https://isamples.org/isamples_202609_release_manifest.json) (human twin: `CANONICAL.md`), which the Explorer cross-checks at boot. :::