Skip to content

feat(alphabets): Emoji alphabet — language-neutral, multi-codepoint nodes - #96

Merged
willwade merged 4 commits into
mainfrom
feat/emoji-alphabet
Sep 19, 2026
Merged

willwade merged 4 commits into
mainfrom
feat/emoji-alphabet

Conversation

@willwade

@willwade willwade commented Sep 18, 2026 •

Copy link
Copy Markdown

@/tmp/opencode/pr-emoji-core.md

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no outstanding functional or repository-rule violations remain.

Findings

  1. P2 Training Corpus Is Not Packaged ▶

Summary

Adds a language-neutral Emoji alphabet with grouped multi-codepoint symbols, a single-code-point training corpus, generated alphabet-index metadata, packaging support, and tests for loading, corpus validity, and atomic multi-codepoint output.

  • Adds 308 emoji and separator nodes across ten groups.
  • Packages training_emoji.txt and narrows the root training-file ignore pattern.
  • Corrects all reviewed symbol-loop upper bounds and strengthens corpus and output coverage.
Diagram
%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Emoji alphabet XML] --> B[Alphabet loader]
  C[Single-codepoint training corpus] --> D[Language model priors]
  B --> E[Grouped emoji nodes]
  D --> E
  E --> F[Text output callback]
  F --> G[Atomic multi-codepoint output]
Loading

Reviews (4) · Last reviewed commit: "fix(emoji): greptile round 2 — single-co..."

…odes

Dasher-Android #61 (option 2): a toolbar switch to an emoji-only
alphabet works for any language via the existing alphabet-switch API —
pure data, no engine changes.

- alphabet.emoji.xml: 9 topical groups (~300 nodes), two zoom levels to
  any emoji. Includes ZWJ sequences (family) and skin-tone modifiers as
  natural regression coverage for multi-codepoint nodes.
- training_emoji.txt: mild ordering priors (common juxtapositions).
- .gitignore: root-anchor the training_*.txt leak rule — the bare
  pattern silently ignored Data/training/ additions.
- tests: emoji loads (309 symbols); ZWJ/skin-tone/space round-trip
  whole through dasher_get_alphabet_symbol_text; shipped training file
  present.

Signed-off-by: will wade <willwade@gmail.com>
Comment thread Data/training/training_emoji.txt Outdated
Comment thread Data/training/training_emoji.txt Outdated
Comment thread Data/training/training_emoji.txt
Comment thread tests/test_alphabet_xml.cpp
…ndex

Review loop (greptile P2s + subagent 7/10 findings):

- Atomicity now proven at EVENT level via dasher_set_output_callback:
  every committed event must equal a whole node text (byte concatenation
  can't distinguish one 👨‍👩‍👧 event from five). Plus a deterministic
  3-node test alphabet (ZWJ, VS16, plain) where multi-codepoint commits
  are unavoidable: 36 events, 23 multi-codepoint, all whole.
- Corpus: single-codepoint nodes only (the trainer looks up one code
  point per symbol — ZWJ/VS16 tokens can never match); expanded to 945
  tokens/4.6KB; new test enforces corpus ⊆ alphabet nodes permanently.
- alphabet_index.json regenerated (475) — Emoji registered for
  index-driven tooling.
- Makefile.am: training_emoji.txt added (keep-consistent-only; the
  automake subtree is not the live packaging path — CMake is).
- Comment fixes (9 groups not 8; ~310 symbols not ~230).

Signed-off-by: will wade <willwade@gmail.com>
Comment thread tests/test_alphabet_xml.cpp
- clang-format over the new tests (CI pins format 18 with --Werror)
- symbol scans use the documented 1..sym_count inclusive range
  (0 = root), matching the file's existing v6 tests
- DASHER_EVENT_OUTPUT macro instead of magic 0 (matches
  test_interaction.cpp style)

Signed-off-by: will wade <willwade@gmail.com>
Comment thread tests/test_alphabet_xml.cpp Outdated
- Corpus test now enforces the single-codepoint invariant explicitly
  (UTF-8 lead-byte count): a multi-codepoint NODE would pass a pure
  membership check, yet the trainer (one code point per symbol lookup)
  can never match it — membership alone was insufficient.
- Symbol scans use the exclusive bound: CAPI rejects index >= iEnd
  (dasher_get_alphabet_symbol_text returns -1), so i < sym_count is the
  correct range; the previous <= probed an invalid index.

Signed-off-by: will wade <willwade@gmail.com>
@willwade
willwade merged commit b00a4c8 into main Sep 19, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant