feat(alphabets): Emoji alphabet — language-neutral, multi-codepoint nodes - #96
Merged
Merged
Conversation
…odes Dasher-Android #61 (option 2): a toolbar switch to an emoji-only alphabet works for any language via the existing alphabet-switch API — pure data, no engine changes. - alphabet.emoji.xml: 9 topical groups (~300 nodes), two zoom levels to any emoji. Includes ZWJ sequences (family) and skin-tone modifiers as natural regression coverage for multi-codepoint nodes. - training_emoji.txt: mild ordering priors (common juxtapositions). - .gitignore: root-anchor the training_*.txt leak rule — the bare pattern silently ignored Data/training/ additions. - tests: emoji loads (309 symbols); ZWJ/skin-tone/space round-trip whole through dasher_get_alphabet_symbol_text; shipped training file present. Signed-off-by: will wade <willwade@gmail.com>
…ndex Review loop (greptile P2s + subagent 7/10 findings): - Atomicity now proven at EVENT level via dasher_set_output_callback: every committed event must equal a whole node text (byte concatenation can't distinguish one 👨👩👧 event from five). Plus a deterministic 3-node test alphabet (ZWJ, VS16, plain) where multi-codepoint commits are unavoidable: 36 events, 23 multi-codepoint, all whole. - Corpus: single-codepoint nodes only (the trainer looks up one code point per symbol — ZWJ/VS16 tokens can never match); expanded to 945 tokens/4.6KB; new test enforces corpus ⊆ alphabet nodes permanently. - alphabet_index.json regenerated (475) — Emoji registered for index-driven tooling. - Makefile.am: training_emoji.txt added (keep-consistent-only; the automake subtree is not the live packaging path — CMake is). - Comment fixes (9 groups not 8; ~310 symbols not ~230). Signed-off-by: will wade <willwade@gmail.com>
- clang-format over the new tests (CI pins format 18 with --Werror) - symbol scans use the documented 1..sym_count inclusive range (0 = root), matching the file's existing v6 tests - DASHER_EVENT_OUTPUT macro instead of magic 0 (matches test_interaction.cpp style) Signed-off-by: will wade <willwade@gmail.com>
- Corpus test now enforces the single-codepoint invariant explicitly (UTF-8 lead-byte count): a multi-codepoint NODE would pass a pure membership check, yet the trainer (one code point per symbol lookup) can never match it — membership alone was insufficient. - Symbol scans use the exclusive bound: CAPI rejects index >= iEnd (dasher_get_alphabet_symbol_text returns -1), so i < sym_count is the correct range; the previous <= probed an invalid index. Signed-off-by: will wade <willwade@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
@/tmp/opencode/pr-emoji-core.md
The PR appears safe to merge; no outstanding functional or repository-rule violations remain.
Findings
Summary
Adds a language-neutral Emoji alphabet with grouped multi-codepoint symbols, a single-code-point training corpus, generated alphabet-index metadata, packaging support, and tests for loading, corpus validity, and atomic multi-codepoint output.
training_emoji.txtand narrows the root training-file ignore pattern.Diagram
%%{init: {'theme': 'neutral'}}%% flowchart LR A[Emoji alphabet XML] --> B[Alphabet loader] C[Single-codepoint training corpus] --> D[Language model priors] B --> E[Grouped emoji nodes] D --> E E --> F[Text output callback] F --> G[Atomic multi-codepoint output]Reviews (4) · Last reviewed commit: "fix(emoji): greptile round 2 — single-co..."