Repository navigation
fix(knowledge): read notes with a UTF-8 BOM correctly - #311
Merged
petertzy merged 2 commits intoOct 2, 2026
Merged
Conversation
A note saved with a BOM was indexed under the wrong title and had its YAML
frontmatter injected verbatim into the AI retrieval context.
U+FEFF is category Cf, not whitespace, so str.strip() does not remove it. It
sits in front of line 1, which breaks both content.startswith('---') in
extract_note_title() and chunk_markdown_document(), plus the first-heading
match. So '\xef\xbb\xbf# My Note' was titled 'my note' (from the filename),
and a frontmattered note kept '---\ntitle: X\n---' as its first chunk, which
reaches the AI system prompt verbatim.
Read as utf-8-sig, which strips a BOM when present and is byte-identical to
utf-8 when it is not. Exposed as NOTE_READ_ENCODING so the behaviour is
named and testable.
Adds 5 cases; verified 3 fail when the encoding is flipped back to utf-8.
Owner
|
Approved. The utf-8-sig change correctly handles UTF-8 BOMs while preserving behavior for regular UTF-8 files. It fixes both incorrect title extraction and frontmatter leakage into indexed chunks. The added tests cover parsing and the end-to-end indexing path, and all 201 tests pass with Ruff checks clean. This change is safe and worthwhile to merge. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
A note saved with a UTF-8 BOM was indexed under the wrong title, and its
YAML frontmatter was injected verbatim into the text sent to the AI as
retrieval context.
Both come from the same thing: the BOM ends up in front of line 1.
Why a BOM matters here
U+FEFFis Unicode categoryCf(format), not whitespace, sostr.strip()does not remove it. Notes with a BOM are routine — anythingauthored or re-saved on Windows picks one up.
Before
Title.
extract_note_title()checkscontent.startswith("---")forfrontmatter and then
line.startswith("# ")for a heading. With a BOM in front,both fail:
# My Notemy note(from the filename)---\ntitle: Real Title\n…Heading(the H1, not the frontmatter title)The wrong title is stored in
indexed_files.titleand in everychunks.title,so it shows up in AI citations and in
build_knowledge_context_for_prompt.Frontmatter leak.
chunk_markdown_document()also gates onbody.startswith("---"), so with a BOM the frontmatter is never stripped andbecomes the first chunk:
That raw YAML then goes into the AI system prompt verbatim.
Changes
backend/knowledge_logic.py— read notes asutf-8-sig, which strips a BOMwhen present and is byte-identical to
utf-8when it is not. This is the oneread site that feeds both the title and the chunker, so it fixes both symptoms
at the source rather than patching each consumer.
The encoding is exposed as a module constant (
NOTE_READ_ENCODING) so thebehaviour is named, documented and testable.
No change for notes without a BOM —
"utf-8-sig"decodes those identically.Testing
tests/test_knowledge_logic.py::TestUtf8BomNotes, five cases. Each writes a realfile with
BOM + bodyas bytes (notencoding="utf-8", which would add asecond BOM) and reads it back through
NOTE_READ_ENCODING, so reverting the fixfails these tests rather than silently passing:
# heading→My NoteReal Titlechunk)
""and chunks without raisingVerified 3 of these fail when
NOTE_READ_ENCODINGis flipped back to"utf-8".Verification
pytest— 200 passed (was 195 onmain; this PR adds 5)ruff check/ruff format --checkon changed files — cleannpm test— 23 passed, unchangednpm run lint/npm run build— clean