Replace the sql dependency with our own lexer and parser - #5
Merged
Merged
Conversation
efsql used elixir-dbvisor/sql only for lexing and parsing, and needed a growing set of token rewrites to work around it: `and x is not null` and any column named after a reserved word (`where day = 'x'`) never returned, and `order by day desc` silently flipped direction. The dependency also brought C NIFs and a runtime that efsql had to keep from starting. Efsql.SQL is a small front end written for the dialect efsql speaks: - Lexer: hand-written over binaries. Folds unquoted names, keeps quoted ones, handles '' and "" escapes, nested /* */ comments, CRLF, Unicode names and whitespace, and reports every error with line and column. - Parser: recursive descent into a documented AST (Efsql.SQL.AST). Keywords are only keywords where the grammar expects one, so day, user, date or value are plain names; the handful that must be quoted are listed and the error says to quote them. Unsupported features (OR, joins, functions, GROUP BY, subqueries...) are rejected by name, and nesting depth is capped. - Efsql.Parser now just translates the AST to Efsql.Logical. It also validates what used to crash: field names too long for an atom, and versionstamps outside 0..2^96-1. Typed literals now also work as BETWEEN bounds. The CLI prints exception messages instead of inspected structs, so syntax errors read cleanly. Tests: lexer and parser suites covering every token and construct with exact positions; the existing translation tests unchanged; every query on the TUI help page; and randomized tests that print random trees with randomized spelling (case, comments, whitespace, quoting, alternate syntax) and require the exact tree back, plus fuzzing of mangled queries and random bytes, which must yield a result or a positioned SyntaxError and never crash. The fuzzing found and fixed one crash: float literals beyond a double's range. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
`1..Enum.random(0..40)` counts down to 0 when the length drawn is 0, so that case produced two characters instead of none and warned on Elixir 1.19. An explicit //1 step makes it an empty range. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Efsql.SQL.Leex and Efsql.SQL.Yecc keep the same contracts as Efsql.SQL.Lexer and Efsql.SQL.Parser, down to error messages and positions. They're built on src/efsql_sql_leex.xrl and src/efsql_sql_yecc.yrl, which Mix compiles automatically. yecc only reports the token where it stopped, so Efsql.SQL.Yecc.Explain works out the wording from the tokens that came before. The lexer, parser and property suites now run against both implementations. A new differential test checks that the two agree on random valid, mangled and garbage input. Also fixes the hand lexer's column for the end of input after a trailing line comment. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
OTP 28's leex tracks columns by UTF-8 width, and its adjust_col/3 has no clause for integers above U+10FFFF, so input with an invalid byte crashed the leex lexer on CI (OTP 25, used locally, has no column tracking). Surrogates never come out of UTF-8 decoding, and leex can count them. The positive character classes skip the surrogate block, so the bytes are still illegal outside strings and comments. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
efsql's SQL is about to grow CTEs, joins and GROUP BY. A grammar that yecc checks for conflicts scales better than more recursive descent, so Efsql.SQL.Parser now drives src/efsql_sql_grammar.yrl with the hand-written lexer's tokens. yecc only reports the token where it stopped. To say what it expected, the parser replays the tokens before that point with each terminal and keeps the ones the grammar accepts: "expected BY, got 'a'". New rules need no extra work to get error messages. A short list of hints still names unsupported features (GROUP BY, JOIN, WITH, functions, ...); each hint goes away when its feature lands. Nesting no longer has a depth limit. Removes the leex lexer, the hand-written parser, the error explainer and the differential test. The lexer stays hand-written, since leex can't do nested comments, Unicode classes or character columns, and its behaviour differs between OTP versions. Error assertions now match the new messages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
We only ever used
elixir-dbvisor/sqlfor lexing and parsing, and we'd piled up a fair number of token rewrites to get around it.... and x is not nulland any column named after a reserved word (where day = 'x') made it hang, andorder by day descquietly flipped to ascending. On top of that, the latest version brings C NIFs and a runtime we had to keep from starting. So this writes a small front end for the dialect efsql actually speaks and drops the dependency.What's in
lib/efsql/sql/''/""escapes, nested/* */comments, CRLF, Unicode names and whitespace, and exponent floats. Every error carries a line and column:syntax error at line 3, column 10: expected an expression, got '='.src/efsql_sql_grammar.yrl, compiled by yecc (Mix does this on its own, nothing to install). Output is a documented AST (Efsql.SQL.AST).day,user,dateandvalueall work as bare column names. The handful of words that can't (order,select,limit,and, ...) are listed, and the error tells you to quote them.expected BY, got 'a',expected NOT or NULL, got 'true'), and new rules need nothing extra to get good errors. On top of that there's a short list of hints for things we don't support yet:GROUP BY, joins,WITH, functions, subqueries,DISTINCT. Each hint goes away when its feature lands.Efsql.Parseris now just the AST →Efsql.Logicaltranslation, and all the old workarounds are gone. It also catches two inputs that used to crash: field names too long to be atoms, and versionstamps outside 0..2^96-1.Why yecc for the parser
I wrote the parser twice, once by hand (recursive descent) and once with leex/yecc, and got them to give identical output on all the tests and 60k fuzzed inputs before comparing. For today's dialect the hand-written one was a bit smaller and twice as fast. But CTEs, joins and GROUP BY are coming, and that's where a grammar yecc checks for conflicts pays off.
(select …)vs(expr)vs a tuple, aliases vs keywords, and join chains are all ambiguities that yecc reports at build time and a hand parser just hides. Speed doesn't matter at this size: about 50µs for a typical query, 0.4s for a 550KB one.The lexer stays hand-written. leex can't do nested comments, Unicode classes or character columns, and it behaves differently on OTP 25 and 28, which broke CI once. Grammar growth only adds keywords to the lexer anyway.
Smaller things
BETWEENbounds, and the README is updated to match.Efsql.lex_and_parse/1is gone. It returned the old library's parse tree and nothing called it.sql,unicode,unicode_setandnimble_parsecare out ofmix.lock. No more C toolchain needed to build.Behaviour changes worth knowing
WHERE,ORDER BY,LIMIT).Tests
!=,CAST(...),ISNULL,timestamp with time zone), and must parse back to the exact tree.SyntaxError, never anything else.The fuzzing found one real crash: float literals too big for a double (
1e400), now a syntax error.🤖 Generated with Claude Code
https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Generated by Claude Code