Skip to content

Replace the sql dependency with our own lexer and parser - #5

Merged
jessestimpson merged 5 commits into
mainfrom
claude/sql-atom-type-parsing-o9gpf8
Sep 26, 2026
Merged

jessestimpson merged 5 commits into
mainfrom
claude/sql-atom-type-parsing-o9gpf8

Conversation

@jessestimpson

@jessestimpson jessestimpson commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

We only ever used elixir-dbvisor/sql for lexing and parsing, and we'd piled up a fair number of token rewrites to get around it. ... and x is not null and any column named after a reserved word (where day = 'x') made it hang, and order by day desc quietly flipped to ascending. On top of that, the latest version brings C NIFs and a runtime we had to keep from starting. So this writes a small front end for the dialect efsql actually speaks and drops the dependency.

What's in lib/efsql/sql/

  • Lexer. Hand-written over binaries. Unquoted names fold to lower case like Postgres, quoted ones keep their case. It handles ''/"" escapes, nested /* */ comments, CRLF, Unicode names and whitespace, and exponent floats. Every error carries a line and column: syntax error at line 3, column 10: expected an expression, got '='.
  • Grammar. src/efsql_sql_grammar.yrl, compiled by yecc (Mix does this on its own, nothing to install). Output is a documented AST (Efsql.SQL.AST). day, user, date and value all work as bare column names. The handful of words that can't (order, select, limit, and, ...) are listed, and the error tells you to quote them.
  • Error messages. yecc only tells you where it stopped. To say what it wanted, the parser replays the tokens up to that point with every terminal and keeps the ones the grammar accepts. So messages come from the grammar itself (expected BY, got 'a', expected NOT or NULL, got 'true'), and new rules need nothing extra to get good errors. On top of that there's a short list of hints for things we don't support yet: GROUP BY, joins, WITH, functions, subqueries, DISTINCT. Each hint goes away when its feature lands.

Efsql.Parser is now just the AST → Efsql.Logical translation, and all the old workarounds are gone. It also catches two inputs that used to crash: field names too long to be atoms, and versionstamps outside 0..2^96-1.

Why yecc for the parser

I wrote the parser twice, once by hand (recursive descent) and once with leex/yecc, and got them to give identical output on all the tests and 60k fuzzed inputs before comparing. For today's dialect the hand-written one was a bit smaller and twice as fast. But CTEs, joins and GROUP BY are coming, and that's where a grammar yecc checks for conflicts pays off. (select …) vs (expr) vs a tuple, aliases vs keywords, and join chains are all ambiguities that yecc reports at build time and a hand parser just hides. Speed doesn't matter at this size: about 50µs for a typical query, 0.4s for a 550KB one.

The lexer stays hand-written. leex can't do nested comments, Unicode classes or character columns, and it behaves differently on OTP 25 and 28, which broke CI once. Grammar growth only adds keywords to the lexer anyway.

Smaller things

  • Typed literals now work as BETWEEN bounds, and the README is updated to match.
  • The CLI prints exception messages instead of inspected structs, so syntax errors are readable.
  • Efsql.lex_and_parse/1 is gone. It returned the old library's parse tree and nothing called it.
  • sql, unicode, unicode_set and nimble_parsec are out of mix.lock. No more C toolchain needed to build.

Behaviour changes worth knowing

  • Clauses have to come in standard order (WHERE, ORDER BY, LIMIT).
  • The reserved words above need quotes when used as column names.

Tests

  • Lexer and parser suites covering every token, construct and error, with exact messages and positions.
  • The existing translation tests pass unchanged. There are new ones for the translator errors, plus every query on the TUI help page.
  • Randomized tests:
    • Random trees get printed with randomized case, comments, whitespace, quoting and alternate spellings (!=, CAST(...), ISNULL, timestamp with time zone), and must parse back to the exact tree.
    • Mangled queries and random bytes must either parse or give a positioned SyntaxError, never anything else.
    • A 20,000-condition query has to parse and translate quickly.

The fuzzing found one real crash: float literals too big for a double (1e400), now a syntax error.

🤖 Generated with Claude Code

https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko


Generated by Claude Code

efsql used elixir-dbvisor/sql only for lexing and parsing, and needed a
growing set of token rewrites to work around it: `and x is not null`
and any column named after a reserved word (`where day = 'x'`) never
returned, and `order by day desc` silently flipped direction. The
dependency also brought C NIFs and a runtime that efsql had to keep from
starting.

Efsql.SQL is a small front end written for the dialect efsql speaks:

- Lexer: hand-written over binaries. Folds unquoted names, keeps quoted
  ones, handles '' and "" escapes, nested /* */ comments, CRLF, Unicode
  names and whitespace, and reports every error with line and column.
- Parser: recursive descent into a documented AST (Efsql.SQL.AST).
  Keywords are only keywords where the grammar expects one, so day,
  user, date or value are plain names; the handful that must be quoted
  are listed and the error says to quote them. Unsupported features (OR,
  joins, functions, GROUP BY, subqueries...) are rejected by name, and
  nesting depth is capped.
- Efsql.Parser now just translates the AST to Efsql.Logical. It also
  validates what used to crash: field names too long for an atom, and
  versionstamps outside 0..2^96-1.

Typed literals now also work as BETWEEN bounds. The CLI prints
exception messages instead of inspected structs, so syntax errors read
cleanly.

Tests: lexer and parser suites covering every token and construct with
exact positions; the existing translation tests unchanged; every query
on the TUI help page; and randomized tests that print random trees with
randomized spelling (case, comments, whitespace, quoting, alternate
syntax) and require the exact tree back, plus fuzzing of mangled queries
and random bytes, which must yield a result or a positioned SyntaxError
and never crash. The fuzzing found and fixed one crash: float literals
beyond a double's range.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
`1..Enum.random(0..40)` counts down to 0 when the length drawn is 0, so
that case produced two characters instead of none and warned on Elixir
1.19. An explicit //1 step makes it an empty range.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Efsql.SQL.Leex and Efsql.SQL.Yecc keep the same contracts as
Efsql.SQL.Lexer and Efsql.SQL.Parser, down to error messages and
positions. They're built on src/efsql_sql_leex.xrl and
src/efsql_sql_yecc.yrl, which Mix compiles automatically. yecc only
reports the token where it stopped, so Efsql.SQL.Yecc.Explain works out
the wording from the tokens that came before.

The lexer, parser and property suites now run against both
implementations. A new differential test checks that the two agree on
random valid, mangled and garbage input. Also fixes the hand lexer's
column for the end of input after a trailing line comment.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
OTP 28's leex tracks columns by UTF-8 width, and its adjust_col/3 has no
clause for integers above U+10FFFF, so input with an invalid byte crashed
the leex lexer on CI (OTP 25, used locally, has no column tracking).
Surrogates never come out of UTF-8 decoding, and leex can count them.
The positive character classes skip the surrogate block, so the bytes
are still illegal outside strings and comments.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
efsql's SQL is about to grow CTEs, joins and GROUP BY. A grammar that
yecc checks for conflicts scales better than more recursive descent, so
Efsql.SQL.Parser now drives src/efsql_sql_grammar.yrl with the
hand-written lexer's tokens.

yecc only reports the token where it stopped. To say what it expected,
the parser replays the tokens before that point with each terminal and
keeps the ones the grammar accepts: "expected BY, got 'a'". New rules
need no extra work to get error messages. A short list of hints still
names unsupported features (GROUP BY, JOIN, WITH, functions, ...); each
hint goes away when its feature lands. Nesting no longer has a depth
limit.

Removes the leex lexer, the hand-written parser, the error explainer
and the differential test. The lexer stays hand-written, since leex
can't do nested comments, Unicode classes or character columns, and its
behaviour differs between OTP versions. Error assertions now match the
new messages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
@jessestimpson
jessestimpson merged commit 315e94c into main Sep 26, 2026
1 check passed
@jessestimpson
jessestimpson deleted the claude/sql-atom-type-parsing-o9gpf8 branch September 26, 2026 12:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants