Skip to content

Parse typed literals in the where clause, starting with atoms - #4

Merged
jessestimpson merged 10 commits into
mainfrom
claude/sql-atom-type-parsing-o9gpf8
Sep 26, 2026
Merged

jessestimpson merged 10 commits into
mainfrom
claude/sql-atom-type-parsing-o9gpf8

Conversation

@jessestimpson

@jessestimpson jessestimpson commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

This started as "let me filter on an atom" and turned into a general pass on non-string types in the where clause, plus a few parser bugs that showed up along the way.

Typed literals

A quoted literal is still always a string. For anything else you annotate it, Postgres style (CAST(x AS type) works too):

where status = 'active'::atom
where inserted_at >= '2024-03-01 12:00:00'::timestamp   -- NaiveDateTime
where seen_at >= '2024-03-01T12:00:00Z'::timestamptz    -- DateTime, UTC
where birthday = '2024-03-01'::date
where opens_at < '09:30:00'::time

naive_datetime and utc_datetime also work, for anyone thinking in Ecto types. A timestamp literal with an offset is rejected instead of quietly dropping the offset. A timestamptz without one is taken as UTC. Adding a new type is a clause or two in Efsql.Types.

Parsing was the easy part. Getting the right rows back needed two more fixes:

  • The executor compared values with plain </==. On date and time structs that compares fields, not time, so 2025-01-02 sorted before 2024-01-03. Filtering, IN and ORDER BY now go through Types.compare/2. The same function fixes numbers against Decimal columns: before, price < 100 never matched a decimal and price > 100 matched all of them.
  • efsql queries without a schema, so the adapter encoded datetime params with term_to_binary and never matched the index keys. That's why the README used to say to query indexed datetimes by their string form. The planner now encodes dates and times the same way the indexer writes them.

Range queries on an indexed decimal still can't work, because the adapter keys decimals by term encoding rather than value. That limit is in the README.

IS NULL

is null, is not null, plus the Postgres isnull/notnull shorthands. x = null now gives an error pointing at IS NULL instead of a FunctionClauseError.

Working around the parser

The sql library hangs or misparses in a few places, so efsql rewrites the token stream before handing it over. Efsql.Parser.parse/1 is now the only way in. Cases handled:

  • ... and x is not null never returned.
  • Any column named after a reserved word (day, at, user, value, ...) hung on where day = 'x'. It also misparsed with in/like/between, and turned order by day desc into ascending. This is what the CI timeouts were.

The sql dependency

I tried hex first, but 0.5.0 is older than the commit we were already on, so this tracks main (f6097ff). That commit fixes a hang on and (...). It also adds a Postgres client with C NIFs, whose startup opens a dets file in priv and sleeps 100ms. We only use the lexer and parser, so sql now goes in included_applications: it's loaded but never started. Building still needs cc/make for the NIFs, which is the main argument for eventually vendoring the lexer and parser.

CI

New test.yml runs mix test on PRs and pushes to main against a real fdbserver, pinned to the same versions as release.yml. It's green.

🤖 Generated with Claude Code

https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko

A quoted literal stays a string; an atom is written as a type-annotated
literal using the PostgreSQL cast operator or standard CAST:

    where status = 'active'::atom
    where status = cast('active' as atom)

Both forms dispatch through Efsql.Types.cast/2, so adding a type is a
new clause there. Reserved type names (e.g. date) are handled as well
as identifiers.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Elixir has two datetime types and Ecto stores whichever a field
declares, so each gets its own literal type:

    '2024-03-01 12:00:00'::timestamp     -- NaiveDateTime (alias naive_datetime)
    '2024-03-01T12:00:00Z'::timestamptz  -- UTC DateTime  (alias utc_datetime)

A timestamp literal with an offset is rejected rather than silently
dropping it; a timestamptz literal without one is taken as UTC; a bare
date means midnight.

Supporting the type end to end also needed:

- Executor filters, IN and sorting compare through Efsql.Types.compare/2.
  Kernel comparison on datetime structs orders by struct fields, not
  time, and == fails when only the precision differs.
- Pushed index predicates encode datetimes the way the adapter's default
  indexer writes them. efsql queries schemaless, so the adapter would
  otherwise term_to_binary the param and never match the index key.

Also corrects the README: Ecto.Enum fields are stored as strings, so they
are queried with a string literal, not an atom.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
    '2024-03-01'::date   -- Date, for :date fields
    '09:30:00'::time     -- Time, for :time / :time_usec fields

Like the timestamp types, these needed more than parsing: term order
compares Date's struct fields with the year last, and Time's microsecond
before its minute, so filters and ORDER BY were wrong. Both now compare
chronologically, and pushed index predicates encode them as the adapter's
default indexer does.

Decimal fields need no literal type, but a number compared against one
used term order, which puts every number below every map: `price < 100`
never matched and `price > 100` matched everything. Types.compare/2 now
compares a Decimal by value against a Decimal, integer or float. A NaN
falls back to term order rather than raising. Indexes on decimal fields
remain unusable for this, since the adapter keys them by term encoding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
    where notes is null
    where notes is not null
    where notes isnull / notnull     -- PostgreSQL shorthands

A nil or absent field is NULL. Both are residual filters in the
executor; an indexed equality alongside still uses the index.

The SQL library only assembles `x is null` into a node that composes
with `and`. It leaves `x is not null` as loose tokens, never returns
from `... and x is not null`, and hangs the same way on its postfix
isnull/notnull nodes. Efsql.Parser.parse/1 now rewrites all of these to
the `is null` shape before parsing, marking negation on the `is`
token's meta, and every lex/parse call site goes through it.

`x = null` (never true in SQL) now raises Unsupported pointing to
IS NULL, instead of a FunctionClauseError. Completion offers `is`,
then `null`/`not`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
0.5.0 is the latest hex release; it predates the git commit efsql was
locked to (da69080, one large refactor after v0.5.0). efsql only uses
SQL.Lexer.lex/1 and SQL.Parser.parse/2, and its tests pass on both.

The one incompatibility: 0.5.0 lexes integers as :numeric rather than
:integer, so `where age > 40` raised in String.to_float/1. A :numeric
literal with no fractional part now parses as an integer.

The git entry is removed from mix.lock; the next `mix deps.get` writes
the hex entry with its checksums (hex.pm was unreachable here).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Back to the git dependency, now at the tip of main, one commit past the
previously locked da69080. efsql uses only SQL.Lexer.lex/1 and
SQL.Parser.parse/2; its tests pass unchanged. The :numeric literal fix
from the hex switch stays, so parsing works on either library version.

What the new commit brings for efsql:

- The parser handles `... and (expr)`, which used to never return.
  efsql now raises Unsupported for a parenthesized predicate instead of
  hanging.
- The library now compiles three C NIFs via a `make` compile alias, so
  building efsql needs make and a C compiler. The release runners
  (GitHub-hosted macOS and Ubuntu) have both.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
efsql uses only sql's lexer and parser. Listing sql in
included_applications keeps its code loaded (and in the release) but
never runs SQL.Application, which since f6097ff opens a :dets file in
its priv dir, sleeps 100 ms on every start, and crashes in stop/1 when
no pool was ever started. Mix also loads sql's own dependencies
without starting them.

The NIFs are still compiled by sql's `make` alias and loaded at boot;
a release in embedded mode fails to boot without them. Removing that
build dependency means vendoring the lexer and parser.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
Runs on pushes to main and on pull requests, on ubuntu-22.04 with the
same Elixir, OTP and FoundationDB versions as release.yml. Installs the
FDB server package as well as the client: EctoFoundationDB.Sandbox runs
the tests against an fdbserver it starts itself. Mirrors
ecto_foundationdb's own CI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
The SQL library lexes a reserved word as its own token even where it
names a column. For `where day = 'x'` or `where at >= ...` its parser
never returns; with in, like, is null or between it misparses, and
`order by day desc` silently became [asc: :day, asc: :desc]. Only
double-quoting the column avoided it.

Efsql.Parser.parse/1 now rewrites a reserved word that directly
precedes a predicate operator or asc/desc into an identifier, except
for words that legitimately sit there (true, null, end, where, ...).

This is why seven parser tests timed out in CI: their `at` and `day`
columns are reserved in the real SQL grammar. New tests cover
reserved-word columns across every predicate form and in order by.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
In the real SQL grammar asc and desc are non-reserved words, lexed as
identifiers that carry the word in their meta's :tag, so the
reserved-word column rewrite never fired before them and
`order by day desc, user asc` reached efsql as a bare comma node.
Followers are now matched by tag for identifier tokens too. The
paren clause moves first so a leading (...) is still normalised.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
@jessestimpson
jessestimpson merged commit a70e3fe into main Sep 26, 2026
1 check passed
@jessestimpson
jessestimpson deleted the claude/sql-atom-type-parsing-o9gpf8 branch September 26, 2026 02:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants