Skip to content

Add GROUP BY and aggregates - #6

Merged
jessestimpson merged 1 commit into
mainfrom
claude/sql-group-by
Sep 26, 2026
Merged

jessestimpson merged 1 commit into
mainfrom
claude/sql-group-by

Conversation

@jessestimpson

Copy link
Copy Markdown
Contributor

efsql can now count and summarize, not just list rows:

select count(*) from acme.orders;
select status, count(*) as n, sum(total) from acme.orders group by status;
select status from acme.orders where total > 10 group by status order by count(*) desc limit 3;

What's supported

  • count(*), count(field), sum, min, max and avg, each optionally named with AS. An aggregate without GROUP BY gives one row for the whole table.
  • ORDER BY a group field, an aggregate or its alias. An aggregate you only sort by doesn't have to be selected; it's computed and dropped. LIMIT limits the groups.
  • SQL semantics: count(*) counts rows, everything else skips NULLs, and NULLs form one group. An aggregate over no rows is still one row (count(*) is 0). sum/avg handle numbers and Decimals, and min/max work on strings and datetimes too.
  • Equal values group together even when their terms differ, like Decimal 1.0 and 1.00, or timestamps stored at different precisions.
  • Mistakes get specific messages: "name must be in GROUP BY or used in an aggregate", "lower() is not supported; the aggregates are count, sum, min, max, avg", "only count takes *".

HAVING is still refused by name.

How it runs

The planner plans the underlying scan the same way as any other query, reading only the fields that are grouped or aggregated. So WHERE keeps its index and primary key pushdown. After that, a new {:aggregate, group_by, aggregates} operator (Efsql.Aggregate) groups the pulled rows, then sort, limit and project run as usual. Efsql.Aggregate is pure, so its tests don't need FoundationDB.

In the grammar this adds GROUP BY, aggregate calls and AS, with no conflicts. The error messages for the new syntax come straight from the grammar. For example, after select a the parser now says expected FROM, '(' or ','.

Also

  • The editor completes group / by, the help page has a Grouping section, and the README documents all of it.
  • The random query generator now produces aggregates and GROUP BY, so the round-trip and never-crash property tests cover them.

Tests

  • New tests for Efsql.Aggregate, for translating and validating grouped queries, and for parsing and completion.
  • test/group_by_test.exs runs grouped queries against FoundationDB. It also checks the plan: WHERE still goes to the index, only the needed fields are read, and the limit isn't pushed to the adapter. This CI run is the first time those tests run, since I didn't have FoundationDB locally. Everything else (199 tests) passes locally.

One thing that hasn't changed: result columns still show up alphabetically rather than in select-list order, the same as for plain queries.

🤖 Generated with Claude Code

https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko


Generated by Claude Code

Supports `select status, count(*) as n, sum(total) from t group by status`
with count(*), count(field), sum, min, max and avg. Aggregates can take an
alias. An aggregate without GROUP BY covers the whole table. ORDER BY can
name a group field, an aggregate or an alias, and LIMIT limits the groups.

The grammar gains GROUP BY, aggregate calls in the select list and ORDER
BY, and AS aliases. The translator checks grouping rules (selected fields
must be grouped, known aggregate names, * only for count) and names the
output columns. The planner plans the underlying scan as before, reading
only the fields that are grouped or aggregated, so WHERE keeps index and
primary-key pushdown. An {:aggregate, ...} operator (Efsql.Aggregate) then
groups the rows, followed by sort, limit and project.

SQL semantics: count(*) counts rows and the other aggregates skip NULLs.
NULLs form one group. Values that compare equal group together even when
their terms differ, like Decimal 1.0 and 1.00. An aggregate over no rows
still returns one row. sum and avg are Decimal when any input is.

HAVING stays unsupported. Completion, the help page and the README cover
the new syntax, and the random query generator produces aggregates and
GROUP BY for the property tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PFBZUrerge6bbWVCHpS5Ko
@jessestimpson
jessestimpson merged commit 3e476aa into main Sep 26, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants