Compress noisy command output before handing it to your LLM.
Tests, diffs, logs, stack traces - thousands of tokens your agent never needed.
qq distills them down to what matters. Save up to ~99% of tokens.
Pipe any command's output through qq with a short question:
# Install
bun add -g quick-question-llm
ollama pull qwen3.5:2b
# Did the tests pass?
bun test 2>&1 | qq "did the tests pass?"
# > 3 tests failed: auth.test.ts lines 42, 78, 91
# Other great use cases
git diff | qq "what changed?"
rg -n "auth|session|token" src/ | qq "where is authentication handled?"
terraform plan 2>&1 | qq "is this safe to apply?"- Token compression - distill verbose command output into tight, signal-only summaries
- Three providers -
ollama(default),openai(any compatible API), andlocal(built-innode-llama-cpp) - No cloud required - Ollama and local modes run entirely on your machine, no API key needed
- Watch mode - repeating output from
watch/tail -fis delta-summarized, not re-summarized every cycle - Interactive passthrough -
[Y/N],Password:, and other prompts pass through raw, no LLM involvement - Bad-summary fallback - if the model produces a poor summary, the original input passes through unchanged
- Per-provider model config -
qq config modelis tracked per provider, so switching providers restores your last model choice - Machine-readable progress -
QQ_PROGRESS_PROTOCOL=stderrfor structured markers in CI or agent harnesses - Built with Effect - typed errors, resource safety, and composable services end to end
Command output is one of the biggest sources of token waste in AI-assisted workflows. A full test run, a large git diff, or a recursive rg search can easily hit tens of thousands of tokens. Your agent doesn't need all of that - it needs an answer.
rg -n "auth|session|token" src/ | qq "where is authentication handled?"| Tokens | Characters | |
|---|---|---|
| Before | ~6,400 | ~25,600 |
| After | ~80 | ~320 |
| Saved | ~99% | ~99% |
The numbers vary by input, but the pattern is consistent: noisy command output compresses aggressively. Short, targeted questions compress even more.
qq is built around Qwen 3.5 - a model family so good at compression and summarization that the 2B variant runs on nearly any computer. No GPU required. If your machine can run a browser, it can run qq.
qq supports three providers. Ollama is the default.
Runs locally via Ollama. Free, private, no account needed.
ollama pull qwen3.5:2b # make sure ollama serve is running
git diff | qq "what changed?"Works with OpenAI or any OpenAI-compatible API (Groq, LM Studio, Together, etc.):
git log --oneline -50 | qq --provider openai --model gpt-4o-mini --api-key sk-... "summarize recent work"# Switch to OpenAI and save credentials
qq config provider openai
qq config api-key "sk-..."
qq config model "gpt-4o-mini"
# Switch to local - your OpenAI model is remembered
qq config provider local
# Switch back - OpenAI model restored automatically
qq config provider openai
# Other useful config
qq config max-tokens 400
qq config thinking falseqq config model is tracked per provider, so switching between ollama, openai, and local always restores the last model you set for that provider.
View the current effective config at any time:
qq configThe local provider runs models directly via node-llama-cpp - no Ollama install, no server process, no cloud account.
qq config provider local
git diff | qq "what changed?"On first run, qq automatically downloads the default model:
hf:unsloth/Qwen3.5-2B-GGUF/Qwen3.5-2B-Q4_K_M.gguf
You can point to any GGUF model on Hugging Face:
qq config model "hf:unsloth/Qwen3.5-2B-GGUF/Qwen3.5-2B-Q4_K_M.gguf"Local provider notes:
- Models are cached under your qq config directory, typically
~/.config/qq/models hostandapi-keyflags are ignored for the local providermax-tokensdefaults to200and applies across all providers- If you installed with Bun and local startup fails, run
bun pm trust node-llama-cppand retry
Add this to your Codex, Claude Code, OpenCode, or Cursor agent instructions:
When running shell commands whose output may be long (tests, builds, diffs, logs), pipe
the output through `qq` with a short question rather than capturing the raw output.
Examples:
bun test 2>&1 | qq "did the tests pass?"
git diff HEAD~1 | qq "summarize the changes"
terraform plan 2>&1 | qq "is this safe to apply?"
This reduces token usage and keeps context focused.
For agent harnesses and CI pipelines, enable machine-readable progress markers:
QQ_PROGRESS_PROTOCOL=stderr bun test 2>&1 | qq "did the tests pass?"This emits structured __QQ_PROGRESS__:phase:collecting, __QQ_PROGRESS__:phase:summarizing, and __QQ_PROGRESS__:stop markers on stderr.
- Watch mode - repeating
watch/tail -foutput is detected automatically. Only the delta between cycles is summarized, not the full dump every time. - Interactive passthrough -
[Y/N],Password:, and similar prompts are forwarded raw with no LLM involvement. - Bad-summary fallback - if the model response is longer than the input or looks like a refusal, the original input passes through unchanged.
- Thinking mode - toggle extended model reasoning with
qq config thinking true/falseor--thinking true.
qq is developed with Bun and Effect.
git clone https://github.com/KyTiXo/quick-question.git
cd quick-question
bun installRun the full check suite (format, typecheck, lint, test):
bun run check:all:qqBuild the distributable package:
bun run buildBuild standalone binaries for all platforms:
bun run build:allManual smoke tests:
bun run test:live
bun run test:explodeContributions are welcome. Please open an issue to discuss larger changes before submitting a PR.
# Fork, clone, branch
bun install
# Make your changes
bun run check:all:qq # must pass before submittingStop wasting tokens. Start asking quick questions.
npm install -g quick-question-llm