Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 18 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,24 @@ All notable changes to this project are documented in this file.

## Unreleased

- Local-first trace viewer:
- `logquill trace <run_id> --file logs.jsonl` reconstructs and prints one
agent run's span tree, annotated with each span's own duration and the
token/cost totals rolled up from everything nested under it. It streams
the file line by line, so tracing one run out of a multi-gigabyte log
costs memory proportional to that run, not the file (a gigabyte-scale
test asserts this). `--json` prints the same tree as nested data instead.
- `logquill serve --file logs.jsonl` (or `--db logs.sqlite` for a
`SQLiteTransport` database) runs a small local web UI — run list, a
combined span-tree/waterfall view, search, and a level filter — built
entirely on the stdlib (`http.server`, `sqlite3`): no new dependency, no
account, nothing leaves the machine. Reading from a SQLite database never
shows token/cost annotations, since that transport's fixed schema doesn't
store the `llm` block.
- `logquill dev logs.jsonl` follows a file like `tail -f`, but live-renders
the current run's span tree (colorized, screen-cleared between redraws)
instead of flat lines, following whichever run is most recently active
unless `--run-id` pins it to one.
- Auto-instrumentation, an OpenAI Agents SDK adapter, and MCP trace propagation:
- `logquill.instrument.anthropic(logger)` / `.openai(logger)` / `.litellm(logger)`
patch the Anthropic, OpenAI, and litellm Python SDKs so every LLM call they
Expand Down
74 changes: 74 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ for what's landed so far.
- **Cheap when idle, precise when it counts** — `logger.opt(lazy=True)` defers expensive `meta` values until a record will really be emitted, `logger.opt(depth=N)` reports the right caller from inside a wrapper, `logquill.disable(__name__)` silences a library's own logs by default, and queued records are flushed automatically at interpreter exit — see [Lazy values, caller depth & disabling a library](#lazy-values-caller-depth--disabling-a-library)
- **Parse any log file** — `parse()` pulls structured fields out of a log file (LogQuill's own or a legacy format) with a regex, streaming line by line — see [Parsing log files](#parsing-log-files)
- **CLI** — `logquill tail app.log --level=warn --json -f` for filtering/following a JSONL log file in local dev, no extra install — see [CLI](#cli)
- **Local-first trace viewer** — `logquill trace <run_id> --file logs.jsonl` prints an annotated span tree (streamed, bounded memory at gigabyte scale); `logquill serve` runs a small offline web UI (stdlib only — run list, span tree/waterfall, search, level filter) reading JSONL or a `SQLiteTransport` database; `logquill dev` live-renders the current run as it happens — no account, nothing leaves your machine

## Install

Expand Down Expand Up @@ -1410,6 +1411,79 @@ colors) when writing to a terminal; pass `--no-color` to disable that, or
that isn't valid JSON, or isn't a JSON object, is skipped with a warning on
stderr rather than aborting the whole tail.

### `logquill trace` — one run's span tree, from the command line

`logquill trace <run_id> --file logs.jsonl` reconstructs and prints one
agent run's span tree, annotated with each span's own duration and the
token/cost totals of everything nested under it — everything a hosted trace
UI shows you, from a plain JSONL file, no account or backend:

```bash
logquill trace run-4f2a --file logs.jsonl
```

```text
└─ [INFO] run (812.5ms, 1540→412 tok, $0.0187)
├─ [INFO] plan the work
├─ [INFO] step (250.0ms, 1200→340 tok, $0.0123)
│ └─ [INFO] chat (1200→340 tok, $0.0123)
└─ [INFO] look it up (40.0ms)
```

It **streams** the file line by line — reconstructing one run out of a
multi-gigabyte log file costs memory proportional to that run, not the file
(tested at gigabyte scale in `benchmarks/test_trace_memory.py`). `--json`
prints the same tree as nested `{record, rollup, children}` objects instead,
for feeding into another tool.

### `logquill serve` — a local, offline trace viewer

`logquill serve --file logs.jsonl` runs a small web UI, entirely on the
stdlib (`http.server`) — no new dependency, no account, and nothing leaves
your machine:

```bash
logquill serve --file logs.jsonl
# logquill serve: 12 run(s) found in logs.jsonl
# logquill serve: listening on http://127.0.0.1:52341/ — Ctrl+C to stop
```

Open the printed URL: a run list on the left (record count, token/cost
totals, an error badge), and a combined span-tree/waterfall view for
whichever run you click — indentation shows nesting, bar position and width
show timing. Search and the level filter both work inside a selected run
(client-side, instant) and, with nothing selected, across every run at once
(via `/api/search`, still streamed rather than loaded into memory).

`--db logs.sqlite` reads from a `SQLiteTransport`-written database instead of
a JSONL file. One limitation, inherent to that transport's fixed table
schema: it doesn't store the `llm` block, so runs served from SQLite show no
token/cost annotations even if the original records had them — trace from
the JSONL file (or a transport that does keep `llm`) to see those.

`logquill serve` computes the run list once at startup by streaming through
the source; it's a snapshot, not a live tail — restart it to pick up runs
logged after it started.

### `logquill dev` — watch an agent run live

`logquill dev logs.jsonl` follows a log file like `tail -f`, but redraws the
current run's span tree — colorized, screen cleared between redraws on a
terminal — every time a new record for it arrives, instead of printing flat
lines:

```bash
logquill dev logs.jsonl
```

It tracks whichever run's records have arrived most recently by default, so
the view follows an agent from one run to the next without restarting; pass
`--run-id` to pin it to one run instead. `--backlog N` (default 5000) caps
how much of the file's *existing* content seeds the very first render, so
pointing it at a large pre-existing file doesn't stall before the first draw
— once running, though, a `dev` session keeps its own growing record list in
memory for as long as it runs, unlike `trace`'s bounded streaming.

## API reference

Every public class and function is documented with a docstring; the full
Expand Down
133 changes: 133 additions & 0 deletions benchmarks/test_trace_memory.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,133 @@
"""The exit criterion for the local-first trace viewer: reconstructing one
run's span tree out of a multi-gigabyte log file costs memory proportional
to that run, not the file. Slow and memory-instrumented on purpose, so it
lives here rather than in the default `pytest` run — see `measure.py`'s
module docstring for why these are a separate CI job.
"""

from __future__ import annotations

import json
import tracemalloc
from pathlib import Path
from typing import Any, Iterator

from logquill.trace_tree import build_trace

#: Comfortably over 1 GB — the scale this test is meant to prove, without
#: depending on exactly hitting it.
TARGET_BYTES = 1_100_000_000

#: However large the file gets, reconstructing one small run out of it must
#: stay nowhere near proportional to the file — a few MB, not gigabytes.
MEMORY_BUDGET_BYTES = 100_000_000

NEEDLE_RUN_ID = "needle-run"

_NOISE_META = {"user_id": 42, "route": "/checkout", "tags": ["a", "b", "c"], "note": "x" * 40}


def _record(*, run_id: str, message: str, **meta: Any) -> dict[str, Any]:
return {
"schema_version": "2.0",
"timestamp": "2026-01-01T00:00:00.000Z",
"level": "INFO",
"logger": "app.agent",
"message": message,
"meta": {"run_id": run_id, **meta},
}


def _needle_run_records() -> list[dict[str, Any]]:
"""A small, ordinary agent run — this is what the test must be able to
reconstruct out of the noise around it."""
span_id, step_id = "a" * 16, "b" * 16
return [
_record(
run_id=NEEDLE_RUN_ID,
message="thought",
kind="thought",
parent_span_id=span_id,
),
{
**_record(run_id=NEEDLE_RUN_ID, message="chat", kind="action", parent_span_id=step_id),
"llm": {"model": "m", "tokens_in": 10, "tokens_out": 5, "cost_usd": 0.01},
},
_record(
run_id=NEEDLE_RUN_ID,
message="step",
kind="span",
span_id=step_id,
parent_span_id=span_id,
duration_ms=5.0,
),
_record(
run_id=NEEDLE_RUN_ID,
message="run",
kind="span",
span_id=span_id,
operation="invoke_agent",
agent_name="planner",
duration_ms=12.0,
),
]


def _write_large_jsonl(path: Path, *, target_bytes: int) -> int:
"""Writes `target_bytes`+ of noise from many unrelated runs, with the
needle run's records inserted partway through — returns how many needle
records were written."""
needle = [json.dumps(r, separators=(",", ":")) for r in _needle_run_records()]
noise_run_ids = [f"noise-{i}" for i in range(1000)]
written = 0
needle_written = False
with path.open("w", encoding="utf-8") as f:
i = 0
while written < target_bytes:
if not needle_written and written > target_bytes // 2:
for line in needle:
f.write(line + "\n")
written += len(line) + 1
needle_written = True
line = json.dumps(
_record(
run_id=noise_run_ids[i % len(noise_run_ids)], message="noise", **_NOISE_META
),
separators=(",", ":"),
)
f.write(line + "\n")
written += len(line) + 1
i += 1
return len(needle)


def _read_jsonl(path: Path) -> Iterator[dict[str, Any]]:
with path.open("r", encoding="utf-8") as f:
for line in f:
yield json.loads(line)


def test_a_gigabyte_scale_log_file_is_traced_with_bounded_memory(tmp_path: Path) -> None:
path = tmp_path / "huge.jsonl"
needle_count = _write_large_jsonl(path, target_bytes=TARGET_BYTES)
file_size = path.stat().st_size
assert file_size >= TARGET_BYTES # the scale this test is actually about

tracemalloc.start()
try:
baseline, _ = tracemalloc.get_traced_memory()
builder = build_trace(_read_jsonl(path), NEEDLE_RUN_ID)
_current, peak = tracemalloc.get_traced_memory()
finally:
tracemalloc.stop()

assert builder.record_count == needle_count
(root,) = builder.roots
assert root.record["message"] == "run"
assert root.rollup().tokens_in == 10

peak_bytes = peak - baseline
assert peak_bytes < MEMORY_BUDGET_BYTES, (
f"reconstructing one run used {peak_bytes:,} bytes against a "
f"{file_size:,}-byte file — memory grew with the file, not the run"
)
Loading
Loading