Export logs-collector and pings-collector metrics to Prometheus - #80
Merged
Merged
Conversation
Serve /metrics on PROMETHEUS_PORT (default 9090) with the essential collection metrics: workers and backlogged workers, log requests by result, logs stored, logs dropped by reason, and storage errors. Every series carries the instance's shard label. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Serve /metrics on PROMETHEUS_PORT (default 9090) with the metrics that correspond to logs-collector's: workers, heartbeat requests by result, heartbeats stored, heartbeats dropped by reason, and storage errors, each with the instance's shard label. Move the /metrics server into collector-utils so both collectors share it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Encoding into a String can't fail, so expect it instead of mapping the error to a 500. Keep collector-utils' dependencies sorted, and start pings-collector's metrics server right before the server, as in logs-collector. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Add the metrics the review asked for: log lag per stored row (with a
bucket at the 1200 s reward cutoff), buffer fill and limit, round and
request durations, and shard_info with total_shards.
Fix the existing ones: split logs_dropped into logs_discarded{reason}
(lost) and logs_deferred (collected again later), label request errors
with a fixed reason taken from the transport error's variant, label
storage errors with the operation, and create every known series up
front so rates read 0 instead of no data. Rename pings' heartbeats_dropped
to heartbeats_discarded to match.
The transport doesn't export FetchLogsError or RequestError, so the
reason is read from the start of the error's Debug form.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
define-null
force-pushed
the
logs-collector-metrics
branch
from
September 24, 2026 10:45
a9095d0 to
1247a2f
Compare
Move workers, requests, request and round durations, shard_info and the error reason lookup into collector_utils::CollectorMetrics, tested once there. observe_request wraps the request future, so call sites keep their original match arms. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds Prometheus metrics to
logs-collectorandpings-collector, served at/metricsonPROMETHEUS_PORT(default9090). Every series carries the instance'sshard. Known label values are created at startup, so rates read 0 instead of no data.Metrics
Both collectors (
logs_collector_*,pings_collector_*):shard_info{total_shards},workersrequests_total{result,reason},request_duration_seconds{result},round_duration_secondsstorage_errors_total(withoperation="read"|"insert"in logs-collector)logs-collector:log_lag_seconds: collector minus worker timestamp, per stored row. Share collected too late for rewards:1 - rate(..._bucket{le="1200.0"}) / rate(..._count)buffer_bytes,buffer_max_bytes,backlogged_workerslogs_stored_total,logs_discarded_total{reason}(lost),logs_deferred_total(buffer full, collected again later)pings-collector:heartbeats_stored_total,heartbeats_discarded_total{reason}Notes
reasonis taken from the variant name at the start of the error'sDebugform, orotherif it isn't recognised.logs-collectorto 2.4.0 andpings-collectorto 2.9.0.metricscontainer port and aPodMonitorper collector in the infra chart.🤖 Generated with Claude Code