Logging and monitoring

The verifier and aggregator export metrics and write logs. This page covers how to collect those metrics, what to watch, and how to read the verifier's startup logs. The policy hook is a service you run, so collect its logs through your own logging stack; the cell reports the policy outcomes it sees as metrics, described below.

Collect the metrics

The verifier and aggregator push metrics over OTLP and expose no scrapable /metrics endpoint, so a Prometheus-style scrape finds nothing. Deploy an OpenTelemetry collector (the upstream OpenTelemetry Collector or Grafana Alloy) to receive the OTLP metrics and forward them to your metrics backend, and set OTEL_SERVICE_NAME on each component so its metrics carry a distinct service name. Then import the kit's Grafana dashboard (the CCV Cell Overview) to visualize them. The off-chain kit documents the collector configuration in Metrics, the metrics to watch and the example dashboard in Monitoring, and its alert catalog in Alerting.

Alert on the policy hook

The RUNBOOK's alert catalog covers the cell's built-in metrics, so follow it for those. It does not cover the policy hook, because the hook is optional, so it is the one alert you add yourself when you run one. The verifier reports each policy outcome on the verifier_message_transitions_total counter with stage="policy":

  • outcome="policy_unavailable" rising means your endpoint is down or timing out and the verifier is retrying.
  • outcome="policy_rejected" counts the messages your endpoint answered FAIL.

Both outcomes leave the message in the verifier's job archive rather than deleting it, with your endpoint's reason string kept for 30 days. The policy hook guide covers listing and replaying archived messages with the verifier's ccv job-queue command.

Read the startup logs

Two log patterns look like faults but are normal, so knowing them saves a false alarm on first deploy.

Right after deploy, the verifier starts before its own aggregator is ready and retries the connection for the first 15 to 30 seconds, logging connection refused and Failed to list message rules until the aggregator comes up. These clear on their own. The catch is that a permanent misconfiguration prints the same lines, so the message alone does not tell the two apart, only how long they last. Treat errors that continue past a minute as a real fault, and check first that the aggregator is running and listening on :::50051.

In steady state with finality_depth: 0, the verifier waits for each source message to reach chain finality before it attests, and logs Healthy while it waits. A log showing only Healthy is a verifier waiting for finality, not a stuck one.

What's next

Get the latest Chainlink content straight to your inbox.