Skip to main content
Scrape /metrics to monitor queues, runs, LLM usage, human review, and worker health. The endpoint is available only on self-hosted, single-tenant deployments, and Prometheus authenticates with an API key that has the Observability (read-only) scope. Each app replica reports metrics for the whole deployment. Configure one scrape target per deployment; you don’t need to scrape workers separately.

Before you begin

You need the following:
  • A Prometheus server with network access to EigenPal.
  • The owner or admin role in your installation’s organization, the organization created when EigenPal was set up.
The endpoint is on by default in self-hosted deployments. Metrics cover every organization on the deployment, so the endpoint serves them only to API keys created by an owner or admin of the installation’s organization, with the Observability (read-only) or Full access scope. Service credentials and keys from other organizations are refused. To turn the endpoint off, go to Settings > Monitoring and turn off Metrics endpoint. While it is off, the endpoint returns 404 for every request.

Create a metrics API key

  1. In EigenPal, go to Settings > Monitoring and click Create an observability key. The form opens with the name Prometheus and the Observability (read-only) access level. You can also go to Developers > API keys, click Create API key, and select Observability (read-only) for Access.
  2. Click Create key, then copy the key. EigenPal shows the key only once.
  3. Save the key to a private file:
  4. Verify access from a host that can reach EigenPal. Replace eigenpal.example.com with your deployment’s hostname:
    A successful request returns HTTP 200 and metrics in text format.
An observability key can read metrics and check its own identity with GET /api/v1/auth/check. Every other API rejects it, so the key can’t run automations or read run data, files, or settings. A Full access key can also read metrics, but use a dedicated observability key for monitoring.

Connect Prometheus

  1. Mount the key file into the Prometheus host or container at /etc/prometheus/secrets/eigenpal-metrics.key. Make it read-only and readable by the Prometheus process. The file must contain only the key, without the Bearer prefix. Keep it out of source control.
  2. Add this job to scrape_configs in prometheus.yml:
    Replace the target with your hostname and port, without a scheme or path. Prometheus sends Authorization: Bearer KEY on each scrape. For more options, see the Prometheus HTTP configuration reference. Use HTTPS with certificate verification. For a private certificate authority, add tls_config.ca_file to the job and mount the CA certificate into Prometheus. Restrict endpoint access to your monitoring network.
  3. Validate the configuration in the Prometheus environment:
    Reload Prometheus with SIGHUP, or restart it.
  4. After the first scrape, run this query in Prometheus:
    The result should be 1. A value of 0 means the scrape failed; check the target’s last error. No result means Prometheus hasn’t scraped the job yet or hasn’t loaded the configuration.

Choose one scrape target

Use one stable ingress, load balancer, or Service address per deployment. Scraping every app pod duplicates deployment metrics and inflates sums. This also applies to a PodMonitor or ServiceMonitor that discovers every app pod. Use a 15-second scrape interval. Workers report their process metrics every 15 seconds, so a longer interval skips some reports and can miss a short event loop stall. Don’t go below 15 seconds: each app replica caches its snapshot for 10 seconds. Keep the scrape timeout shorter than the interval. If you use Grafana, set the Prometheus data source’s Scrape interval to the same value so $__rate_interval picks the right window.

Rotate the key

Create a new observability key, replace the Prometheus key file, and reload Prometheus. After up{job="eigenpal"} returns 1, revoke the old key in Developers > API keys. Both keys work until you revoke the old one, so scrapes don’t fail during rotation.

Response formats

The endpoint returns OpenMetrics 1.0 when requested in the Accept header; otherwise, it returns Prometheus text format 0.0.4. Prometheus negotiates the format automatically. If metrics collection fails, the endpoint returns 503 Service Unavailable instead of partial data. Prometheus records up == 0 for the target.

How values are collected

  • Deployment counters and histograms (_total, _seconds_bucket) are maintained by the database as runs, steps, LLM calls, and reviews change state. They don’t reset when you deploy or restart services, and data retention doesn’t decrease them, so rate() and increase() work as expected. Counting starts when you upgrade to the release that introduced these metrics.
  • Queue gauges are read when the cached snapshot refreshes.
  • Worker process gauges come from a heartbeat that each worker publishes every 15 seconds.
Labels that come from your configuration (model, provider, and step_type) are capped at 128 characters. Each database-backed metric holds at most 500 label combinations; after that, new combinations report those labels as other.

Metrics reference

All metric names start with eigenpal_ and use base units: seconds and bytes. The tables list labels emitted by EigenPal. Prometheus adds job and instance labels to each series. Worker CPU counters reset when the worker process restarts.

Queue

Runs and steps

type is workflow or agent.

LLM and OCR

Token counts match the Usage page. On that page, input tokens include cache reads and writes; here, total input is input + cache_read + cache_write. Agent runs report their tokens when the run reports its totals, not after each model call.

Human review

Worker processes

Each worker process has a unique worker label that changes on restart. After 45 seconds without a heartbeat, eigenpal_worker_up becomes 0. Unresponsive workers remain listed for 10 minutes; workers that shut down cleanly disappear immediately. Offline workers report only identity, uptime status, heartbeat age, and start time. Event loop lag measures timer delays. High lag can delay lease renewals and heartbeats. Use the maximum to detect individual stalls that the mean or 99th percentile can hide.

Endpoint

Use the up series that Prometheus records for each target to check endpoint health.

Histogram buckets

Example queries

These queries use the eigenpal job configured above. Apply rate() or increase() to each counter series before aggregating so Prometheus can detect counter resets. Use histogram_quantile() for histogram percentiles. See Prometheus query functions.

Example alerts

Save these rules as eigenpal-alerts.yml, add the file to rule_files in prometheus.yml, and reload Prometheus. Validate the file first:
Adjust thresholds and for durations to your workload. Configure Alertmanager to route notifications. See Prometheus alerting practices.

Troubleshoot