> ## Documentation Index
> Fetch the complete documentation index at: https://docs.eigenpal.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Prometheus metrics

> Scrape queue, throughput, latency, LLM token, and worker process metrics from a self-hosted EigenPal deployment.

Scrape `/metrics` to monitor queues, runs, LLM usage, human review, and worker
health. The endpoint is available only on self-hosted, single-tenant
deployments, and Prometheus authenticates with an API key that has the
**Observability (read-only)** scope.

Each app replica reports metrics for the whole deployment. Configure one scrape
target per deployment; you don't need to scrape workers separately.

## Before you begin

You need the following:

* A Prometheus server with network access to EigenPal.
* The owner or admin role in your installation's organization, the organization
  created when EigenPal was set up.

The endpoint is on by default in self-hosted deployments. Metrics cover every
organization on the deployment, so the endpoint serves them only to API keys
created by an owner or admin of the installation's organization, with the
**Observability (read-only)** or **Full access** scope. Service credentials
and keys from other organizations are refused.

To turn the endpoint off, go to **Settings** > **Monitoring** and turn off
**Metrics endpoint**. While it is off, the endpoint returns `404` for every
request.

## Create a metrics API key

1. In EigenPal, go to **Settings** > **Monitoring** and click **Create an
   observability key**. The form opens with the name `Prometheus` and the
   **Observability (read-only)** access level. You can also go to
   **Developers** > **API keys**, click **Create API key**, and select
   **Observability (read-only)** for **Access**.

2. Click **Create key**, then copy the key. EigenPal shows the key only once.

3. Save the key to a private file:

   ```bash theme={null}
   (umask 077; printf '%s' 'PASTE_KEY_HERE' > eigenpal-metrics.key)
   ```

4. Verify access from a host that can reach EigenPal. Replace
   `eigenpal.example.com` with your deployment's hostname:

   ```bash theme={null}
   curl --fail-with-body --silent --show-error \
     --header "Authorization: Bearer $(cat eigenpal-metrics.key)" \
     https://eigenpal.example.com/metrics
   ```

   A successful request returns HTTP `200` and metrics in text format.

An observability key can read metrics and check its own identity with
`GET /api/v1/auth/check`. Every other API rejects it, so the key can't run
automations or read run data, files, or settings. A **Full access** key can also
read metrics, but use a dedicated observability key for monitoring.

## Connect Prometheus

1. Mount the key file into the Prometheus host or container at
   `/etc/prometheus/secrets/eigenpal-metrics.key`. Make it read-only and readable
   by the Prometheus process. The file must contain only the key, without the
   `Bearer` prefix. Keep it out of source control.

2. Add this job to `scrape_configs` in `prometheus.yml`:

   ```yaml theme={null}
   scrape_configs:
     - job_name: eigenpal
       scheme: https
       metrics_path: /metrics
       scrape_interval: 15s
       scrape_timeout: 10s
       authorization:
         type: Bearer
         credentials_file: /etc/prometheus/secrets/eigenpal-metrics.key
       static_configs:
         - targets: ['eigenpal.example.com:443']
   ```

   Replace the target with your hostname and port, without a scheme or path.
   Prometheus sends `Authorization: Bearer KEY` on each scrape. For more
   options, see the [Prometheus HTTP configuration reference](https://prometheus.io/docs/prometheus/latest/configuration/configuration/#http_config).

   Use HTTPS with certificate verification. For a private certificate authority,
   add `tls_config.ca_file` to the job and mount the CA certificate into
   Prometheus. Restrict endpoint access to your monitoring network.

3. Validate the configuration in the Prometheus environment:

   ```bash theme={null}
   promtool check config /etc/prometheus/prometheus.yml
   ```

   Reload Prometheus with `SIGHUP`, or restart it.

4. After the first scrape, run this query in Prometheus:

   ```promql theme={null}
   up{job="eigenpal"}
   ```

   The result should be `1`. A value of `0` means the scrape failed; check the
   target's last error. No result means Prometheus hasn't scraped the job yet or
   hasn't loaded the configuration.

### Choose one scrape target

Use one stable ingress, load balancer, or Service address per deployment.
Scraping every app pod duplicates deployment metrics and inflates sums. This
also applies to a `PodMonitor` or `ServiceMonitor` that discovers every app pod.

Use a 15-second scrape interval. Workers report their process metrics every 15
seconds, so a longer interval skips some reports and can miss a short event loop
stall. Don't go below 15 seconds: each app replica caches its snapshot for 10
seconds. Keep the scrape timeout shorter than the interval.

If you use Grafana, set the Prometheus data source's **Scrape interval** to the
same value so `$__rate_interval` picks the right window.

### Rotate the key

Create a new observability key, replace the Prometheus key file, and reload
Prometheus. After `up{job="eigenpal"}` returns `1`, revoke the old key in
**Developers** > **API keys**. Both keys work until you revoke the old one, so
scrapes don't fail during rotation.

## Response formats

The endpoint returns OpenMetrics 1.0 when requested in the `Accept` header;
otherwise, it returns Prometheus text format 0.0.4. Prometheus negotiates the
format automatically.

If metrics collection fails, the endpoint returns `503 Service Unavailable`
instead of partial data. Prometheus records `up == 0` for the target.

## How values are collected

* **Deployment counters and histograms** (`_total`, `_seconds_bucket`) are maintained by the
  database as runs, steps, LLM calls, and reviews change state. They don't reset
  when you deploy or restart services, and data retention doesn't decrease them,
  so `rate()` and `increase()` work as expected. Counting starts when you upgrade
  to the release that introduced these metrics.
* **Queue gauges** are read when the cached snapshot refreshes.
* **Worker process gauges** come from a heartbeat that each worker publishes
  every 15 seconds.

Labels that come from your configuration (`model`, `provider`, and `step_type`)
are capped at 128 characters. Each database-backed metric holds at most 500 label combinations;
after that, new combinations report those labels as `other`.

## Metrics reference

All metric names start with `eigenpal_` and use base units: seconds and bytes.
The tables list labels emitted by EigenPal. Prometheus adds `job` and `instance`
labels to each series. Worker CPU counters reset when the worker process restarts.

### Queue

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_queue_jobs` | Gauge | `queue`, `state` | Runs that haven't finished. `queue` is `workflow` or `agent`. `state` is `created`, `scheduled` (waiting for a retry or a delayed start), `pending` (ready to run), `running`, `waiting` (paused, for example on human review), or `finalizing`. |
| `eigenpal_queue_oldest_pending_job_age_seconds` | Gauge | `queue` | How long the oldest ready run has waited for a worker, including any time it spent in `created` while its files uploaded. `0` when no run is ready. |
| `eigenpal_queue_expired_lease_jobs` | Gauge | `queue` | Runs whose worker stopped renewing its lease, usually because the worker crashed. EigenPal retries them automatically. |
| `eigenpal_queue_busy_slots` | Gauge | `queue` | Worker slots that are running a run of this queue. |

### Runs and steps

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_executions_started_total` | Counter | `type` | Runs that a worker picked up for the first time. Retries don't count again. |
| `eigenpal_executions_finished_total` | Counter | `type`, `status` | Runs that finished. `status` is `completed`, `failed`, `cancelled`, or `rejected`. |
| `eigenpal_execution_queue_wait_seconds` | Histogram | `type` | Time from when a run became ready to when a worker picked it up. Includes time a run spends in `created` while its files upload. |
| `eigenpal_execution_duration_seconds` | Histogram | `type`, `status` | Time from pickup to finish, including retries and time paused on human review. |
| `eigenpal_step_duration_seconds` | Histogram | `step_type`, `outcome` | Duration of each workflow step attempt. `outcome` is `ok` or `error`. |

`type` is `workflow` or `agent`.

### LLM and OCR

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_llm_tokens_total` | Counter | `source`, `provider`, `model`, `token_type` | LLM tokens consumed. `source` is `workflow` or `agent`. `token_type` is `input` (uncached input), `cache_read`, `cache_write`, or `output`. |
| `eigenpal_llm_calls_total` | Counter | `source`, `provider`, `model`, `outcome` | LLM calls made by workflow steps. Automatic retries of one call count once. |
| `eigenpal_llm_call_duration_seconds` | Histogram | `source`, `provider`, `model` | Duration of LLM calls made by workflow steps, including retries. |
| `eigenpal_ocr_calls_total` | Counter | `provider`, `model`, `outcome` | OCR requests made while parsing documents. |
| `eigenpal_ocr_pages_total` | Counter | `provider`, `model` | Pages processed by OCR. |

Token counts match the **Usage** page. On that page, input tokens include cache
reads and writes; here, total input is `input + cache_read + cache_write`. Agent
runs report their tokens when the run reports its totals, not after each model
call.

### Human review

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_human_review_pending_tasks` | Gauge | `source_kind` | Review tasks waiting for a reviewer. `source_kind` is `workflow_step` or `agent_tool`. |
| `eigenpal_human_review_oldest_pending_task_age_seconds` | Gauge | `source_kind` | Age of the oldest waiting review task. `0` when none are waiting. |
| `eigenpal_human_review_pending_required_fields` | Gauge | `source_kind` | Fields that waiting tasks ask a reviewer to confirm. |
| `eigenpal_human_review_pending_confirmed_fields` | Gauge | `source_kind` | Fields that reviewers have already confirmed on waiting tasks. |
| `eigenpal_human_review_tasks_finished_total` | Counter | `source_kind`, `outcome` | Review tasks that were `approved`, `rejected`, or `cancelled`. |
| `eigenpal_human_review_task_duration_seconds` | Histogram | `source_kind`, `outcome` | Time from task creation to its outcome. |

### Worker processes

Each worker process has a unique `worker` label that changes on restart. After
45 seconds without a heartbeat, `eigenpal_worker_up` becomes `0`. Unresponsive
workers remain listed for 10 minutes; workers that shut down cleanly disappear
immediately. Offline workers report only identity, uptime status, heartbeat age,
and start time.

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_workers_online` | Gauge | None | Worker processes that sent a heartbeat in the last 45 seconds. |
| `eigenpal_worker_up` | Gauge | `worker` | `1` while the worker sends heartbeats, `0` after it stops. |
| `eigenpal_worker_info` | Gauge | `worker`, `hostname`, `version` | Identity of the worker process. Always `1`. |
| `eigenpal_worker_heartbeat_age_seconds` | Gauge | `worker` | Time since the worker last published its metrics. |
| `eigenpal_worker_start_time_seconds` | Gauge | `worker` | Start time of the process, in seconds since the Unix epoch. |
| `eigenpal_worker_slots` | Gauge | `worker` | Runs the worker can process at once (`WORKER_POOL_CONCURRENCY`). |
| `eigenpal_worker_busy_slots` | Gauge | `worker` | Slots that are running a run. |
| `eigenpal_worker_cpu_seconds_total` | Counter | `worker`, `mode` | CPU time used by the process. `mode` is `user` or `system`. `rate()` gives the CPU cores in use. |
| `eigenpal_worker_resident_memory_bytes` | Gauge | `worker` | Resident memory of the process. |
| `eigenpal_worker_heap_used_bytes` | Gauge | `worker` | JavaScript heap in use. |
| `eigenpal_worker_heap_total_bytes` | Gauge | `worker` | JavaScript heap reserved. |
| `eigenpal_worker_external_memory_bytes` | Gauge | `worker` | Memory held outside the JavaScript heap, such as buffers. |
| `eigenpal_worker_event_loop_lag_mean_seconds` | Gauge | `worker` | Average event loop delay over the last 15 seconds. |
| `eigenpal_worker_event_loop_lag_p99_seconds` | Gauge | `worker` | 99th percentile event loop delay over the last 15 seconds. |
| `eigenpal_worker_event_loop_lag_max_seconds` | Gauge | `worker` | Longest event loop stall over the last 15 seconds. |

Event loop lag measures timer delays. High lag can delay lease renewals and
heartbeats. Use the maximum to detect individual stalls that the mean or 99th
percentile can hide.

### Endpoint

| Metric | Type | Labels | Description |
| - | - | - | - |
| `eigenpal_build_info` | Gauge | `version`, `revision` | Release of the app serving the endpoint. Always `1`. |
| `eigenpal_metrics_collection_duration_seconds` | Gauge | None | Time the last snapshot query took. |

Use the `up` series that Prometheus records for each target to check endpoint
health.

### Histogram buckets

| Histogram | Buckets (seconds) |
| - | - |
| `eigenpal_execution_queue_wait_seconds` | 0.1, 0.5, 1, 2.5, 5, 10, 30, 60, 120, 300, 600, 1800, 3600 |
| `eigenpal_execution_duration_seconds` | 1, 5, 10, 30, 60, 120, 300, 600, 1200, 1800, 3600, 7200, 21600 |
| `eigenpal_step_duration_seconds` | 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, 60, 120, 300, 600, 1800 |
| `eigenpal_llm_call_duration_seconds` | 0.5, 1, 2.5, 5, 10, 20, 30, 60, 120, 300, 600 |
| `eigenpal_human_review_task_duration_seconds` | 60, 300, 900, 1800, 3600, 7200, 14400, 28800, 86400, 259200, 604800 |

## Example queries

These queries use the `eigenpal` job configured above. Apply `rate()` or
`increase()` to each counter series before aggregating so Prometheus can detect
counter resets. Use `histogram_quantile()` for histogram percentiles.
See [Prometheus query functions](https://prometheus.io/docs/prometheus/latest/querying/functions/).

```promql theme={null}
# Tokens per second, by model and token type
sum by (model, token_type) (rate(eigenpal_llm_tokens_total{job="eigenpal"}[5m]))

# Total tokens in the last 24 hours, by source and model
sum by (source, model) (increase(eigenpal_llm_tokens_total{job="eigenpal"}[24h]))

# Average run duration
sum by (type) (rate(eigenpal_execution_duration_seconds_sum{job="eigenpal"}[5m]))
  / sum by (type) (rate(eigenpal_execution_duration_seconds_count{job="eigenpal"}[5m]))

# 95th percentile run duration
histogram_quantile(0.95, sum by (le, type) (rate(eigenpal_execution_duration_seconds_bucket{job="eigenpal"}[5m])))

# Failure ratio
sum by (type) (rate(eigenpal_executions_finished_total{job="eigenpal",status="failed"}[15m]))
  / sum by (type) (rate(eigenpal_executions_finished_total{job="eigenpal"}[15m]))

# Ready runs waiting for a worker
eigenpal_queue_jobs{job="eigenpal",state="pending"}

# Worker slot utilization
sum(eigenpal_worker_busy_slots{job="eigenpal"}) / sum(eigenpal_worker_slots{job="eigenpal"})

# CPU cores used by each worker
sum by (worker) (rate(eigenpal_worker_cpu_seconds_total{job="eigenpal"}[5m]))
```

## Example alerts

Save these rules as `eigenpal-alerts.yml`, add the file to `rule_files` in
`prometheus.yml`, and reload Prometheus. Validate the file first:

```bash theme={null}
promtool check rules eigenpal-alerts.yml
```

Adjust thresholds and `for` durations to your workload. Configure Alertmanager
to route notifications. See [Prometheus alerting practices](https://prometheus.io/docs/practices/alerting/).

```yaml theme={null}
groups:
  - name: eigenpal
    rules:
      - alert: EigenPalMetricsDown
        expr: up{job="eigenpal"} == 0
        for: 5m
        annotations:
          summary: EigenPal metrics endpoint is unreachable or rejecting scrapes
      - alert: EigenPalNoWorkers
        expr: eigenpal_workers_online{job="eigenpal"} == 0
        for: 2m
        annotations:
          summary: EigenPal has no online workers
      - alert: EigenPalWorkerDown
        expr: eigenpal_worker_up{job="eigenpal"} == 0
        for: 1m
        annotations:
          summary: EigenPal worker stopped sending heartbeats
      - alert: EigenPalQueueBacklog
        expr: max by (queue) (eigenpal_queue_oldest_pending_job_age_seconds{job="eigenpal"}) > 300
        for: 5m
        annotations:
          summary: EigenPal ready runs have waited more than 5 minutes
      - alert: EigenPalHighFailureRate
        expr: |
          sum by (type) (rate(eigenpal_executions_finished_total{job="eigenpal",status="failed"}[15m]))
            / sum by (type) (rate(eigenpal_executions_finished_total{job="eigenpal"}[15m])) > 0.2
        for: 15m
        annotations:
          summary: EigenPal run failure ratio exceeds 20%
      - alert: EigenPalWorkerEventLoopBlocked
        expr: max_over_time(eigenpal_worker_event_loop_lag_max_seconds{job="eigenpal"}[5m]) > 1
        for: 10m
        annotations:
          summary: EigenPal worker event loop stalls exceed 1 second
      - alert: EigenPalWorkerMemoryHigh
        expr: eigenpal_worker_resident_memory_bytes{job="eigenpal"} > 3e9
        for: 15m
        annotations:
          summary: EigenPal worker resident memory exceeds 3 GB
```

## Troubleshoot

| Symptom | Cause and fix |
| - | - |
| `404 Not Found` | The endpoint is turned off in **Settings** > **Monitoring**, or the deployment runs in multi-tenant mode. The endpoint is available only on self-hosted, single-tenant deployments. |
| `401 Unauthorized` | The scraper sends no key, or the key is revoked or expired. Check the key file and make sure any proxy forwards `Authorization`. |
| `403 Forbidden` | The key lacks the **Observability (read-only)** or **Full access** scope, belongs to another organization, or its creator isn't an owner or admin of the installation's organization. |
| `503` and `up == 0` | The database didn't respond while EigenPal checked the key or collected metrics. Check the app logs for `Metrics collection failed` or `Metrics scrape key lookup failed`. |
| TLS certificate error | Use a hostname covered by the certificate and configure `tls_config.ca_file` for a private CA. |
| Key file permission error | Check that the file is mounted at `credentials_file` inside Prometheus and readable by its process. |
| HTML parse error in Prometheus | The scraper reached a dashboard page. Set `metrics_path` to `/metrics`. |
| Values are a multiple of the expected values | Prometheus scrapes each app pod. Scrape one stable address instead. |
| `eigenpal_workers_online` is `0` | No worker is running, or workers can't write to the database. Check the worker logs for `Worker heartbeat failed`. |
| No `eigenpal_llm_*` series | No LLM calls have run since the upgrade. The series appear after the first call. |


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.