/metrics to monitor queues, runs, LLM usage, human review, and worker
health. The endpoint is available only on self-hosted, single-tenant
deployments, and Prometheus authenticates with an API key that has the
Observability (read-only) scope.
Each app replica reports metrics for the whole deployment. Configure one scrape
target per deployment; you don’t need to scrape workers separately.
Before you begin
You need the following:- A Prometheus server with network access to EigenPal.
- The owner or admin role in your installation’s organization, the organization created when EigenPal was set up.
404 for every
request.
Create a metrics API key
-
In EigenPal, go to Settings > Monitoring and click Create an
observability key. The form opens with the name
Prometheusand the Observability (read-only) access level. You can also go to Developers > API keys, click Create API key, and select Observability (read-only) for Access. - Click Create key, then copy the key. EigenPal shows the key only once.
-
Save the key to a private file:
-
Verify access from a host that can reach EigenPal. Replace
eigenpal.example.comwith your deployment’s hostname:A successful request returns HTTP200and metrics in text format.
GET /api/v1/auth/check. Every other API rejects it, so the key can’t run
automations or read run data, files, or settings. A Full access key can also
read metrics, but use a dedicated observability key for monitoring.
Connect Prometheus
-
Mount the key file into the Prometheus host or container at
/etc/prometheus/secrets/eigenpal-metrics.key. Make it read-only and readable by the Prometheus process. The file must contain only the key, without theBearerprefix. Keep it out of source control. -
Add this job to
scrape_configsinprometheus.yml:Replace the target with your hostname and port, without a scheme or path. Prometheus sendsAuthorization: Bearer KEYon each scrape. For more options, see the Prometheus HTTP configuration reference. Use HTTPS with certificate verification. For a private certificate authority, addtls_config.ca_fileto the job and mount the CA certificate into Prometheus. Restrict endpoint access to your monitoring network. -
Validate the configuration in the Prometheus environment:
Reload Prometheus with
SIGHUP, or restart it. -
After the first scrape, run this query in Prometheus:
The result should be
1. A value of0means the scrape failed; check the target’s last error. No result means Prometheus hasn’t scraped the job yet or hasn’t loaded the configuration.
Choose one scrape target
Use one stable ingress, load balancer, or Service address per deployment. Scraping every app pod duplicates deployment metrics and inflates sums. This also applies to aPodMonitor or ServiceMonitor that discovers every app pod.
Use a 15-second scrape interval. Workers report their process metrics every 15
seconds, so a longer interval skips some reports and can miss a short event loop
stall. Don’t go below 15 seconds: each app replica caches its snapshot for 10
seconds. Keep the scrape timeout shorter than the interval.
If you use Grafana, set the Prometheus data source’s Scrape interval to the
same value so $__rate_interval picks the right window.
Rotate the key
Create a new observability key, replace the Prometheus key file, and reload Prometheus. Afterup{job="eigenpal"} returns 1, revoke the old key in
Developers > API keys. Both keys work until you revoke the old one, so
scrapes don’t fail during rotation.
Response formats
The endpoint returns OpenMetrics 1.0 when requested in theAccept header;
otherwise, it returns Prometheus text format 0.0.4. Prometheus negotiates the
format automatically.
If metrics collection fails, the endpoint returns 503 Service Unavailable
instead of partial data. Prometheus records up == 0 for the target.
How values are collected
- Deployment counters and histograms (
_total,_seconds_bucket) are maintained by the database as runs, steps, LLM calls, and reviews change state. They don’t reset when you deploy or restart services, and data retention doesn’t decrease them, sorate()andincrease()work as expected. Counting starts when you upgrade to the release that introduced these metrics. - Queue gauges are read when the cached snapshot refreshes.
- Worker process gauges come from a heartbeat that each worker publishes every 15 seconds.
model, provider, and step_type)
are capped at 128 characters. Each database-backed metric holds at most 500 label combinations;
after that, new combinations report those labels as other.
Metrics reference
All metric names start witheigenpal_ and use base units: seconds and bytes.
The tables list labels emitted by EigenPal. Prometheus adds job and instance
labels to each series. Worker CPU counters reset when the worker process restarts.
Queue
Runs and steps
type is workflow or agent.
LLM and OCR
Token counts match the Usage page. On that page, input tokens include cache
reads and writes; here, total input is
input + cache_read + cache_write. Agent
runs report their tokens when the run reports its totals, not after each model
call.
Human review
Worker processes
Each worker process has a uniqueworker label that changes on restart. After
45 seconds without a heartbeat, eigenpal_worker_up becomes 0. Unresponsive
workers remain listed for 10 minutes; workers that shut down cleanly disappear
immediately. Offline workers report only identity, uptime status, heartbeat age,
and start time.
Event loop lag measures timer delays. High lag can delay lease renewals and
heartbeats. Use the maximum to detect individual stalls that the mean or 99th
percentile can hide.
Endpoint
Use the
up series that Prometheus records for each target to check endpoint
health.
Histogram buckets
Example queries
These queries use theeigenpal job configured above. Apply rate() or
increase() to each counter series before aggregating so Prometheus can detect
counter resets. Use histogram_quantile() for histogram percentiles.
See Prometheus query functions.
Example alerts
Save these rules aseigenpal-alerts.yml, add the file to rule_files in
prometheus.yml, and reload Prometheus. Validate the file first:
for durations to your workload. Configure Alertmanager
to route notifications. See Prometheus alerting practices.