CosmicAC Logo
Platform management

Connect CosmicAC to your Prometheus and Loki

Connect CosmicAC to your Prometheus and Loki so job pages show charts and log history.

Connect CosmicAC to your Prometheus and Loki so job pages show charts and log history. Job metrics go to Prometheus, and job logs go to Loki.

You run and maintain both, including their storage, retention, and access control. CosmicAC doesn't install or operate either one. To install them, see the Prometheus installation guide and the Loki installation guide.

Prerequisites

Before you start, make sure that you have the following.

  • A running CosmicAC deployment. See Set up CosmicAC.
  • The platform administrator role. See Teams and roles.
  • A running Prometheus and a running Loki.
  • The host and port that cosmicac-wrk-monitor serves on. It listens on port 9110 unless your deployment sets another.
  • Telemetry switched on. A service sends its logs and metrics to cosmicac-wrk-monitor only after you turn on that service's producer. See Telemetry.

Allow the following network paths between your Prometheus, your Loki, and cosmicac-wrk-monitor.

AllowWithout it
Your Prometheus to reach cosmicac-wrk-monitor on port 9110Prometheus scrapes nothing, and every chart stays empty
cosmicac-wrk-monitor to reach your LokiLog history stays empty, though live tail still works
cosmicac-wrk-monitor to reach your PrometheusThe charts on a job page stay empty, even while Prometheus scrapes normally

CosmicAC connects to your Prometheus and Loki at the URLs that you save in the following steps. If your Prometheus or Loki uses a port, include it in the URL.

Steps

Open the Observability page

In the left navigation, click Settings, and then under Instance, click Observability.

Connect Prometheus

In the Prometheus query URL field, enter the base URL of your Prometheus, such as http://prometheus:9090, and then click Save.

CosmicAC reads from this URL to draw the charts on a job page. It never writes to your Prometheus. Add CosmicAC as a Prometheus scrape target sets up the opposite direction, where your Prometheus scrapes metrics from CosmicAC.

Connect Loki

In the Loki query URL field, enter the base URL of your Loki, such as http://loki:3100, and then click Save.

CosmicAC doesn't send Loki's X-Scope-OrgID tenant header. On a single-tenant Loki, set auth_enabled: false. On a multi-tenant Loki, add a proxy in front that sets the header.

Add CosmicAC as a Prometheus scrape target

In your Prometheus configuration, add cosmicac-wrk-monitor as a scrape target.

scrape_configs:
  - job_name: cosmicac-jobs
    scrape_interval: 5s
    honor_labels: true
    static_configs:
      - targets: ['<monitor-host>:9110']

A scrape that omits a required token fails with 401.

honor_labels must be on

Without honor_labels: true, Prometheus overwrites the job ID in the instance label, so the CPU and memory charts show no data. The GPU charts still work because GPU series use the job_id label.

A scrape stores one value per series for each interval, so scrape_interval controls how much detail the charts keep. Job agents sample about once a second, so a scrape_interval of 5s keeps about one sample in every five. A lower interval keeps more samples and stores more data.

Confirm the connection

A card shows Connected as soon as you save a URL. That status doesn't confirm that CosmicAC reached Prometheus or Loki.

To confirm that Prometheus scrapes CosmicAC, click Status > Targets in the Prometheus web interface. The cosmicac-jobs target reports UP.

To confirm that CosmicAC queries Prometheus, request the metrics for a running job.

curl http://<monitor-host>:9110/job-metrics/<job-id>

The response carries cpu and gpu values. Wait about 5 seconds after a job starts so that Prometheus scrapes at least once.

To confirm that CosmicAC reads your Loki, request the stored logs for the same job.

curl "http://<monitor-host>:9110/logs/history?job_id=<job-id>"

The response carries log lines. An empty response with a running job that has already produced logs means CosmicAC isn't reaching your Loki. Live tail keeps working in that case because it doesn't read from Loki. See Get log history.

Restrict access to cosmicac-wrk-monitor

Only the /metrics and /metrics/<endpoint-name> endpoints on cosmicac-wrk-monitor can require a token. The other endpoints, including /logs, /logs/history, /logs/export, and /job-metrics, require no authentication.

cosmicac-wrk-monitor is reachable at the following addresses.

  • Port 9110 on the host machine.
  • The /monitor/ path on the web interface address, such as https://<base-url>/monitor/.

Anyone who can reach either address can read the logs and metrics of every job. Allow only trusted networks to reach both addresses, and add authentication before you expose them.

To require a token on /metrics, do the following.

  1. In .env, set the token.

    MONITOR_COMMON_CONFIG__metricsScrapeToken=<scrape-token>
  2. Apply the change.

    task apply-wrk-monitor-common-config
  3. Restart cosmicac-wrk-monitor.

    task restart SERVICES="cosmicac-wrk-monitor"
  4. In your Prometheus configuration, use the With a scrape token tab in Add CosmicAC as a Prometheus scrape target.

Disconnect Prometheus or Loki

To disconnect one, click Disconnect on its card, and then confirm. CosmicAC clears that URL and leaves the other in place.

To disconnect both, click Disconnect all at the top of the page, and then confirm.

Disconnecting Loki stops CosmicAC from sending new log lines to Loki and from reading log history. Disconnecting Prometheus doesn't stop your Prometheus from scraping. To stop the scrape, remove the target from your Prometheus configuration.

Help and troubleshooting

The target reports UP but no job series arrive

Run a job. Series exist only while a job runs, and they leave the target about a minute after it ends.

GPU charts populate but CPU and memory charts stay empty

Add honor_labels: true to your scrape job, and then reload Prometheus.

The charts show fewer points than you expect

Lower scrape_interval. Prometheus stores one value per series per scrape.

The charts fail with a 503 ERR_PROMETHEUS_NOT_CONFIGURED error

Save a Prometheus query URL in step 2, and allow cosmicac-wrk-monitor to reach your Prometheus.

The charts fail with a 502 ERR_PROMETHEUS_UNAVAILABLE error

CosmicAC couldn't reach your Prometheus, or your Prometheus returned a server error. Check that the saved URL is the base URL of your Prometheus, and that cosmicac-wrk-monitor can reach it.

The charts fail with a 400 ERR_PROMETHEUS_QUERY_REJECTED error

Your Prometheus rejected the query. The response includes the error message from Prometheus. Check that the saved URL is the base URL of your Prometheus.

Log history fails with a 503 ERR_LOKI_NOT_CONFIGURED error

Save a Loki query URL in step 3. Live tail keeps working without it.

Log history fails with a 400 ERR_LOKI_QUERY_REJECTED error

Narrow the time range, or raise max_query_length in your Loki configuration.

Log history fails with a 502 ERR_LOKI_UNAVAILABLE error

Check that the saved URL is the root that serves /loki/api/v1/. If Loki sits behind a proxy that adds a path prefix, include the prefix.

Scraping fails with a 401 error

Use the With a scrape token configuration when your deployment requires a scrape token.

A finished job's charts are empty

Query Prometheus directly over a time range that covers when the job ran.

Next steps

On this page