Check the health of a Managed Inference endpoint
Check the status, success rate, and latency of a Managed Inference endpoint from the web interface or the CLI.
Check the status, success rate, and latency of your Managed Inference endpoints, and probe an endpoint when you need new results.
Prerequisites
Before you start, make sure that you have the following.
- A running Managed Inference Job with a serving endpoint. See Create a vLLM Managed Inference Job with the CLI.
- Access to the CosmicAC web interface, for the web interface method.
- The CosmicAC CLI installed and configured, for the CLI method. See Install the CLI.
Steps
Open the health view
CosmicAC probes the replicas every five minutes by default, and both methods show the result of the latest health check.
In the left navigation, click Model Health. While the page is open, it refreshes the results every 30 seconds.
To change the time range, select 1H, 6H, 24H, 7D, or 30D at the top of the page. The default is 1H.
The time range changes the success rate, traffic, failures, and average response. It doesn't change the status, which CosmicAC always calculates over the last 24 hours.
Read the results
Each endpoint reports the following values.
- Status: Healthy, Degraded, or Down. For what each value means, see Model health.
- Success rate: the percentage of requests that succeeded.
- Traffic: the total number of requests that the endpoint handled.
- Failures: the number of those requests that failed.
- Avg response: the average response time in milliseconds. The CLI shows it as Latency (avg).
- Last health check: when CosmicAC last probed the endpoint. The CLI shows it as Last Check At.
- Last updated: when CosmicAC generated the results. The CLI shows it as Timestamp.
- Last Check: whether the last probe succeeded. Only the CLI shows this value.
- Last Check Latency: how long the endpoint took to answer the last probe. Only the CLI shows this value.
If the endpoint handled no requests in the time range, the success rate and the average response are empty. The web interface shows a dash, and the CLI shows N/A.
To find an unhealthy replica, check the replica details.
- Web interface: each endpoint card shows how many replicas are healthy. Below the count, the card lists each unhealthy replica by ID. A degraded replica also shows its success rate, traffic, failures, average response, and last check. A down replica shows Down · out of rotation.
- CLI: the Replicas table lists every replica, with its ID, status, traffic, failures, and average latency.
Probe an endpoint on demand
To get new results before the next scheduled probe, probe the endpoint yourself.
CosmicAC probes a replica by sending it a real inference request, so a probe uses serving capacity. Probing every endpoint sends a request to every replica that you run. On a large deployment, probe one endpoint at a time.
To probe every endpoint, click Run health check at the top of the Model Health page. To probe one endpoint, click Health check on its card.
The results don't change right away. They update on the next refresh, within 30 seconds.