Create a vLLM Managed Inference Job in the web interface
Serve a language model behind an OpenAI-compatible endpoint by creating a vLLM Managed Inference Job in the CosmicAC web interface.
Create a vLLM Managed Inference Job to serve a language model behind an OpenAI-compatible chat endpoint. The form has six sections, and you click Continue to move from one section to the next. For a description of each field in the form, see vLLM Managed Inference Job configuration.
Changing some of the prefilled serving values can cause the deployment to fail. These values come from the model's recommended configuration, which holds the settings that the model needs to deploy.
CosmicAC provides recommended configurations for a set of models. For the list, see Recommended configuration values. For other models, register a vLLM model to create a recommended configuration from its vLLM recipe, or add a recommended configuration with your own values.
Prerequisites
Before you start, make sure that you have the following.
- A running CosmicAC deployment. See Set up CosmicAC.
- Access to the CosmicAC web interface.
Steps
Open the new job form
In the left navigation, click Jobs, and then click New Job.
Select the job type
In the What kind of job? section, select Managed Inference, and then click Continue.
Enter the basics
In the Basics section, enter a Job name. Use lowercase letters and hyphens.
In Tags, add at least one tag. To add a tag, type it, and then press Enter or type a comma.
Click Continue.
Select a model
In the Model to serve section, select a vLLM Model. CosmicAC prefills the Serving configuration from the model's recommended configuration, and labels it Prefilled from with the model name. The configuration has the following fields.
- Runtime image (CUDA): the vLLM serving image.
- Data type: the numeric precision the model runs at.
- Quantisation: how to compress the model weights.
- GPU memory utilisation: the fraction of GPU memory to use, with 85%, 90%, and 95% presets.
- Max model length: the maximum context length, with 8k, 16k, 32k, and 128k presets.
- Max concurrent sequences: the maximum requests handled at once, with 64, 128, 256, and 512 presets.
- Reasoning parser: the parser that separates thinking tokens from the final response.
- Video & image input: whether the model accepts multimodal input, for vision-language models only.
You can edit every field except the runtime image, which comes from the model's recommended configuration. To serve a different image, update the recommended model configuration.
Name the endpoint
Under Endpoint, enter an Endpoint name. The name must be unique across the active Managed Inference Jobs in the deployment. A failed job keeps its name until you delete the job.
Below the name, Will be reachable at shows the endpoint URL.
Set the root disk and environment variables
Under Instance resources, set the Root disk (GB). Select a preset or enter a value. Increase it for a large checkpoint that exceeds the cluster default.
Under Environment variables, review the prefilled variables, and then edit them if needed. To add another, click Add variable, and then enter a Name and Value. To remove one, click the X beside it. For the names CosmicAC reads, see vLLM serving options.
Require an API key
Under API key required, select Require Authorization header, and then click Continue. To create an API key that authenticates requests to the endpoint, see Create an API key in the web interface.
Select the hardware
In the Hardware section, select a Location first. The GPU list stays empty until you select one.
Select a GPU from the ones available in that location. Each card shows the GPU's VRAM, CPU, and RAM. Set the GPU count, and then set the CUDA / driver. Below the count, Recommended for this model shows the model's GPU count. On one node is the most free GPUs on any single node, and In this region is the total free across the location. A GPU count of 16 spreads one replica over two nodes. See Multi-node replicas.
Set Replicas to 1, 2, or 4. CosmicAC doesn't autoscale replicas, and a job keeps the count it starts with. Total shows the GPUs the job claims, the GPU count multiplied by the replica count. For how requests reach the replicas, see Replicas.
Click Continue.
Select the notification events
In the Notifications section, turn on each job lifecycle event you want this job to report. CosmicAC turns all four on by default.
- job.failed: the job transitions to Failed, and the event carries the failure reason.
- job.degraded: healthy replicas drop below the count you set, and the endpoint stays live.
- job.recovered: the job returns to Active from Degraded or Failed.
- job.restart_storm: any replica restarts three times within 10 minutes.
These preferences cover this job alone. An event you turn on here reaches your webhook only if it's also turned on in Settings > Notifications, which also controls the model health and usage window events for the whole deployment. See Set up webhook notifications.
Click Continue.
Review and create the job
In the Review & launch section, check that it reports Ready to create, and then click Create job. If it reports issues instead, click Edit on the row that names the problem, fix it, and then return to this section.
If your configuration departs from the model's recommended parameters, the form shows a Job creation warning that names the values that differ, and Create job stays disabled until you select Create anyway with this configuration. A job created that way can still fail to start.
The model's recommended model configuration sets which parameters warn and in which direction. See Set the job creation warnings.
Open the endpoint
In the Provisioning status dialog, click Open job. When the job is running, click the Endpoint tab to view the endpoint URL. To send it a request, see Connect to a vLLM Managed Inference endpoint.
Help and troubleshooting
Job stuck in Creating or Starting
If a job stays in Creating or Starting, check the status of its KubeVirt virtual machine instance (VMI).
-
Find the job's container ID. Click Jobs in the left navigation, and then click the job. The Containers section lists the Container ID for each container.
-
Find the VMI for the container.
CosmicAC creates one VMI for each container and names it
<container-id>-n0. A multi-node job has one VMI per node.From a machine with
kubectlaccess to your Kubernetes cluster, run the following command.kubectl get vmi -n <namespace>Replace
<namespace>with the namespace configured inK8S_NAMESPACE. -
Check the VMI status.
-
If the VMI is not Running, inspect its events.
kubectl describe vmi <container-id>-n0 -n <namespace> -
If the VMI is Running but the job stays in Creating or Starting, cosmicac-wrk-agent-inference cannot reach cosmicac-wrk-server-k8s-nvidia. These two components connect directly, and some cluster network configurations can block the connection.
To route the connection through a relay, see Set up a relay for CosmicAC.
-
Next steps
- Create a vLLM Managed Inference Job with the CLI
- Create an API key in the web interface
- Connect to a vLLM Managed Inference endpoint
- Create a Parakeet Managed Inference Job in the web interface to serve speech-to-text instead
- vLLM Managed Inference Job configuration
- Recommended configuration values
- What a Managed Inference Job is