Welcome
Run GPU Container Jobs and Managed Inference Jobs for language and speech-to-text models on your Kubernetes cluster.
CosmicAC is a self-hosted platform for running GPU workloads on your Kubernetes cluster. You deploy CosmicAC on your host machine, where it connects to your cluster and runs GPU Container Jobs and Managed Inference Jobs for language and speech-to-text models.
Get started
To deploy CosmicAC and create your first job, complete these steps in order.
- Check the requirements.
- Deploy CosmicAC.
- Set up recommended model configurations.
- Install the CLI.
- Create your first job with a GPU Container Job or a vLLM Managed Inference Job.
What you can run
CosmicAC runs two types of jobs. Choose a GPU Container Job to run code on a GPU machine with shell access. Choose a Managed Inference Job to serve a model through an API.
Understand and manage CosmicAC
See how CosmicAC works, then find guides for operating the platform, managing teams and access, and managing models.
Architecture
See how CosmicAC connects to your Kubernetes cluster and runs each job type.
CosmicAC jobs
Understand the job types and the statuses a job moves through.
Recommended model configurations
Manage the default serving values for the models you serve.
Accounts and settings
Manage teams, team members, roles, and node reservations.
Platform management
Upgrade your deployment, manage racks, and configure notifications and observability.
Call a model
Each Managed Inference Job exposes an endpoint for its model. Create an API key, send requests to the endpoint, and check its health.
Create an API key
Authorize requests to a Managed Inference endpoint with the CLI.
Call a vLLM endpoint
Send a chat completion request from the CLI or an OpenAI-compatible client.
Transcribe audio
Send audio to a Parakeet endpoint and get a transcription with the CLI or curl.
Check endpoint health
Check the status, success rate, and latency of a Managed Inference endpoint.
Reference
Look up exact commands, routes, fields, and values, and see what changed in each release.
CLI commands
Every CosmicAC CLI command, with usage, arguments, and options.
Task deployment commands
Commands for deploying, upgrading, and operating the stack.
API reference
HTTP routes for the inference API, monitor API, and observability settings API.
Configuration reference
Deployment environment variables, kubeconfig requirements, and job configuration fields.
Recommended configuration values
Recommended values for a set of vLLM models and the Parakeet model.
Changelog
User-facing changes in each release.