CosmicAC Logo

Connect to a GPU Container Job

Open an interactive shell in a running GPU Container Job with the CLI.

Open a shell in a running GPU Container Job to run commands directly on the container.

Prerequisites

Before you start, make sure that you have the following.

Steps

Find the job ID

List your jobs, and copy the ID of the job you want to connect to.

cosmicac jobs list

Find the container index

View the job details. Replace <jobId> with the job ID.

cosmicac jobs detail <jobId>

In the Containers section, each container has an index that starts at 0, such as Container 0. Note the index of the container you want to connect to.

Open the container shell

Make sure that the job's status is running. If the job is still starting, wait until it's running.

Open a shell in the container. Replace <jobId> with the job ID and <containerId> with the container index.

cosmicac jobs shell <jobId> <containerId>

The shell opens as appuser. This user can run commands as root with sudo. To close the shell, run exit.

Help and troubleshooting

nvidia-smi fails with Failed to initialize NVML: Unknown Error

If nvidia-smi displays Failed to initialize NVML: Unknown Error, the NVIDIA device files under /dev and /proc/driver/nvidia/version are present, but the GPU is not available to nvidia-smi.

  1. Restart the container from the shell.

    kill 1

    If you need root access, run sudo kill 1 instead.

  2. Reconnect to the container with cosmicac jobs shell and run nvidia-smi again.

The shell doesn't open when the job is Running

cosmicac-cli connects to cosmicac-wrk-agent-instance over hyperswarm-ssh. The connection then goes directly to the job's virtual machine.

Some cluster network configurations can block this connection, so the job can stay Running while the shell connection fails.

To route the connection through a relay, see Set up a relay for CosmicAC.

Next steps

On this page