Skip to main content

Slurm Clusters

On a Slurm cluster the Stealthium Client runs as a systemd service on every compute node, next to slurmd. It is installed from the same stealthium Debian package used for bare metal installs, and it sees every job the node runs without any change to your Slurm configuration.

Run it as a service, not a job

Don't launch the client through sbatch or srun. A job is confined to its allocation's cgroup and is killed when the allocation ends. The client needs the whole node and has to keep running between jobs.

Before you start​

  • An API key: follow Creating an API Key. Use one key for the whole cluster. Each node gets its own identity automatically.
  • Ubuntu on the compute nodes: packages are built for Ubuntu 22.04 and 24.04 on amd64, and Ubuntu 24.04 on arm64 (for example, Grace-based nodes). For other distributions, see the Docker install.
  • Outbound network access from every compute node to cnc.backend.stealthium.io on TCP port 443. The client connects directly and doesn't use an HTTP proxy, so a node that can only reach the internet through a proxy can't report.
  • The NVIDIA driver loaded on GPU nodes, as it already is for GPU jobs.

You install the client on compute nodes. Login and controller nodes don't run GPU workloads, but you can install it there too if you want to monitor them.

Get the package​

Choose the source that fits how your nodes reach the network:

  • Apt repository (Ubuntu 24.04, amd64): add the Stealthium repository as described in Install the Stealthium Repo. Nodes can then install and upgrade with apt-get.
  • Release .deb files (every supported Ubuntu version and architecture, or nodes that can't reach the apt repository): download the file for your nodes from the Stealthium releases page. Files are named stealthium_<version>-1_<arch>_<ubuntu-version>.deb, for example stealthium_1.1.37-1_amd64_ubuntu-22.04.deb. Put the file on storage every node can read, such as your shared filesystem.
Install stealthium, not stealthium-agent

The stealthium-agent package is the stand-alone evaluation agent and doesn't report to Stealthium Cloud. The two packages conflict: installing one removes the other.

Option 1: Bake it into the node image​

Most Slurm clusters boot compute nodes from a shared image (Warewulf, xCAT, Bright/Base Command Manager, or a golden disk image). Install the client into that image once, and every node that boots it starts reporting.

Run these commands inside the image, for example in the image chroot or the image build script:

# From the apt repository
apt-get install -y stealthium
# ...or from a release file
apt-get install -y ./stealthium_<version>-1_amd64_ubuntu-24.04.deb

# Configure the API key
install -d -m 700 /etc/stealthium
cat > /etc/stealthium/config.toml <<'EOF'
[orion]
api_key = "YOUR-KEY-HERE"
EOF
chmod 600 /etc/stealthium/config.toml

# Start on boot
systemctl enable stealthium

The config file must contain only the API key when the image is captured. On first start, the client writes a host_id line into this file to identify the node. If that line is baked into the image, every node reports as the same machine. Writing the file as the last step, as shown above, avoids that problem even if the client ran on the build host. For the same reason, follow your image tool's usual practice of clearing /etc/machine-id before capture.

Stateless nodes

On GPU nodes the host_id comes from the GPU UUIDs and the board's DMI ID. It stays the same across reboots even when /etc is rebuilt from the image each time. A CPU-only node falls back to /etc/machine-id, so on stateless nodes keep that file stable across reboots, or the node shows up as a new machine after each boot.

Option 2: Install on running nodes​

To add the client to nodes that are already up, run the install on every node in parallel with pdsh (or clush, or your configuration management tool). This example targets every node Slurm knows about. Run it as root from a host that has SSH access to the nodes:

NODES=$(sinfo -h -N -o '%N' | sort -u | paste -sd,)
DEB=/shared/stealthium/stealthium_<version>-1_amd64_ubuntu-24.04.deb

pdsh -w "$NODES" "
apt-get install -y $DEB &&
install -d -m 700 /etc/stealthium &&
printf '[orion]\napi_key = \"YOUR-KEY-HERE\"\n' > /etc/stealthium/config.toml &&
chmod 600 /etc/stealthium/config.toml &&
systemctl enable --now stealthium
"

Only run this on nodes that don't have the client yet. Rewriting config.toml on a node that is already reporting removes its saved host_id. On a GPU node the client derives the same ID again. A CPU-only node falls back to /etc/machine-id.

Installing the package doesn't restart slurmd or touch running jobs, so you don't need to drain nodes first.

Verify​

On a single node:

systemctl status stealthium
journalctl -u stealthium -n 50 --no-pager

The client also serves a health check on port 28080. It returns OK once the node is connected and sending data:

curl -s http://localhost:28080/health
# OK - Last send: 3s ago, CnC: connected

Across the cluster, list any node that isn't healthy:

pdsh -w "$NODES" 'curl -sf http://localhost:28080/health >/dev/null || echo UNHEALTHY' | grep UNHEALTHY

The nodes' GPUs then appear in the Compute Index of your workspace.

Upgrades and removal​

Nodes that use the apt repository upgrade with the rest of the system:

pdsh -w "$NODES" 'apt-get update && apt-get install -y --only-upgrade stealthium'

When you install from release files, run the same apt-get install -y ./stealthium_<new-version>...deb as before, or rebuild the node image. Upgrades keep /etc/stealthium/config.toml, so nodes keep their API key and identity.

To remove the client:

pdsh -w "$NODES" 'systemctl disable --now stealthium && apt-get remove -y stealthium'