Slurm Clusters
On a Slurm cluster the Stealthium Client runs as a systemd service on every compute node, next to slurmd. It is installed from the same stealthium Debian package used for bare metal installs, and it sees every job the node runs without any change to your Slurm configuration.
Don't launch the client through sbatch or srun. A job is confined to its allocation's cgroup and is killed when the allocation ends. The client needs the whole node and has to keep running between jobs.
Before you start
- An API key: follow Creating an API Key. Use one key for the whole cluster. Each node gets its own identity automatically.
- Ubuntu on the compute nodes: packages are built for Ubuntu 22.04 and 24.04 on
amd64, and Ubuntu 24.04 onarm64(for example, Grace-based nodes). For other distributions, see the Docker install. - Outbound network access from every compute node to
cnc.backend.stealthium.ioon TCP port 443. The client connects directly and doesn't use an HTTP proxy, so a node that can only reach the internet through a proxy can't report. - The NVIDIA driver loaded on GPU nodes, as it already is for GPU jobs.
You install the client on compute nodes. Login and controller nodes don't run GPU workloads, but you can install it there too if you want to monitor them.
Get the package
Choose the source that fits how your nodes reach the network:
- Apt repository (Ubuntu 24.04,
amd64): add the Stealthium repository as described in Install the Stealthium Repo. Nodes can then install and upgrade withapt-get. - Release
.debfiles (every supported Ubuntu version and architecture, or nodes that can't reach the apt repository): download the file for your nodes from the Stealthium releases page. Files are namedstealthium_<version>-1_<arch>_<ubuntu-version>.deb, for examplestealthium_1.1.37-1_amd64_ubuntu-22.04.deb. Put the file on storage every node can read, such as your shared filesystem.
stealthium, not stealthium-agentThe stealthium-agent package is the stand-alone evaluation agent and doesn't report to Stealthium Cloud. The two packages conflict: installing one removes the other.
Option 1: Bake it into the node image
Most Slurm clusters boot compute nodes from a shared image (Warewulf, xCAT, Bright/Base Command Manager, or a golden disk image). Install the client into that image once, and every node that boots it starts reporting.
Run these commands inside the image, for example in the image chroot or the image build script:
# From the apt repository
apt-get install -y stealthium
# ...or from a release file
apt-get install -y ./stealthium_<version>-1_amd64_ubuntu-24.04.deb
# Configure the API key
install -d -m 700 /etc/stealthium
cat > /etc/stealthium/config.toml <<'EOF'
[orion]
api_key = "YOUR-KEY-HERE"
EOF
chmod 600 /etc/stealthium/config.toml
# Start on boot
systemctl enable stealthium
The config file must contain only the API key when the image is captured. On first start, the client writes a host_id line into this file to identify the node. If that line is baked into the image, every node reports as the same machine. Writing the file as the last step, as shown above, avoids that problem even if the client ran on the build host. For the same reason, follow your image tool's usual practice of clearing /etc/machine-id before capture.
On GPU nodes the host_id comes from the GPU UUIDs and the board's DMI ID. It stays the same across reboots even when /etc is rebuilt from the image each time. A CPU-only node falls back to /etc/machine-id, so on stateless nodes keep that file stable across reboots, or the node shows up as a new machine after each boot.
Option 2: Install on running nodes
To add the client to nodes that are already up, run the install on every node in parallel with pdsh (or clush, or your configuration management tool). This example targets every node Slurm knows about. Run it as root from a host that has SSH access to the nodes:
NODES=$(sinfo -h -N -o '%N' | sort -u | paste -sd,)
DEB=/shared/stealthium/stealthium_<version>-1_amd64_ubuntu-24.04.deb
pdsh -w "$NODES" "
apt-get install -y $DEB &&
install -d -m 700 /etc/stealthium &&
printf '[orion]\napi_key = \"YOUR-KEY-HERE\"\n' > /etc/stealthium/config.toml &&
chmod 600 /etc/stealthium/config.toml &&
systemctl enable --now stealthium
"
Only run this on nodes that don't have the client yet. Rewriting config.toml on a node that is already reporting removes its saved host_id. On a GPU node the client derives the same ID again. A CPU-only node falls back to /etc/machine-id.
Installing the package doesn't restart slurmd or touch running jobs, so you don't need to drain nodes first.
Verify
On a single node:
systemctl status stealthium
journalctl -u stealthium -n 50 --no-pager
The client also serves a health check on port 28080. It returns OK once the node is connected and sending data:
curl -s http://localhost:28080/health
# OK - Last send: 3s ago, CnC: connected
Across the cluster, list any node that isn't healthy:
pdsh -w "$NODES" 'curl -sf http://localhost:28080/health >/dev/null || echo UNHEALTHY' | grep UNHEALTHY
The nodes' GPUs then appear in the Compute Index of your workspace.
Upgrades and removal
Nodes that use the apt repository upgrade with the rest of the system:
pdsh -w "$NODES" 'apt-get update && apt-get install -y --only-upgrade stealthium'
When you install from release files, run the same apt-get install -y ./stealthium_<new-version>...deb as before, or rebuild the node image. Upgrades keep /etc/stealthium/config.toml, so nodes keep their API key and identity.
To remove the client:
pdsh -w "$NODES" 'systemctl disable --now stealthium && apt-get remove -y stealthium'