SLURM (Simple Linux Utility for Resource Management) is an open-source workload manager used on the majority of HPC clusters. It handles job scheduling, resource allocation, and queue management across hundreds or thousands of nodes.
Key Concepts
- Job — a script submitted to the cluster that runs one or more tasks
- Node — a physical or virtual machine in the cluster
- Partition — a logical grouping of nodes (similar to a queue) with its own limits and policies
- Account — a billing/accounting group your jobs are charged against
Essential Commands
| Command | What it does |
|---|---|
sbatch job.sh |
Submit a batch job |
squeue -u $USER |
List your queued/running jobs |
scancel <jobid> |
Cancel a job |
sinfo |
Show partition and node status |
sacct -j <jobid> |
Show accounting info for a completed job |
srun |
Run a command interactively on a node |
A Minimal Job Script
#!/bin/bash
#SBATCH --job-name=my-job
#SBATCH --output=my-job-%j.out
#SBATCH --error=my-job-%j.err
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=4
#SBATCH --mem=8G
#SBATCH --time=01:00:00
#SBATCH --partition=short
module load python/3.11
python my_script.py
Submit it with:
sbatch job.sh
The %j in the output filename is replaced with the job ID automatically.
Requesting GPUs
#SBATCH --gres=gpu:1
#SBATCH --partition=gpu
Useful Tips
- Use
squeue --start -j <jobid>to see the estimated start time for a pending job. seff <jobid>(if available) gives a summary of CPU and memory efficiency after a job completes — handy for right-sizing future submissions.- Keep wall-time estimates tight; shorter jobs typically start sooner because they fit more easily into scheduling gaps.
- Use
--arrayfor parameter sweeps instead of submitting hundreds of individual jobs.
Job Arrays
#SBATCH --array=1-10
python process.py --index $SLURM_ARRAY_TASK_ID
This submits 10 jobs, each receiving a different SLURM_ARRAY_TASK_ID value (1 through 10).
Further Reading
- SLURM documentation
man sbatchon any cluster that runs SLURM