Skip to content

Slurm Submission Scripts

Slurm submission scripts have three required parts:

  1. A shebang line, such as #!/bin/bash, which tells the shell to use Bash to run the script
  2. #SBATCH directives, which list the computational resources needed to run the job
  3. The commands that need to be executed

Example Task

Extract reads aligned to chromosome 21 from a BAM file.

Command:

samtools view INPUT_BAM_FILE -o OUTPUT_BAM_FILE chr21

Basic Submission Script

File name example: extract_chromosome_from_bam_file_ssubmit.sh

#!/bin/bash

#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

samtools view SEA-3-469_WGBS.bam \
         -o SEA-3-469_WGBS.chr21.bam \
         chr21

The suffix _ssubmit.sh can make Slurm submission scripts easier to identify.

Understanding #SBATCH Directives

#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
  • --job-name sets a name for the job
  • --cpus-per-task sets the number of cores or CPUs to request for the job
  • --mem-per-cpu sets the amount of memory to request; 4 GB is a reasonable default
  • --output sets the path for standard output; this naming scheme produces JOBNAME-JOBID.out
  • --error sets the path for standard error; this naming scheme produces JOBNAME-JOBID.err

If --error is omitted, the standard error stream is included with standard output in the --output file.

Refined Script

This version uses variables, full paths, and threads.

#!/bin/bash
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

INPUT_BAM_FILE="/analysis/cloud_projects/zirconium/GRCh38/methyl/SEA-3-469_WGBS/SEA-3-469_WGBS.bam"
OUTPUT_BAM_FILE="~/rcc_demo/outputs/SEA-3-469_WGBS.chr21.bam"

samtools view --threads 4 \
              ${INPUT_BAM_FILE} \
              -o ${OUTPUT_BAM_FILE} \
              chr21

Parameterized Script

This version accepts the BAM file path and chromosome as command-line arguments. The output file name is based on the input file name, which reduces the chance of typos.

#!/bin/bash
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err

TARGET_CHR=${2}
INPUT_BAM_FILE=${1}
OUTPUT_DIR="~/rcc_demo/outputs"

OUTPUT_FILE_PREFIX=$(basename ${INPUT_BAM_FILE} .bam)
OUTPUT_BAM_FILE="${OUTPUT_DIR}/${OUTPUT_FILE_PREFIX}.${TARGET_CHR}.bam"

samtools view --threads ${SLURM_CPUS_PER_TASK} \
              ${INPUT_BAM_FILE} \
              -o ${OUTPUT_BAM_FILE} \
              ${TARGET_CHR}

Resource requirements are saved in environment variables, which lets a script dynamically adjust when needed. In this example, ${SLURM_CPUS_PER_TASK} is used as the number of samtools threads.

Submitting the Script

Use sbatch:

sbatch extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21

You can override #SBATCH arguments with command-line arguments provided to sbatch.

# request 8 cores
sbatch --cpus-per-task=8 extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21

# set a different job name
sbatch --job-name="extract_chr21_SEA-3-469_WGBS" extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21

# use a high memory node
sbatch --partition="highmem" --mem-per-cpu=16G extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr7