Slurm Submission Scripts¶
Slurm submission scripts have three required parts:
- A shebang line, such as
#!/bin/bash, which tells the shell to use Bash to run the script #SBATCHdirectives, which list the computational resources needed to run the job- The commands that need to be executed
Example Task¶
Extract reads aligned to chromosome 21 from a BAM file.
Command:
samtools view INPUT_BAM_FILE -o OUTPUT_BAM_FILE chr21
Basic Submission Script¶
File name example: extract_chromosome_from_bam_file_ssubmit.sh
#!/bin/bash
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
samtools view SEA-3-469_WGBS.bam \
-o SEA-3-469_WGBS.chr21.bam \
chr21
The suffix _ssubmit.sh can make Slurm submission scripts easier to identify.
Understanding #SBATCH Directives¶
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
--job-namesets a name for the job--cpus-per-tasksets the number of cores or CPUs to request for the job--mem-per-cpusets the amount of memory to request; 4 GB is a reasonable default--outputsets the path for standard output; this naming scheme producesJOBNAME-JOBID.out--errorsets the path for standard error; this naming scheme producesJOBNAME-JOBID.err
If --error is omitted, the standard error stream is included with standard output in the --output file.
Refined Script¶
This version uses variables, full paths, and threads.
#!/bin/bash
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
INPUT_BAM_FILE="/analysis/cloud_projects/zirconium/GRCh38/methyl/SEA-3-469_WGBS/SEA-3-469_WGBS.bam"
OUTPUT_BAM_FILE="~/rcc_demo/outputs/SEA-3-469_WGBS.chr21.bam"
samtools view --threads 4 \
${INPUT_BAM_FILE} \
-o ${OUTPUT_BAM_FILE} \
chr21
Parameterized Script¶
This version accepts the BAM file path and chromosome as command-line arguments. The output file name is based on the input file name, which reduces the chance of typos.
#!/bin/bash
#SBATCH --job-name=extract_chrom
#SBATCH --cpus-per-task=4
#SBATCH --mem-per-cpu=4G
#SBATCH --output=logs/%x-%j.out
#SBATCH --error=logs/%x-%j.err
TARGET_CHR=${2}
INPUT_BAM_FILE=${1}
OUTPUT_DIR="~/rcc_demo/outputs"
OUTPUT_FILE_PREFIX=$(basename ${INPUT_BAM_FILE} .bam)
OUTPUT_BAM_FILE="${OUTPUT_DIR}/${OUTPUT_FILE_PREFIX}.${TARGET_CHR}.bam"
samtools view --threads ${SLURM_CPUS_PER_TASK} \
${INPUT_BAM_FILE} \
-o ${OUTPUT_BAM_FILE} \
${TARGET_CHR}
Resource requirements are saved in environment variables, which lets a script dynamically adjust when needed. In this example, ${SLURM_CPUS_PER_TASK} is used as the number of samtools threads.
Submitting the Script¶
Use sbatch:
sbatch extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21
You can override #SBATCH arguments with command-line arguments provided to sbatch.
# request 8 cores
sbatch --cpus-per-task=8 extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21
# set a different job name
sbatch --job-name="extract_chr21_SEA-3-469_WGBS" extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr21
# use a high memory node
sbatch --partition="highmem" --mem-per-cpu=16G extract_chromosome_from_bam_file_ssubmit.sh SEA-3-469_WGBS.bam chr7