The Batch System

Interactive sessions are for looking at data. Batch jobs are for processing it. You write a short script describing what to run and what resources it needs, hand it to the scheduler, and the scheduler finds a compute node and runs it — whether or not you are still logged in.

This is the whole point of a cluster. Eight subjects processed one at a time on a laptop takes most of a work week; the same eight submitted as eight batch jobs finish in the time one of them takes.

The SCC uses the Sun Grid Engine (SGE) scheduler.

Anatomy of a Job Script

A job script is an ordinary shell script with scheduler directives at the top. Every directive begins with #$, which the shell treats as a comment and the scheduler reads as an instruction.

Directive Does If you leave it out
#$ -P your_project Charge the job to this project. Required. No default — the job is rejected
#$ -l h_rt=12:00:00 Hard time limit. The job is killed when it expires. 12 hours
#$ -N bet Name the job The name of the script file
#$ -j y Merge the error log into the output log Separate output and error files
#$ -o /path/to/logs/ Where to write the log The directory you submitted from
#$ -m ea Email when the job ends or aborts No email
#$ -pe omp 4 Request CPU cores 1 core
#$ -l mem_per_core=4G Request memory per core 4 GB per core

An Example Job

This script skull-strips one image with FSL's Brain Extraction Tool.

#!/bin/bash -l

# Set the SCC project the job is charged to
#$ -P your_project

# Specify hard time limit for the job. The job is killed if it runs longer.
#$ -l h_rt=12:00:00

# Send an email when the job finishes or if it aborts
#$ -m ea

# Give the job a name
#$ -N bet

# Merge the error stream into the output stream
#$ -j y

# Write the job log here. This directory must already exist -- the scheduler
# will not create it, and the job fails immediately if it is missing.
#$ -o /projectnb/your_project/logs/

# Keep track of information related to the current job
echo "=========================================================="
echo "Start date : $(date)"
echo "Job name : $JOB_NAME"
echo "Job ID : $JOB_ID"
echo "=========================================================="

# Load the software this job needs. A batch job does NOT inherit the modules
# you loaded in your terminal, so load them here every time.
module load fsl/6.0.7.8

# Specify the image to skull-strip and where to write the result
INPUT='/projectnb/your_project/your_username/nifti/anat.nii.gz'
OUTPUT='/projectnb/your_project/your_username/nifti/anat_brain.nii.gz'

# Run the FSL Brain Extraction Tool. Without -v it succeeds silently and
# writes nothing to the log; -v makes it report what it is doing.
bet $INPUT $OUTPUT -v

Trying It Yourself

Copy the script into a directory of your own, so you can edit it:

[scc4]$ cp /project/scv/examples/imaging/sccintro/bet.qsub .

You also need a brain image to run it on. There is one ready to use in the workshop materials -- a T1-weighted anatomical scan. Copy it in next to the script:

[scc4]$ cp /project/scv/examples/imaging/tut_cnc_intro_scc/materials/T1w.nii.gz .

The . on the end of both commands means "into the directory I am standing in right now," so run them from wherever you want the files to land.

Now point the script at the image you just copied: set INPUT to T1w.nii.gz and OUTPUT to T1w_brain.nii.gz. You can type those into the Input image and Output image boxes above and copy the filled-in script, or open the file in a text editor and change the two lines by hand.

One last step before you submit: the log directory named in the -o line has to already exist. The scheduler will not create it for you. Make it now, substituting your own project:

[scc4]$ mkdir -p /projectnb/your_project/logs

The -p flag creates any missing parent directories along the way, and says nothing if the directory already exists -- so it is safe to run again.

Whatever path you use here must match the -o line in the script, and the Log directory box above. If they disagree, qsub accepts the job without complaint and it then lands in the Eqw error state described below, never having run your code.

Submitting

[scc4]$ qsub bet.qsub
Your job 1041461 ("bet.qsub") has been submitted

The number is the job ID. You will need it to check on or cancel the job.

Checking on a Job

[scc4]$ qstat -u your_username
7238828 0.14071 ood-deskto tuta1 r  09/12/2022 09:46:27 b@scc-bc2.scc.bu.edu 1
7243957 0.00000 bet.qsub   tuta1 qw 09/12/2022 14:47:20                      1

The fifth column is the state:

State Means
qw queued, waiting for resources
r running
Eqw error -- the job could not start

Note the first line: that is an OnDemand Desktop session, which is a batch job like any other.

Cancelling a Job

[scc4]$ qdel 7243957

Reading the Output

With #$ -j y and #$ -o, everything the job printed lands in one file named after the job and its ID:

[scc4]$ cat /projectnb/your_project/logs/bet.o1041461

Requesting Resources

A job asks the scheduler for four kinds of thing:

Resource Directive What it controls
Wall clock time #$ -l h_rt=HH:MM:SS How long the job may run before it is killed
CPU cores #$ -pe omp N How many cores the job may use at once
Memory #$ -l mem_per_core=NG Memory per core — total memory is this times the number of cores
GPUs #$ -l gpus=N How many GPUs the job gets

So What Should I Request?

It depends — on the software you are running, and on the data you are running it on.

  • Cores. Some software is written to spread its work across many cores and gets steadily faster the more you give it. Plenty of software cannot do that at all, and finishes in the same time on sixteen cores as on one, with the other fifteen sitting idle.
  • Memory. Some datasets fit comfortably in a couple of gigabytes; a high-resolution, multi-session study can need many times that. A job that runs out of memory is killed.
  • GPUs. A few tools can use a GPU and run dramatically faster with one. Most neuroimaging software cannot use one at all.

Asking for more than you need is not free, either: the larger the request, the longer the scheduler takes to find a machine that can satisfy it, so a big job often starts later than a small one.

The guides for individual packages give resource recommendations for the software they cover:

Running Many Subjects at Once

Processing twenty subjects does not mean writing twenty scripts. An array job runs many copies of one script at once — one per subject — and the scheduler places each copy on whatever node comes free.

You ask for one with the -t directive, which numbers the tasks:

#$ -t 1-20

That submits twenty tasks. Each is an independent job running the same script, and each can tell which one it is by reading the $SGE_TASK_ID variable: task 3 sees SGE_TASK_ID=3. The usual pattern is to keep a text file with one subject per line, and use $SGE_TASK_ID to pick out the line belonging to that task.

The MRIQC guide works through a complete array job, from the subject list to the submitted script: Running MRIQC on the SCC.

RCS documents array jobs in full on the advanced batch usage page.

Further Reading