The Batch System
Interactive sessions are for looking at data. Batch jobs are for processing it. You write a short script describing what to run and what resources it needs, hand it to the scheduler, and the scheduler finds a compute node and runs it — whether or not you are still logged in.
This is the whole point of a cluster. Eight subjects processed one at a time on a laptop takes most of a work week; the same eight submitted as eight batch jobs finish in the time one of them takes.
The SCC uses the Sun Grid Engine (SGE) scheduler.
Anatomy of a Job Script
A job script is an ordinary shell script with scheduler directives at the
top. Every directive begins with #$, which the shell treats as a comment
and the scheduler reads as an instruction.
| Directive | Does | If you leave it out |
|---|---|---|
#$ -P your_project |
Charge the job to this project. Required. | No default — the job is rejected |
#$ -l h_rt=12:00:00 |
Hard time limit. The job is killed when it expires. | 12 hours |
#$ -N bet |
Name the job | The name of the script file |
#$ -j y |
Merge the error log into the output log | Separate output and error files |
#$ -o /path/to/logs/ |
Where to write the log | The directory you submitted from |
#$ -m ea |
Email when the job ends or aborts | No email |
#$ -pe omp 4 |
Request CPU cores | 1 core |
#$ -l mem_per_core=4G |
Request memory per core | 4 GB per core |
An Example Job
This script skull-strips one image with FSL's Brain Extraction Tool.
#!/bin/bash -l
# Set the SCC project the job is charged to
#$ -P your_project
# Specify hard time limit for the job. The job is killed if it runs longer.
#$ -l h_rt=12:00:00
# Send an email when the job finishes or if it aborts
#$ -m ea
# Give the job a name
#$ -N bet
# Merge the error stream into the output stream
#$ -j y
# Write the job log here. This directory must already exist -- the scheduler
# will not create it, and the job fails immediately if it is missing.
#$ -o /projectnb/your_project/logs/
# Keep track of information related to the current job
echo "=========================================================="
echo "Start date : $(date)"
echo "Job name : $JOB_NAME"
echo "Job ID : $JOB_ID"
echo "=========================================================="
# Load the software this job needs. A batch job does NOT inherit the modules
# you loaded in your terminal, so load them here every time.
module load fsl/6.0.7.8
# Specify the image to skull-strip and where to write the result
INPUT='/projectnb/your_project/your_username/nifti/anat.nii.gz'
OUTPUT='/projectnb/your_project/your_username/nifti/anat_brain.nii.gz'
# Run the FSL Brain Extraction Tool. Without -v it succeeds silently and
# writes nothing to the log; -v makes it report what it is doing.
bet $INPUT $OUTPUT -v
Trying It Yourself
Copy the script into a directory of your own, so you can edit it:
[scc4]$ cp /project/scv/examples/imaging/sccintro/bet.qsub .
You also need a brain image to run it on. There is one ready to use in the workshop materials -- a T1-weighted anatomical scan. Copy it in next to the script:
[scc4]$ cp /project/scv/examples/imaging/tut_cnc_intro_scc/materials/T1w.nii.gz .
The . on the end of both commands means "into the directory I am standing
in right now," so run them from wherever you want the files to land.
Now point the script at the image you just copied: set INPUT to
T1w.nii.gz and OUTPUT to T1w_brain.nii.gz. You can type those into the
Input image and Output image boxes above and copy the filled-in
script, or open the file in a text editor and change the two lines by hand.
One last step before you submit: the log directory named in the -o line has
to already exist. The scheduler will not create it for you. Make it now,
substituting your own project:
[scc4]$ mkdir -p /projectnb/your_project/logs
The -p flag creates any missing parent directories along the way, and says
nothing if the directory already exists -- so it is safe to run again.
Whatever path you use here must match the -o line in the script, and the
Log directory box above. If they disagree, qsub accepts the job without
complaint and it then lands in the Eqw error state described below, never
having run your code.
Submitting
[scc4]$ qsub bet.qsub
Your job 1041461 ("bet.qsub") has been submitted
The number is the job ID. You will need it to check on or cancel the job.
Checking on a Job
[scc4]$ qstat -u your_username
7238828 0.14071 ood-deskto tuta1 r 09/12/2022 09:46:27 b@scc-bc2.scc.bu.edu 1
7243957 0.00000 bet.qsub tuta1 qw 09/12/2022 14:47:20 1
The fifth column is the state:
| State | Means |
|---|---|
qw |
queued, waiting for resources |
r |
running |
Eqw |
error -- the job could not start |
Note the first line: that is an OnDemand Desktop session, which is a batch job like any other.
Cancelling a Job
[scc4]$ qdel 7243957
Reading the Output
With #$ -j y and #$ -o, everything the job printed lands in one file named
after the job and its ID:
[scc4]$ cat /projectnb/your_project/logs/bet.o1041461
Requesting Resources
A job asks the scheduler for four kinds of thing:
| Resource | Directive | What it controls |
|---|---|---|
| Wall clock time | #$ -l h_rt=HH:MM:SS |
How long the job may run before it is killed |
| CPU cores | #$ -pe omp N |
How many cores the job may use at once |
| Memory | #$ -l mem_per_core=NG |
Memory per core — total memory is this times the number of cores |
| GPUs | #$ -l gpus=N |
How many GPUs the job gets |
So What Should I Request?
It depends — on the software you are running, and on the data you are running it on.
- Cores. Some software is written to spread its work across many cores and gets steadily faster the more you give it. Plenty of software cannot do that at all, and finishes in the same time on sixteen cores as on one, with the other fifteen sitting idle.
- Memory. Some datasets fit comfortably in a couple of gigabytes; a high-resolution, multi-session study can need many times that. A job that runs out of memory is killed.
- GPUs. A few tools can use a GPU and run dramatically faster with one. Most neuroimaging software cannot use one at all.
Asking for more than you need is not free, either: the larger the request, the longer the scheduler takes to find a machine that can satisfy it, so a big job often starts later than a small one.
The guides for individual packages give resource recommendations for the software they cover:
Running Many Subjects at Once
Processing twenty subjects does not mean writing twenty scripts. An array job runs many copies of one script at once — one per subject — and the scheduler places each copy on whatever node comes free.
You ask for one with the -t directive, which numbers the tasks:
#$ -t 1-20
That submits twenty tasks. Each is an independent job running the same
script, and each can tell which one it is by reading the $SGE_TASK_ID
variable: task 3 sees SGE_TASK_ID=3. The usual pattern is to keep a text
file with one subject per line, and use $SGE_TASK_ID to pick out the line
belonging to that task.
The MRIQC guide works through a complete array job, from the subject list to the submitted script: Running MRIQC on the SCC.
RCS documents array jobs in full on the advanced batch usage page.