Submitting Jobs from MATLAB
Everything the CONN interface can do can also be described in a plain MATLAB
structure and run with no interface at all. That structure is called BATCH,
and you hand it to the conn_batch function.
Scripting is worth the effort when you want a reproducible record of an analysis, when a dataset is too large to babysit, or when you expect to run the same pipeline again on new subjects.
The SCC uses the SGE scheduler, but many other clusters use SLURM. The scheduler-specific blocks below have a toggle so you can read either.
The BATCH Structure
BATCH is a nested structure whose fields mirror the tabs of the CONN
interface. You fill in the parts you need and leave the rest alone.
| Field | Corresponds to |
|---|---|
BATCH.filename |
The CONN project file to create or open |
BATCH.parallel |
How the work is distributed — see below |
BATCH.Setup |
The Setup tab: subjects, functionals, structurals, conditions |
BATCH.Denoising |
The Denoising tab |
BATCH.Analysis |
First-level analyses |
BATCH.Results |
Second-level (group) statistics |
Each section runs only if you set its done flag, which is what lets one
script build a project and a later one carry it forward.
An Example Script
This script builds a new CONN project from a BIDS dataset and runs Setup.
% ----------------------------------------------------------------------- %
% Example CONN batch script
%
% Everything the CONN GUI can do can also be described in a BATCH structure
% and run without the interface. This script builds a new CONN project from
% a BIDS dataset and runs the Setup step.
%
% Run it from MATLAB with: run('conn_batch_example.m')
% ----------------------------------------------------------------------- %
%% Paths
% your BIDS-valid input dataset
bidsDir = '/projectnb/your_project/bids';
% where CONN should save the project it builds. CONN creates a companion
% directory alongside this file, so give it a directory of its own
connProject = '/projectnb/your_project/conn_projects/conn_project.mat';
%% Parallelization
%
% parallel.N is the number of jobs CONN submits to the scheduler.
% 0 = run everything in this MATLAB session, submitting nothing
% >0 = split the work into N cluster jobs
%
% parallel.profile names one of CONN's built-in scheduler profiles. On the
% SCC that is the Grid Engine profile.
BATCH.filename = connProject;
BATCH.parallel.N = 0;
BATCH.parallel.profile = 'Grid Engine computer cluster';
% TODO: CONN's built-in Grid Engine profile does NOT pass an SCC project, so
% jobs submitted with parallel.N > 0 will be rejected unless you add one:
%
% BATCH.parallel.cmd_submitoptions = '-P your_project -l h_rt=12:00:00';
%
% Confirm the exact string needed before recommending it to users.
%% Setup
BATCH.Setup.isnew = 1; % 1 = create a new project, 0 = use an existing one
BATCH.Setup.done = 1; % 1 = actually run Setup, 0 = only fill in the fields
BATCH.Setup.overwrite = 1; % 1 = overwrite an existing project of this name
BATCH.Setup.add = 0;
% find the participants
subjects = spm_select('List', bidsDir, 'dir', 'sub-.*');
subjects = cellstr(subjects);
BATCH.Setup.nsubjects = length(subjects);
BATCH.Setup.RT = NaN; % repetition time; NaN reads it from the BIDS sidecars
BATCH.Setup.acquisitiontype = 1; % 1 = continuous
% ------------------------- functionals --------------------------------- %
for isub = 1:length(subjects)
csubDir = fullfile(bidsDir, subjects{isub});
scans = spm_select('FPListRec', csubDir, 'sub.*_bold\.nii\.gz');
scans = cellstr(scans);
for iscan = 1:length(scans)
BATCH.Setup.functionals{isub}{iscan} = scans{iscan};
end
end
% ------------------------- anatomicals --------------------------------- %
for isub = 1:length(subjects)
csubDir = fullfile(bidsDir, subjects{isub});
BATCH.Setup.structurals{isub} = spm_select('FPListRec', csubDir, '^sub.*_T1w\.nii\.gz$');
end
% ------------------------- conditions ---------------------------------- %
BATCH.Setup.conditions.missingdata = 1;
for isub = 1:length(subjects)
csubDir = fullfile(bidsDir, subjects{isub});
eventsFile = spm_select('FPListRec', csubDir, 'sub.*_events\.tsv');
eventsFile = cellstr(eventsFile);
BATCH.Setup.conditions.importfile{isub}{1} = eventsFile{1};
end
BATCH.Setup.conditions.add = 0;
% ------------------------- misc ---------------------------------------- %
BATCH.Setup.analyses = 1;
BATCH.Setup.analysisunits = 2;
BATCH.Setup.outputfiles = [0 0 0 0 0 0];
% TODO: this script stops after Setup. Add BATCH.Denoising, BATCH.Analysis
% and BATCH.Results sections to carry the project through the rest of the
% pipeline.
%% Run it
conn_batch(BATCH);
Copy it to work from:
[scc4]$ cp /project/scv/examples/imaging/conn/conn_batch_example.m .
Running It
Interactively
In a MATLAB session — see Launching CONN — run:
>> run('conn_batch_example.m')
Fine for a quick test. The session has to stay open until the script finishes, so it is a poor fit for anything long.
As a job
For a real run, submit a job that starts MATLAB with no display, runs the script, and exits. Copy the wrapper:
[scc4]$ cp /project/scv/examples/imaging/conn/conn_batch.qsub .
[scc4]$ cp /project/scv/examples/imaging/conn/conn_batch.sbatch .
Fill in the wrapper
#!/bin/bash -l
# Set your SCC project
#$ -P your_project
# Specify hard time limit for the job.
# The job will be aborted if it runs longer than this time.
#$ -l h_rt=12:00:00
# Send an email when the job finishes or if it is aborted (by default no email is sent).
#$ -m ea
# Give job a name
#$ -N conn
# Combine output and error files into a single file
#$ -j y
# Write the job logs here. This directory must already exist -- the scheduler
# will not create it, and the job fails immediately if it is missing.
#$ -o /projectnb/your_project/logs/
#$ -pe omp 8
#$ -l mem_per_core=4G
# Keep track of information related to the current job
echo "=========================================================="
echo "Start date : $(date)"
echo "Job name : $JOB_NAME"
echo "Job ID : $JOB_ID"
echo "=========================================================="
# CONN needs MATLAB and SPM, and the modules must be loaded in this order
module load matlab/2024b
module load spm/25.01.02
module load conn/25b
# the CONN batch script to run
CONN_SCRIPT='conn_batch_example.m'
# -batch runs the script with no display and exits when it finishes. MATLAB's
# exit status becomes the job's, so a MATLAB error fails the job
matlab -batch "run('$CONN_SCRIPT')"
#!/bin/bash -l
# Set the account the job is charged to (SGE calls this the project)
#SBATCH --account=your_project
# Specify hard time limit for the job.
# The job will be aborted if it runs longer than this time.
#SBATCH --time=12:00:00
# Send an email when the job finishes or if it fails. Add --mail-user=<address>
# if your site does not default to the submitting user.
#SBATCH --mail-type=END,FAIL
# Give job a name
#SBATCH --job-name=conn
# Write the job logs here. This directory must already exist -- the scheduler
# will not create it, and the job fails immediately if it is missing.
# %j is the job ID. Naming no --error file means errors are written into this
# same file.
#SBATCH --output=/projectnb/your_project/logs/conn-%j.out
#SBATCH --nodes=1
#SBATCH --ntasks=1
#SBATCH --cpus-per-task=8
#SBATCH --mem-per-cpu=4G
# Keep track of information related to the current job
echo "=========================================================="
echo "Start date : $(date)"
echo "Job name : $SLURM_JOB_NAME"
echo "Job ID : $SLURM_JOB_ID"
echo "=========================================================="
# CONN needs MATLAB and SPM, and the modules must be loaded in this order.
# Module names differ between sites -- adjust these to your own cluster
module load matlab/2024b
module load spm/25.01.02
module load conn/25b
# the CONN batch script to run
CONN_SCRIPT='conn_batch_example.m'
# -batch runs the script with no display and exits when it finishes. MATLAB's
# exit status becomes the job's, so a MATLAB error fails the job
matlab -batch "run('$CONN_SCRIPT')"
Create the log directory and submit
The scheduler will not create the log directory for you, and the job fails straight away if it is missing:
[scc4]$ mkdir -p /projectnb/your_project/logs
[scc4]$ qsub conn_batch.qsub
[scc4]$ mkdir -p /projectnb/your_project/logs
[scc4]$ sbatch conn_batch.sbatch
The job writes its log there, named conn.o<job ID>conn-<job ID>.out. Watch it with:
[scc4]$ qstat -u <your BU login name>
[scc4]$ squeue -u <your login name>
Letting CONN Split the Work
The wrapper above runs everything inside one job, because the example script sets:
BATCH.parallel.N = 0;
Set N to a number greater than zero and CONN stops doing the work itself.
Instead it submits N jobs to the scheduler, tracks them, and merges the
results — the same machinery described in
Submitting Jobs from the GUI, driven from a script:
BATCH.parallel.N = 20;
BATCH.parallel.profile = <span data-sched="sge">'Grid Engine computer cluster';</span><span data-sched="slurm">'Slurm computer cluster';</span>
As in the GUI, CONN's built-in profile does not pass an SCC project, so add one:
BATCH.parallel.cmd_submitoptions = '-P your_project -l h_rt=12:00:00';