Skip to content
Merged

Fixes #213

Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
97 commits
Select commit Hold shift + click to select a range
f80b261
Register cellposev4 in benchmark run scripts
dariarom94 Jul 19, 2026
1a2fa09
fix anndata version mismatch with txsim
dariarom94 Jul 19, 2026
82add80
add segger to workflow (test)
dariarom94 Jul 19, 2026
53e1728
duplicates when FOV stiching cleaned up
dariarom94 Jul 19, 2026
1186b7a
chunks issue atera
dariarom94 Jul 20, 2026
18644d7
segger update image
dariarom94 Jul 20, 2026
ecb302d
claude fix for segger
dariarom94 Jul 20, 2026
7d66898
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
d400ebe
atera version fix
dariarom94 Jul 20, 2026
64d7b4e
wf for the custom rnaseq scripts
dariarom94 Jul 20, 2026
3edfbf1
adjust the loader image name
dariarom94 Jul 20, 2026
cbd2f12
adjust the memory
dariarom94 Jul 20, 2026
184260e
troubleshootig edges
dariarom94 Jul 20, 2026
9fa9a33
Merge branch 'main' into fixes
dariarom94 Jul 20, 2026
36631c4
segger update
dariarom94 Jul 21, 2026
0626127
cell type label correction
dariarom94 Jul 21, 2026
3186435
fix boundaries
dariarom94 Jul 21, 2026
d8a7d93
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
3505718
OOM fixes
dariarom94 Jul 21, 2026
d6e110a
fix code
dariarom94 Jul 21, 2026
4660f26
RCTD
dariarom94 Jul 21, 2026
5abd651
segger to RAPIDS
dariarom94 Jul 21, 2026
fe2e90a
Merge branch 'main' into fixes
dariarom94 Jul 21, 2026
0b23474
fix rctd
dariarom94 Jul 22, 2026
196ff1f
segger debug (torchvision)
dariarom94 Jul 22, 2026
4be7bd4
Merge branch 'main' into fixes
dariarom94 Jul 22, 2026
b8d3d7b
save the xenium version
dariarom94 Jul 22, 2026
202ac49
add atera to datasets
dariarom94 Jul 22, 2026
14be8d0
Add gene efficiency correction as a separate pipeline stage (#183)
dariarom94 Jul 22, 2026
0cf0243
moscot to pca and segger troubleshooting
dariarom94 Jul 22, 2026
d7afb84
added fastreseg
dariarom94 Jul 23, 2026
f87a1d9
segger bug new fix
dariarom94 Jul 23, 2026
123e112
fastreseg to workflow
dariarom94 Jul 23, 2026
7549589
add fastreseg test
dariarom94 Jul 23, 2026
9fa9604
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
ff04467
optimized fastreseg build
dariarom94 Jul 23, 2026
7ffc514
Merge branch 'main' into fixes
dariarom94 Jul 23, 2026
a7404d8
add s3 paths
dariarom94 Jul 23, 2026
19e5f83
troubleshoot comseg/segger
dariarom94 Jul 24, 2026
aaca151
segger update
dariarom94 Jul 25, 2026
c9bdb91
data loader bug
dariarom94 Jul 25, 2026
1df9834
Merge branch 'main' into fixes
dariarom94 Jul 25, 2026
4467d32
rctd adjustment (raw counts)
dariarom94 Jul 26, 2026
58912e4
fix segger and comseg
dariarom94 Jul 26, 2026
0a5aa99
optimize cosmx
dariarom94 Jul 26, 2026
298e666
Merge branch 'main' into fixes
dariarom94 Jul 26, 2026
c921937
parameter test for cellpose4
dariarom94 Jul 26, 2026
57c2d79
add atera
dariarom94 Jul 26, 2026
cf67e09
add a test in pciseq and dynamic memory for bruker
dariarom94 Jul 27, 2026
9f71692
add test to vizgen data
dariarom94 Jul 28, 2026
d795f33
Merge branch 'main' into fixes
dariarom94 Jul 28, 2026
fefaadc
param sweep
dariarom94 Jul 28, 2026
4dc08d3
add params to segmentation
dariarom94 Jul 28, 2026
e9505f1
adjust segger mem
dariarom94 Jul 29, 2026
3a155be
update fastreseg to tacco
dariarom94 Jul 29, 2026
acdd6c7
Add annotation + expression-correction parameter sweeps (rctd, ssam, …
dariarom94 Jul 30, 2026
fa462b8
Add moscot + split parameter sweeps (annotation, expression correction)
dariarom94 Jul 30, 2026
fd37aee
singler: read par['celltype_key'] instead of hardcoding "cell_type"
dariarom94 Jul 30, 2026
3449bd0
fastreseg
dariarom94 Jul 30, 2026
331d219
Merge branch 'main' into fixes
dariarom94 Jul 30, 2026
5961d0a
adjust labels
dariarom94 Jul 30, 2026
857e16a
bruker nsclc
dariarom94 Jul 31, 2026
54f023e
adjust bruker nsclc loader
dariarom94 Jul 31, 2026
6e8b9ea
setup
dariarom94 Jul 31, 2026
4eec493
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
44ad7af
method correction
dariarom94 Jul 31, 2026
b351148
Merge branch 'main' into fixes
dariarom94 Jul 31, 2026
2b8177c
pin anndata
dariarom94 Aug 1, 2026
6c16707
mirror nsclc
dariarom94 Aug 1, 2026
5618fb8
sync the vizgen files
dariarom94 Aug 2, 2026
b68c100
adjust mem for allen brain
dariarom94 Aug 2, 2026
98dbc69
claude notes
dariarom94 Aug 2, 2026
fa23294
claude notes
dariarom94 Aug 2, 2026
ac9492c
fix mirror script
dariarom94 Aug 2, 2026
26d292e
Merge branch 'main' into fixes
dariarom94 Aug 2, 2026
1ee371a
adapt fastreseg requirements
dariarom94 Aug 3, 2026
c53016a
merscope kuppe script update
dariarom94 Aug 3, 2026
d479db6
fix nsclc loader
dariarom94 Aug 3, 2026
6c54220
adjust mem for new test resources
dariarom94 Aug 4, 2026
9494fc5
adjust the nsclc loader for test resources
dariarom94 Aug 4, 2026
a459afc
change processor to avoid spatialdata 0.8.0 bug
dariarom94 Aug 4, 2026
8f7b0c1
pin spatialdata version
dariarom94 Aug 4, 2026
f18296b
memory fix
dariarom94 Aug 5, 2026
fc96dcb
Merge branch 'main' into fixes
dariarom94 Aug 5, 2026
e75042f
subsampling code
dariarom94 Aug 5, 2026
594bb25
remove unexisting dataset
dariarom94 Aug 5, 2026
90207aa
fix data extraction bag
dariarom94 Aug 5, 2026
a94ec38
transcript assignment edits
dariarom94 Aug 5, 2026
905b166
segger: simplify transcript-assignment OOB handling to an edge clamp
dariarom94 Aug 5, 2026
b368d5e
modify proseg (param sweep)
dariarom94 Aug 5, 2026
0f57e2f
add scale0
dariarom94 Aug 6, 2026
5f09575
Merge branch 'main' into fixes
dariarom94 Aug 6, 2026
c64f7d2
fix code bug
dariarom94 Aug 6, 2026
7317d57
expand test dataset space
dariarom94 Aug 6, 2026
fe3ad75
Merge branch 'main' into fixes
dariarom94 Aug 6, 2026
a0a3d43
fix stardist params
dariarom94 Aug 6, 2026
6669752
process_dataset: opt-in tissue-centered crop (fixes ABCA whole-brain …
dariarom94 Aug 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
108 changes: 108 additions & 0 deletions scripts/create_resources/combine/process_datasets_allen_nebius.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,108 @@
#!/bin/bash

# Process ONLY the Allen Brain Cell Atlas (ABCA) MERFISH datasets (combine step) — FOUR
# whole-brain sections, ALL paired with the SAME mouse-brain SC reference (like the LTX/MPII
# combine scripts, and unlike Kuppe which uses condition-specific references):
# mouse1_coronal, mouse2_coronal, mouse3_sagittal, mouse4_sagittal
# <-> allen_brain_cell_atlas/2023_yao_mouse_brain_scrnaseq_10xv2 (ABCA 2023 Yao 10xv2 atlas)
#
# Reads each spatial input (process_allen_brain_cell_atlas_merfish loader output) and the SC
# reference from the local /scratch raw folder, and writes the combined datasets to the
# /scratch datasets folder (same layout the other *_nebius.sh combine scripts use).
#
# Prerequisites (both publish to the same /scratch raw folder used below):
# 1. scripts/create_resources/spatial/process_allen_brain_cell_atlas_merfish_nebius.sh
# -> /scratch/.../raw/allen_brain_cell_atlas_merfish/mouse{1,2,3,4}_*/rep1/dataset.zarr
# 2. scripts/create_resources/sc/process_allen_brain_cell_atlas_brain_nebius.sh (log_cp -> hvg
# -> pca -> knn, i.e. it must carry a 'normalized' layer)
# -> /scratch/.../raw/allen_brain_cell_atlas/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad

# get the root of the directory
REPO_ROOT=$(git rev-parse --show-toplevel)

# ensure that the command below is run from the root of the repository
cd "$REPO_ROOT"

set -e

# process_allen_brain_cell_atlas_merfish_nebius.sh + process_allen_brain_cell_atlas_brain_nebius.sh
# both publish here.
raw_dir='/scratch/task_ist_preprocessing/raw'
publish_dir='/scratch/task_ist_preprocessing/datasets'

# shared SC reference for all four sections (ABCA 2023 Yao whole-brain 10xv2 atlas)
sc_ref="$raw_dir/allen_brain_cell_atlas/2023_yao_mouse_brain_scrnaseq_10xv2/dataset.h5ad"

launch_batch() {
local params_file="$1"
local label="$2"
tw launch https://github.com/openproblems-bio/task_ist_preprocessing.git \
--revision build/main \
--pull-latest \
--main-script target/nextflow/workflows/process_datasets/main.nf \
--workspace 167877437119966 \
--compute-env 5hfmdCBxMRd4nHZaJKYEQZ \
--params-file "$params_file" \
--config src/base/labels_nebius.config \
--labels "task_ist_preprocessing,process_datasets,$label"
}

cat > /tmp/params_allen.yaml << HERE
param_list:

- id: "allen_brain_cell_atlas_merfish_combined/mouse1_coronal/rep1"
input_sp: "$raw_dir/allen_brain_cell_atlas_merfish/mouse1_coronal/rep1/dataset.zarr"
input_sc: "$sc_ref"
dataset_id: "allen_brain_cell_atlas_merfish_combined/mouse1_coronal/rep1"
dataset_name: "Mouse brain combined ABCA MERFISH mouse1 coronal + 2023 Yao scRNAseq"
dataset_url: "https://download.brainimagelibrary.org/29/3c/293cc39ceea87f6d/"
dataset_reference: "10.1038/s41586-023-06812-z"
dataset_summary: "Allen Brain Cell Atlas whole-brain MERFISH mouse 1 (coronal, ~1100-gene panel) + 2023 Yao mouse-brain scRNAseq"
dataset_description: "Brain-wide MERFISH spatial transcriptomics (Zhuang lab, Allen Brain Cell Atlas), mouse 1 coronal section, paired with the ABCA 2023 Yao whole-mouse-brain 10xv2 scRNAseq reference."
dataset_organism: "mus_musculus"

- id: "allen_brain_cell_atlas_merfish_combined/mouse2_coronal/rep1"
input_sp: "$raw_dir/allen_brain_cell_atlas_merfish/mouse2_coronal/rep1/dataset.zarr"
input_sc: "$sc_ref"
dataset_id: "allen_brain_cell_atlas_merfish_combined/mouse2_coronal/rep1"
dataset_name: "Mouse brain combined ABCA MERFISH mouse2 coronal + 2023 Yao scRNAseq"
dataset_url: "https://download.brainimagelibrary.org/29/3c/293cc39ceea87f6d/"
dataset_reference: "10.1038/s41586-023-06812-z"
dataset_summary: "Allen Brain Cell Atlas whole-brain MERFISH mouse 2 (coronal, ~1100-gene panel) + 2023 Yao mouse-brain scRNAseq"
dataset_description: "Brain-wide MERFISH spatial transcriptomics (Zhuang lab, Allen Brain Cell Atlas), mouse 2 coronal section, paired with the ABCA 2023 Yao whole-mouse-brain 10xv2 scRNAseq reference."
dataset_organism: "mus_musculus"

- id: "allen_brain_cell_atlas_merfish_combined/mouse3_sagittal/rep1"
input_sp: "$raw_dir/allen_brain_cell_atlas_merfish/mouse3_sagittal/rep1/dataset.zarr"
input_sc: "$sc_ref"
dataset_id: "allen_brain_cell_atlas_merfish_combined/mouse3_sagittal/rep1"
dataset_name: "Mouse brain combined ABCA MERFISH mouse3 sagittal + 2023 Yao scRNAseq"
dataset_url: "https://download.brainimagelibrary.org/29/3c/293cc39ceea87f6d/"
dataset_reference: "10.1038/s41586-023-06812-z"
dataset_summary: "Allen Brain Cell Atlas whole-brain MERFISH mouse 3 (sagittal, ~1100-gene panel) + 2023 Yao mouse-brain scRNAseq"
dataset_description: "Brain-wide MERFISH spatial transcriptomics (Zhuang lab, Allen Brain Cell Atlas), mouse 3 sagittal section, paired with the ABCA 2023 Yao whole-mouse-brain 10xv2 scRNAseq reference."
dataset_organism: "mus_musculus"

- id: "allen_brain_cell_atlas_merfish_combined/mouse4_sagittal/rep1"
input_sp: "$raw_dir/allen_brain_cell_atlas_merfish/mouse4_sagittal/rep1/dataset.zarr"
input_sc: "$sc_ref"
dataset_id: "allen_brain_cell_atlas_merfish_combined/mouse4_sagittal/rep1"
dataset_name: "Mouse brain combined ABCA MERFISH mouse4 sagittal + 2023 Yao scRNAseq"
dataset_url: "https://download.brainimagelibrary.org/29/3c/293cc39ceea87f6d/"
dataset_reference: "10.1038/s41586-023-06812-z"
dataset_summary: "Allen Brain Cell Atlas whole-brain MERFISH mouse 4 (sagittal, ~1100-gene panel) + 2023 Yao mouse-brain scRNAseq"
dataset_description: "Brain-wide MERFISH spatial transcriptomics (Zhuang lab, Allen Brain Cell Atlas), mouse 4 sagittal section, paired with the ABCA 2023 Yao whole-mouse-brain 10xv2 scRNAseq reference."
dataset_organism: "mus_musculus"

# ABCA whole-brain images are enormous (~83k x 102k px) and the tissue sits off-centre,
# so an image-centred crop can miss it (mouse1: 0 of 42M transcripts). Centre the crop
# on the transcript density instead (opt-in flag; default is image-centred).
tissue_centered_crop: true

output_sc: "\$id/output_sc.h5ad"
output_sp: "\$id/output_sp.zarr"
output_state: "\$id/state.yaml"
publish_dir: "$publish_dir"
HERE

launch_batch /tmp/params_allen.yaml "allen_brain_cell_atlas_merfish"
96 changes: 96 additions & 0 deletions scripts/run_benchmark/param_sweep/run_test_baysor_nebius.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
#!/bin/bash

# Nebius test run: all default methods + baysor at the transcript-assignment stage,
# with a parameter sweep over baysor's tuning knobs.
# See src/methods_transcript_assignment/baysor/NOTES.md ("Optimization / tuning").
#
# The stage default (basic_transcript_assignment) stays enabled alongside baysor so the run
# also produces the baseline the sweep is scored against — the workflow allows at most ONE
# non-default variant at a time, and baysor is that one non-default method here; every OTHER
# stage stays on its single default.
#
# PARAMS-FILE CAVEAT: `method_parameters_yaml` is opened by the WORKFLOW at runtime on the cloud
# (readYaml -> Nextflow file()), so a local /tmp path won't exist there and /scratch is read-only
# from the launch host. file() DOES stage http(s):// and this repo is public, so the sweep lives
# in a COMMITTED file read from GitHub via its raw URL => commit AND PUSH
# scripts/run_benchmark/param_sweep/baysor_params.yaml to $params_branch before launching.
# This is independent of --revision (which selects the pipeline CODE).

# get the root of the directory
REPO_ROOT=$(git rev-parse --show-toplevel)

# ensure that the command below is run from the root of the repository
cd "$REPO_ROOT"

set -e

resources_test_s3="/scratch/task_ist_preprocessing/resources_test/task_ist_preprocessing/"
# Results publish to /scratch — created and written by the cloud compute env (read-only here).
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_baysor"

# The sweep lives in a committed file, read from GitHub at runtime. $params_branch defaults to
# the branch you are on; the file must be pushed there. (Independent of --revision below.)
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweep/baysor_params.yaml"

cat > /tmp/params_settings.yaml << HERE
default_methods:
- custom_segmentation
- basic_transcript_assignment
- basic_count_aggregation
- basic_qc_filter
- alpha_shapes
- normalize_by_volume
- tacco
- no_correction
segmentation_methods:
- custom_segmentation
transcript_assignment_methods:
- basic_transcript_assignment
- baysor
count_aggregation_methods:
- basic_count_aggregation
qc_filtering_methods:
- basic_qc_filter
volume_calculation_methods:
- alpha_shapes
normalization_methods:
- normalize_by_volume
celltype_annotation_methods:
- tacco
expression_correction_methods:
- no_correction
gene_efficiency_correction_methods:
- no_correction
method_parameters_yaml: $params_url
HERE

# Write the parameters to file (input_states version, NOTE: enable `-entry_name auto` for this)
cat > /tmp/params.yaml << HERE
input_states: $resources_test_s3/**/state.yaml
rename_keys: 'input_sc:output_sc;input_sp:output_sp'
save_spatial_data: false
settings: '$(yq -o json /tmp/params_settings.yaml | jq -c .)'
output_state: "state.yaml"
publish_dir: "$publish_dir"
HERE

# Fail early with a clear message if the params file isn't reachable on GitHub yet.
if ! curl -fsSL -o /dev/null "$params_url"; then
echo "ERROR: params file not reachable at:" >&2
echo " $params_url" >&2
echo "Commit and push scripts/run_benchmark/param_sweep/baysor_params.yaml to '$params_branch' first." >&2
exit 1
fi

tw launch https://github.com/openproblems-bio/task_ist_preprocessing.git \
--revision build/main \
--pull-latest \
--main-script target/nextflow/workflows/run_benchmark/main.nf \
--workspace 167877437119966 \
--compute-env 5hfmdCBxMRd4nHZaJKYEQZ \
--params-file /tmp/params.yaml \
--entry-name auto \
--config src/base/labels_nebius.config \
--labels task_ist_preprocessing,test,baysor
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@ cd "$REPO_ROOT"

set -e

resources_test_s3=s3://openproblems-data/resources_test/task_ist_preprocessing
resources_test_s3="/scratch/task_ist_preprocessing/resources_test/task_ist_preprocessing/"
# Results publish to /scratch — created and written by the cloud compute env, so
# the launcher does NOT create it here (it is read-only from the launch host).
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_binning"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ cd "$REPO_ROOT"

set -e

resources_test_s3=s3://openproblems-data/resources_test/task_ist_preprocessing
resources_test_s3="/scratch/task_ist_preprocessing/resources_test/task_ist_preprocessing/"
# Results publish to /scratch — created and written by the cloud compute env, so
# the launcher does NOT create it here (it is read-only from the launch host).
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_cellpose"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@
# publish) is READ-ONLY from the launch host — which is why the binning
# method_params block is commented out in run_test_nebius.sh. file() does stage
# http(s):// though, and this repo is public, so we keep the sweep in a COMMITTED
# file (scripts/run_benchmark/cellposev4_params.yaml) and read it from GitHub via
# file (scripts/run_benchmark/param_sweep/cellposev4_params.yaml) and read it from GitHub via
# its raw URL. => the params file must be committed AND PUSHED to $params_branch
# before launching (edit the file there, not here, to change the sweep).

Expand All @@ -27,7 +27,7 @@ cd "$REPO_ROOT"

set -e

resources_test_s3=s3://openproblems-data/resources_test/task_ist_preprocessing
resources_test_s3="/scratch/task_ist_preprocessing/resources_test/task_ist_preprocessing/"
# Results publish to /scratch — created and written by the cloud compute env, so
# the launcher does NOT create it here (it is read-only from the launch host).
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_cellposev4"
Expand All @@ -37,7 +37,7 @@ publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_cellposev4"
# is independent of --revision below, which selects the pipeline CODE to run.)
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/cellposev4_params.yaml"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweep/cellposev4_params.yaml"

cat > /tmp/params_settings.yaml << HERE
default_methods:
Expand Down Expand Up @@ -89,7 +89,7 @@ HERE
if ! curl -fsSL -o /dev/null "$params_url"; then
echo "ERROR: params file not reachable at:" >&2
echo " $params_url" >&2
echo "Commit and push scripts/run_benchmark/cellposev4_params.yaml to '$params_branch' first." >&2
echo "Commit and push scripts/run_benchmark/param_sweep/cellposev4_params.yaml to '$params_branch' first." >&2
exit 1
fi

Expand Down
96 changes: 96 additions & 0 deletions scripts/run_benchmark/param_sweep/run_test_clustermap_nebius.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,96 @@
#!/bin/bash

# Nebius test run: all default methods + clustermap at the transcript-assignment stage,
# with a parameter sweep over clustermap's tuning knobs.
# See src/methods_transcript_assignment/clustermap/NOTES.md ("Optimization / tuning").
#
# The stage default (basic_transcript_assignment) stays enabled alongside clustermap so the run
# also produces the baseline the sweep is scored against — the workflow allows at most ONE
# non-default variant at a time, and clustermap is that one non-default method here; every OTHER
# stage stays on its single default.
#
# PARAMS-FILE CAVEAT: `method_parameters_yaml` is opened by the WORKFLOW at runtime on the cloud
# (readYaml -> Nextflow file()), so a local /tmp path won't exist there and /scratch is read-only
# from the launch host. file() DOES stage http(s):// and this repo is public, so the sweep lives
# in a COMMITTED file read from GitHub via its raw URL => commit AND PUSH
# scripts/run_benchmark/param_sweep/clustermap_params.yaml to $params_branch before launching.
# This is independent of --revision (which selects the pipeline CODE).

# get the root of the directory
REPO_ROOT=$(git rev-parse --show-toplevel)

# ensure that the command below is run from the root of the repository
cd "$REPO_ROOT"

set -e

resources_test_s3="/scratch/task_ist_preprocessing/resources_test/task_ist_preprocessing/"
# Results publish to /scratch — created and written by the cloud compute env (read-only here).
publish_dir="/scratch/results/runs/$(date +%Y-%m-%d_%H-%M-%S)_clustermap"

# The sweep lives in a committed file, read from GitHub at runtime. $params_branch defaults to
# the branch you are on; the file must be pushed there. (Independent of --revision below.)
params_repo="openproblems-bio/task_ist_preprocessing"
params_branch="$(git rev-parse --abbrev-ref HEAD)"
params_url="https://raw.githubusercontent.com/${params_repo}/${params_branch}/scripts/run_benchmark/param_sweep/clustermap_params.yaml"

cat > /tmp/params_settings.yaml << HERE
default_methods:
- custom_segmentation
- basic_transcript_assignment
- basic_count_aggregation
- basic_qc_filter
- alpha_shapes
- normalize_by_volume
- tacco
- no_correction
segmentation_methods:
- custom_segmentation
transcript_assignment_methods:
- basic_transcript_assignment
- clustermap
count_aggregation_methods:
- basic_count_aggregation
qc_filtering_methods:
- basic_qc_filter
volume_calculation_methods:
- alpha_shapes
normalization_methods:
- normalize_by_volume
celltype_annotation_methods:
- tacco
expression_correction_methods:
- no_correction
gene_efficiency_correction_methods:
- no_correction
method_parameters_yaml: $params_url
HERE

# Write the parameters to file (input_states version, NOTE: enable `-entry_name auto` for this)
cat > /tmp/params.yaml << HERE
input_states: $resources_test_s3/**/state.yaml
rename_keys: 'input_sc:output_sc;input_sp:output_sp'
save_spatial_data: false
settings: '$(yq -o json /tmp/params_settings.yaml | jq -c .)'
output_state: "state.yaml"
publish_dir: "$publish_dir"
HERE

# Fail early with a clear message if the params file isn't reachable on GitHub yet.
if ! curl -fsSL -o /dev/null "$params_url"; then
echo "ERROR: params file not reachable at:" >&2
echo " $params_url" >&2
echo "Commit and push scripts/run_benchmark/param_sweep/clustermap_params.yaml to '$params_branch' first." >&2
exit 1
fi

tw launch https://github.com/openproblems-bio/task_ist_preprocessing.git \
--revision build/main \
--pull-latest \
--main-script target/nextflow/workflows/run_benchmark/main.nf \
--workspace 167877437119966 \
--compute-env 5hfmdCBxMRd4nHZaJKYEQZ \
--params-file /tmp/params.yaml \
--entry-name auto \
--config src/base/labels_nebius.config \
--labels task_ist_preprocessing,test,clustermap
Loading
Loading