Joint masked pretraining across histopathology (WSI), transcriptomics (RNA-seq) and DNA methylation.
mtcp implements a multimodal ViT-MAE architecture that learns a shared, process-level
representation of tumor data. Omics modalities are encoded with GO-guided tokenization
(biologically interpretable tokens), and the model learns cross-modal interactions through
masked-reconstruction objectives. The same backbone is used for unimodal and
multimodal training and supports survival prognosis and cancer-type stratification.
- Highlights
- Repository structure
- Installation
- Data preparation
- Quickstart
- Training recipes
- Configuration (Hydra)
- Outputs and experiment tracking
- Hyperparameter studies and batch runs
- Troubleshooting
- Citation
- License
- One backbone, two regimes. Train a single modality (RNA-seq or WSI) or fuse several (RNA-seq + DNA methylation + WSI) with the same entrypoint.
- GO-guided omics tokenization. RNA-seq and DNA methylation share a process-level token space, so the full transcriptome is represented without ad-hoc gene preselection.
- Masked pretraining + downstream heads. Self-supervised MAE pretraining, then survival or stratification fine-tuning.
- Missing-modality handling in the multimodal pipeline.
- Config-driven (Hydra). Every experiment is a config plus optional command-line overrides.
mtcp/
├── run.py # main entrypoint (Hydra) for all experiments
├── run_study.py # single hyperparameter study
├── run_multistage_study.py # multi-stage study (e.g. pretrain -> fine-tune)
├── run_schedule.py # scheduled / sequential runs
├── parallel_run.py # launch runs in parallel
├── requirements.txt
├── src/ # models, datasets, training loops, configs
│ └── configs/ # Hydra configs (base/, model/, ...)
├── sources/ # supporting assets / scripts
├── notebooks/ # exploratory and analysis notebooks
└── README.md
Tip: configuration groups live under
src/configs/(for examplesrc/configs/base/rna_base.yamlandsrc/configs/model/rna_mae.yaml). Looking through this folder is the fastest way to see every available option.
Prerequisites
- Python 3.10+
- A CUDA-capable GPU (training is GPU-first)
- libvips system library (required by
pyvipsfor reading whole-slide images)
# 1) system dependency for WSI reading (Debian/Ubuntu)
sudo apt-get update && sudo apt-get install -y libvips
# macOS: brew install vips
# conda (any): conda install -c conda-forge libvips
# 2) clone
git clone https://github.com/maryjis/mtcp.git
cd mtcp
# 3) environment
python -m venv .venv && source .venv/bin/activate
# or: conda create -n mtcp python=3.10 && conda activate mtcp
# 4) Python dependencies
pip install -r requirements.txtPyTorch / CUDA.
requirements.txtpinstorch==2.9.1/torchvision==0.24.1. If the default wheel does not match your CUDA driver, install the matching build from https://pytorch.org first, then runpip install -r requirements.txtfor the rest.
The model is trained on multi-omics TCGA cohorts (e.g. UCEC, GBM-LGG, LUAD, BLCA,
BRCA) across up to three modalities:
| Modality | What it is | Encoder |
|---|---|---|
wsi |
whole-slide histopathology images | ViT / ViT-MAE |
rna |
RNA-seq expression | GO-tokenized transformer |
dnam |
DNA methylation | GO-tokenized transformer (shared tokens with RNA) |
Data locations, cohorts and preprocessing are set in the Hydra configs, not hard-coded. To run on your machine:
- Obtain the per-cohort modality files (e.g. from TCGA).
- Open the relevant base config (for RNA:
src/configs/base/rna_base.yaml) and point the data-path fields to your local directories.<!-- TODO(maintainer): document the exact expected folder/file layout per modality here --> - Select cohorts at run time with
base.project_ids, e.g.base.project_ids='["UCEC"]'.
Minimal end-to-end run (RNA-seq survival on a single cohort, no experiment tracking):
python run.py --config-name unimodal_config \
base.project_ids='["UCEC"]' \
base.n_epochs=50 \
base.device=cuda:0 \
base.log.logging=FalseThis trains per-fold models and writes them to outputs/models/ (see
Outputs).
The entrypoint for all experiments is run.py. Select a scenario with --config-name.
python run.py --config-name unimodal_configDefaults: base.type=unimodal, base.modalities=["rna"], base.strategy=survival,
with presets from src/configs/model/rna_mae.yaml and src/configs/base/rna_base.yaml.
Common overrides:
python run.py --config-name unimodal_config base.n_epochs=50 base.device=cuda:0
python run.py --config-name unimodal_config base.project_ids='["UCEC"]' base.batch_size=32# survival baseline
python run.py --config-name unimodal_config_wsi_base
# MAE pretraining
python run.py --config-name unimodal_config_wsi_mae
# survival with an MAE-pretrained encoder
python run.py --config-name unimodal_config_wsi_mae_surv
# embedding variant
python run.py --config-name unimodal_wsi_embedWSI override examples:
python run.py --config-name unimodal_config_wsi_mae base.n_epochs=200 base.batch_size=32
python run.py --config-name unimodal_config_wsi_base base.device=cuda:1 base.available_gpus='[1]'python run.py --config-name multimodal_config
# additional presets:
python run.py --config-name multimodal_config_2
python run.py --config-name multimodal_config_3
python run.py --config-name multimodal_config_4Override the fused modalities and budget:
python run.py --config-name multimodal_config \
base.n_epochs=100 base.modalities='["rna","dnam","wsi"]'All settings come from a named config (--config-name) and can be overridden inline as
key=value. The most useful keys:
| Key | Meaning | Example |
|---|---|---|
base.type |
pipeline: unimodal or multimodal |
base.type=multimodal |
base.modalities |
active modalities | base.modalities='["rna","dnam","wsi"]' |
base.strategy |
objective: mae, survival, or boosting_survival (multimodal) |
base.strategy=mae |
base.project_ids |
TCGA cohorts to use | base.project_ids='["UCEC"]' |
base.n_epochs |
training epochs | base.n_epochs=100 |
base.batch_size |
batch size | base.batch_size=32 |
base.device |
training device | base.device=cuda:0 |
base.available_gpus |
visible GPUs | base.available_gpus='[1]' |
base.log.logging |
enable/disable tracking | base.log.logging=False |
Quote list values so the shell does not split them:
base.modalities='["rna","dnam"]'.
- Models. Per-fold checkpoints are saved to
outputs/models/with the suffix_split_{fold}.pth. - Metrics. Printed to the console; optionally logged to Weights & Biases or Neptune when tracking is enabled in the config.
- Disable tracking. Add
base.log.logging=Falseto any command.
Beyond run.py, several helper entrypoints are provided (Optuna is available via
hydra-optuna-sweeper):
run_study.py— run a single hyperparameter study.run_multistage_study.py— multi-stage studies (e.g. pretrain then fine-tune).run_schedule.py— run a sequence of experiments.parallel_run.py— launch multiple runs in parallel.
Inspect each script's arguments/config before launching, since they wrap run.py with a
study or scheduling layer.