Completed
2023 · Transcriptomics

Cancer Immunotherapy Machine Learning Challenge

Cancer immunotherapy seeks to harness the body’s immune system, and most often T cells, to recognize and kill cancer cells while leaving healthy cells alone.
Learn more
Overview

In the last decade, there have been many breakthroughs in cancer immunotherapy, yet treatments still only work for some cancer patients some of the time. To address this gap, the Schmidt Center held a machine learning competition in January 2023, in which participants developed algorithms to uncover new ways to modify, or “perturb,” T cells to make them better cancer-cell killers. The challenge centered on Perturb-seq measurements that connect CRISPR genetic perturbations with single-cell transcriptional state changes, giving teams a way to reason about which perturbations might push T cells toward more effective anti-tumor behavior. Scientists in the Hacohen Lab at the Broad then tested their predictions in mouse models, making this the first challenge that the Schmidt Center knows of in which new experiments were performed based on the output of machine-learning models developed in the challenge. More than 1,000 people from 87 countries registered for the competition.

CHALLENGE 1 DATA

Includes all 73 perturbations for training / validation / testing

Can be accessed at GSM9664310

CHALLENGE 2 DATA

Includes the experimental validation for the participant-proposed 59 perturbations

Can be accessed at GSM9664311

ADDITIONAL DATA

All raw and processed data can be accessed at GSE327731

REFERENCE

More information, including detailed descriptions of the experiments, analyses, and code, can be found in the paper (DOI: 2026.05.21.726863). If you use the dataset in your research, please cite this work.

Challenge 1

More information, including the full archive, code, and writeup of each method below, can be found in the paper.

Challenge 2
Challenge 3

Starter Kits, Baselines, Docs

Core Challenge Materials
  • Read about the challenge specifications at TopCoder
  • Read about the design, analyses, and results in the paper (DOI: 2026.05.21.726863)

Public Reference Resources

Publicly available datasets, models, ontologies, and tooling that may be useful for exploration and external benchmarking.

Gene Mapping & Ontology Resources
  • BiomaRt: tool for mapping homologous genes between mouse and human, useful when integrating human and mouse T-cell studies.
  • Gene Ontology Resource: knowledge base of gene function (molecular function, cellular component, biological process) that can be used to build gene embeddings and functional priors.
Multi-Omic & Epigenetic Tooling
  • MuData / muon: Python frameworks extending AnnData/scanpy for multimodal analysis (RNA + ATAC from the same cells).
  • ArchR: R package for large-scale single-cell ATAC-seq analysis and integration with gene expression.
  • chromVAR: R package for analyzing single-cell ATAC-seq bias and transcription factor motif variation.
  • JASPAR: database of transcription factor binding profiles, useful for linking open chromatin peaks to specific regulators.
Perturbation Modeling Methods & Benchmarks
Single-Cell Atlases & Data Repositories
  • Human Cell Atlas Data Portal: large-scale human single-cell datasets for building reference atlases and cross-tissue comparisons.
  • Tabula Sapiens: human transcriptome reference at single-cell resolution across multiple organs.
  • Tabula Muris: mouse transcriptome reference at single-cell resolution.
  • Jingle Bells: repository of standardized single-cell RNA-seq datasets for analysis and visualization.