Every human cell contains the same ~20,000 genes, yet the regulatory networks connecting these genes drive the extraordinary diversity of cellular behaviors. Modern perturbation technologies such CRISPR, combined with single-cell sequencing and high-content imaging, now allow us to systematically probe these networks at scale. By observing how cells respond to perturbations, we can begin to map the rules that govern cell identity, plasticity, and fate transitions.
The Cell Perturbation Prediction Challenge (CPPC), organized by the Eric and Wendy Schmidt Center at the Broad Institute, brings together high-throughput experimental data generation and innovative machine learning methods to advance this goal. Our mission is to help build a computation-informed foundation for understanding the ways in which perturbations reshape cell states across modalities, contexts, and scales.

Integrating diverse, unpaired, multi-modal data to enable the prediction of missing modalities and translating between different cell lines and species;

Guiding the experimental design by predicting the outcome of unseen perturbations and identifying, through active learning, the most informative perturbations to test experimentally;

Identifying regulatory (causal) networks within and among cells.
In 2023, the Eric and Wendy Schmidt Center at the Broad Institute launched the first Cancer Immunotherapy Machine Learning Challenge, providing a single-cell Perturb-seq dataset covering 73 genetic perturbations and experimentally validating 61 additional predictions in T cells transferred into a mouse melanoma model to evaluate their potential to enhance anti-tumor activity for cancer immunotherapy. The challenge revealed promising directions while highlighting the need for sustained, community-driven progress.
Going forward, CPPC will evolve into a continuous, large-scale, therapeutically relevant benchmark for perturbation prediction, much like
ImageNet for computer vision or
CASP for protein structure prediction. Each year, we will release new experimental datasets aligned with our flagship projects, enabling objective measurement of progress and accelerating biological discovery.