
In the last decade, there have been many breakthroughs in cancer immunotherapy, yet treatments still only work for some cancer patients some of the time. To address this gap, the Schmidt Center held a machine learning competition in January 2023, in which participants developed algorithms to uncover new ways to modify, or “perturb,” T cells to make them better cancer-cell killers. The challenge centered on Perturb-seq measurements that connect CRISPR genetic perturbations with single-cell transcriptional state changes, giving teams a way to reason about which perturbations might push T cells toward more effective anti-tumor behavior. Scientists in the Hacohen Lab at the Broad then tested their predictions in mouse models, making this the first challenge that the Schmidt Center knows of in which new experiments were performed based on the output of machine-learning models developed in the challenge. More than 1,000 people from 87 countries registered for the competition.






Challenge 1 focused on predicting how specific CRISPR perturbations reshaped the distribution of T cell states within the tumor microenvironment. The underlying dataset consisted of 71,388 tumor-infiltrating T cells, including 4,978 unperturbed cells and 26,031 perturbed cells across 73 single-gene knockouts. After stringent preprocessing to remove doublets and cells with ambiguous guide assignments, the training data included all unperturbed cells and 23,719 perturbed cells corresponding to 66 of the 73 knockouts.
The remaining seven knockouts were withheld for evaluation: Aqr, Bach2, and Bhlhe40 for validation, and Ets1, Fosb, Mafk, and Stat3 for testing. For each cell, participants received a normalized 15,077 gene expression vector, its perturbation condition, and an expert-annotated T cell state label chosen from progenitor, effector, terminal exhausted, cycling, and other. The task was to build a model that, given only the name of a held-out gene knockout and not its expression profile, predicted a five-dimensional probability vector representing the proportions of cells in these five states, with entries summing to one.
Challenge 2 extended beyond the original set of 73 knockouts and asked participants to propose novel CRISPR targets among all 15077 measured genes that were predicted to induce T cell state distributions favorable for cancer immunotherapy. Building on their Challenge 1 models, participants developed algorithms that, given the name of any gene among the 15077 measured genes passing quality control, output a predicted five-state cell proportion vector representing the effect of knocking out that gene.
These predictions underpinned three related sub-tasks corresponding to different therapeutic strategies, and top-ranked perturbations were experimentally validated in vivo in a mouse tumor model. In Challenge 2A, focused on immune checkpoint blockade, the goal was to identify knockout targets that maximized the proportion of progenitor T cells while ensuring that at least five percent of cells were in the cycling state to maintain sufficient cell expansion. Conceptually, predictions were ranked by their proximity, in L1 distance, to the idealized cell state vector (1, 0, 0, 0, 0), corresponding to 100% progenitor cells.
Challenge 3 moved beyond predicting cell state outcomes and instead focused on how perturbations should be scored and prioritized for downstream applications in cancer immunotherapy. Participants were asked to design new metrics that more effectively evaluated the ability of a gene knockout to move cells from undesired to desired states. Formally, let P0 denote the empirical gene expression distribution of unperturbed cells and Pi the distribution resulting from knocking out gene i. Participants proposed a statistic s(·) that summarized Pi (for example, but not limited to, cell state proportions) and a scoring function that combined P0, an optional desired distribution Q, and the predicted statistic Ŝ(Pi) to assign a score to each perturbation, with higher scores indicating more desirable knockouts.
The design had to balance biological richness with statistical feasibility: while higher-dimensional statistics can be more informative, they require larger sample sizes to estimate reliably, whereas lower-dimensional summaries may be easier to predict but less expressive.
Includes all 73 perturbations for training / validation / testing
Can be accessed at GSM9664310
Includes the experimental validation for the participant-proposed 59 perturbations
Can be accessed at GSM9664311
All raw and processed data can be accessed at GSE327731
More information, including detailed descriptions of the experiments, analyses, and code, can be found in the paper (DOI: 2026.05.21.726863). If you use the dataset in your research, please cite this work.
More information, including the full archive, code, and writeup of each method below, can be found in the paper.
Participant(s): Marios Gavrielatos and Konstantinos Kyriakidis
Details: Greece
Topcoder: Konstantinos Kyriakidis
Full archive →
Code →
Wiriteup →
Participant(s): Yuzhou Gu, Anzo Teh, Yanjun Han, and Brandon Wang
Details: U.S. (MIT)
Topcoder: Yuzhou Gu
Full archive →
Code →
Wiriteup →
Participant(s): Peter Novotný
Details: Poland
Topcoder: Peter Novotný (link: https://profiles.topcoder.com/nofto)
Full archive →
Code →
Wiriteup →
Participant(s): Thu Huyen Nguyen
Details: N/A
Topcoder: Thu Huyen Nguyen
Full archive →
Code →
Wiriteup →
Participant(s): Yul Young Park
Details: N/A
Topcoder: Yul Young Park
Full archive →
Code →
Wiriteup →
Participant(s): Elizabeth Hudson
Details: N/A
Topcoder: Elizabeth Hudson
Full archive →
Code →
Wiriteup →
Participant(s): NA
Details: N/A
Topcoder: _mahcih_
Full archive →
Code →
Wiriteup →
Participant(s): Dean Ninalga
Details: N/A
Topcoder: Dean Ninalga
Full archive →
Code →
Wiriteup →
Participant(s): Roman Piankov
Details: N/A
Topcoder: Roman Piankov
Full archive →
Code →
Wiriteup →
Participant(s): Siqi Liu
Details: N/A
Full archive →
Code →
Wiriteup →
Participant(s): Brody Langille, Jordan Trajkovski, and Elizabeth Hudson
Details: N/A
Topcoder: Brody Langille, Elizabeth Hudson
Full archive →
Code →
Wiriteup →
Participant(s): Marc Glettig
Details: Archive folder: marc_glettig
Topcoder: Marc Glettig
Full archive →
Code →
Wiriteup →
Participant(s): Ai Vu Hong
Details: Researcher at Genethon, France
Topcoder: Ai Vu Hong
Full archive →
Code →
Wiriteup →
Participant(s): Saket Kunwar
Details: Independent researcher, Nepal
Topcoder: Saket Kunwar
Full archive →
Code →
Wiriteup →
Participant(s): Xiang L
Details: N/A
Topcoder: lxastro0
Full archive →
Code →
Wiriteup →
Participant(s): John Gardner
Details: Freelance data scientist; Tecumseh, United States
Topcoder: John Gardner
Full archive →
Code →
Wiriteup →
Participant(s): Agilidade S
Details: N/A
Topcoder: Agilidade S
Full archive →
Code →
Wiriteup →
Participant(s): Basak Eraslan
Details: Postdoctoral researcher at the Regev Lab in Genentech and Kundaje Lab at Stanford University
Topcoder: Basak Eraslan
Full archive →
Code →
Wiriteup →
Participant(s): Zhengmei L
Details: N/A
Topcoder: Zhengmei L
Full archive →
Code →
Wiriteup →
Participant(s): Liu Xindi
Details: Freelance programmer
Full archive →
Code →
Wiriteup →
Participant(s): Haoyue Dai, Kun Zhang, Ignavier Ng, Yujia Zheng, Xinshuai Dong, Yewen Fan, Petar Stojanov, Gongxu Luo, and Biwei Huang
Details: 1. Haoyue Dai, Kun Zhang, Ignavier Ng, Yujia Zheng, Xinshuai Dong, and Yewen Fan: Carnegie Mellon University. 2. Petar Stojanov: Eric and Wendy Schmidt Center. 3. Gongxu Luo: Mohamed bin Zayed University of AI. 4. Biwei Huang: UC San Diego.
Full archive →
Code →
Wiriteup →
Participant(s): Johnson Zhou, Camille Sayoc, and Yi-Cheng Peng
Details: University of Melbourne, Australia
Topcoder: Johnson Zhou
Full archive →
Code →
Wiriteup →
Participant(s): Dariusz Brzeziński and Wojciech Kotlowski
Details: Poznań University of Technology, Poland
Topcoder: Dariusz Brzeziński
Full archive →
Code →
Wiriteup →
Participant(s): Salil Bhate
Details: Postdoctoral fellow at the Eric and Wendy Schmidt Center
Topcoder: Salil Bhate
Full archive →
Code →
Wiriteup →
Participant(s): Irene Bonafonte Pardàs, Artur Szalata, Benjamin Schubert, and Miriam Lyzotte
Details: Helmholtz Center Munich and Mila - Quebec AI Institute
Topcoder: Irene Bonafonte Pardàs
Full archive →
Code →
Wiriteup →
9 Short background lectures and dataset walkthroughs from the challenge series.
Publicly available datasets, models, ontologies, and tooling that may be useful for exploration and external benchmarking.