ResearchAI Research

New Transformer Model Generates Realistic Single-Cell Gene Expression Data

Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog· August 5, 2026 View original

Key takeaways

  • A new autoregressive transformer can generate high-fidelity single-cell gene expression data.
  • The model uses a VAE tokenizer and is trained with a cross-entropy loss.
  • Scaling laws for single-cell foundation models were identified for the first time.
  • The pretrained model has potential for predicting perturbation responses.

Who benefits

BiotechnologyPharmaceuticalsHealthcareAcademic Research

Summary

Researchers developed an autoregressive transformer with a VAE tokenizer to generate single-cell gene expression vectors, demonstrating high biological fidelity and identifying scaling laws for this type of foundation model. The study also explores its potential for predicting perturbation responses.

This research introduces a novel autoregressive transformer architecture designed for generating synthetic single-cell gene expression data. The model, which incorporates a learned quantized Variational Autoencoder (VAE) tokenizer, is trained to produce gene expression vectors that closely mimic real biological data from specific cell types. The study rigorously evaluates the model's ability to generate biologically faithful data and investigates its scaling properties by varying parameters and training data. A key finding is the identification of the first jointly-fit two-exponent scaling law and compute-optimal frontier for a single-cell foundation model. This work also suggests future applications, such as fine-tuning the pretrained model for predicting cellular responses to perturbations.

Why it matters

This research offers a powerful new tool for generating synthetic biological data, which can accelerate drug discovery, disease modeling, and personalized medicine by providing more data for analysis and experimentation.

How to implement this in your domain

  1. 1Integrate the model into bioinformatics pipelines for generating synthetic control groups in drug screens.
  2. 2Utilize the model to augment limited real single-cell datasets for training downstream machine learning models.
  3. 3Explore fine-tuning the model for specific perturbation prediction tasks relevant to drug development.
  4. 4Apply the identified scaling laws to optimize computational resources when developing similar biological foundation models.

Original post by Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog

"arXiv:2608.02961v1 Announce Type: new Abstract: We study a self-supervised generation task for single-cell gene expression vectors: given a set of vectors from a cell type, we aim to generate additional gene expression vectors of that cell type. For this task we characterize both…"

View on X

Originally posted by Aleksandr Sharipov, Yusif Mukhtarov, Igor Molybog on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses