Universal Cell Embedding: A Foundation Model for Cell Biology

Single-cell RNA sequencing has generated massive datasets across tissues, experiments, and species, but integrating and analyzing them remains challenging due to batch effects, species differences, and the need for extensive annotations.
Researchers introduce UCE (Universal Cell Embedding), a foundation model trained self-supervised on 36 million cells. By representing cells as "bags of RNA" ordered by genomic location and leveraging protein language models (ESM2), UCE learns a unified latent space that captures biological variation while remaining robust to experimental noise.
The model enables true zero-shot embedding of new cells and datasets without any fine-tuning or species-specific adjustments. When applied to novel species (e.g., green monkey, naked mole rat, chicken), it successfully aligns cell types to a reference atlas. The embedding space also exhibits emergent biological organization — cells cluster according to developmental lineages and show consistent identities across tissues (e.g., macrophages) — without being explicitly trained for these properties.
This Ledger Entry expands how readers think about foundation models in biology by showing that self-supervised learning on single-cell data can produce a universal embedding space that generalizes across tissues, experiments, and species — revealing emergent organizational principles such as developmental lineages and cross-tissue cell homogeneity without explicit supervision.