OmicsFM
{{ sceneIndex }}{{ sceneModality }}{{ sceneReadoutRest }}COMPOMICS · VIB-UGENT · APACHE 2.0
{{ sceneAxisX }} → · {{ sceneAxisY }} ↑
OmicsFM
{{ sceneIndex }}·{{ sceneTag }}
{{ sceneTitle }}

{{ sceneText }}

Proteomics48,837 profiles Bulk transcriptomics680,216 profiles Single-cell transcriptomics4,550,106 cells
1,143 PRIDE projects · DDA & DIA

Public PRIDE proteomes, reprocessed and quality-filtered, with tissue, disease and instrument metadata recovered from the manuscripts.

ARCHS4 · 7,684 GEO series

ARCHS4 bulk RNA-seq profiles projected onto the shared protein vocabulary, so every gene can be compared to its protein.

CELLxGENE Census · 722 cell types

CELLxGENE Census cells across 722 cell types, sampled inversely to cell-type frequency so rare types are seen as often as common ones.

01 / 03

What is OmicsFM?

Public repositories hold tens of thousands of omics experiments. Put together, they can reveal biological structure that no single study shows.

For transcriptomics this has already happened. Atlases with millions of single cells let foundation models such as scGPT, Geneformer and scPRINT learn what genes and cells look like, simply by hiding part of each expression profile and reconstructing it from the rest. Those representations now power cell-type annotation, integration and gene-network inference.

But a transcript is not a protein. Between the two lie translation, folding, modification, trafficking, complex assembly and degradation, so mRNA is often an unreliable proxy for the proteins that actually do the work. Proteomics measures those proteins directly. Yet no foundation model existed for it, because the public proteomics data were raw spectra with scattered metadata, not analysis-ready matrices.

OmicsFM changes that. We reprocessed 1,397 public PRIDE projects into 48,837 quality-filtered protein abundance profiles, rebuilt their sample metadata, and pretrained a transformer on them with the same trick: hide some proteins, predict them from the others. No labels needed. We then trained the identical architecture on bulk RNA-seq (ARCHS4) and single-cell RNA-seq (CELLxGENE Census), so the three modalities can be compared head-to-head. Despite far fewer training profiles, the proteomics model matches or outperforms its transcriptomics counterparts on most tasks. Our preprint is out: for the full story, press Read the paper in the top-right corner.

This website has one aim: to let anybody explore the biological structure the model learned on its own. During training it saw zero annotations. No tissue labels, no protein families, no pathways, no interactions; only abundance values. Every label you can toggle in the explorers (tissue, disease, protein family, curated complex membership) is an overlay drawn from external sources such as PRIDE metadata, CORUM, Gene Ontology, STRING and Reactome, placed on top of the model’s representations after the fact. Where an overlay lines up with the structure, the model found it by itself; where it does not, that is worth a look too.

Three biologically meaningful representations fall out of pretraining. This site lets you explore each of them.

01 · SamplesSST
One vector per sample.

The sample summary token (SST) condenses an entire abundance profile into a single vector. Without ever seeing a tissue label, these vectors cluster samples by tissue across independent studies, and keep more biological structure than alternative representations.

→ Clustering · nearest-neighbour search · tissue separation
02 · AttentionNetwork
Which proteins the model relies on.

Attention weights show which proteins the model uses to predict which others. These networks recover known relationships, from complexes to functional interactions, and clustering them exposes pathway-level organisation. They are context-dependent: a protein’s partners change between tissues, matching known condition-specific interactions.

→ Partner discovery · pathway context · tissue comparison
03 · ProteinsIdentity
One learned vector per protein.

During training the model learns an identity embedding for every protein, shaped only by co-occurrence and co-variation. These embeddings predict protein properties such as gene essentiality, and combining them across proteomics, RNA and protein-language models (ESM-C) predicts better than any single source.

→ Family discovery · essentiality · cross-modality comparison
02 / 03

What can I do here?

03 / 03

One architetcure , three modalities.

Each modality learns its own identity space for the same 20,272 proteins. Pick a family to see where each model puts it. The three maps are separate embeddings, not one aligned space.

Proteomics{{ statProt }}
UMAP 1 → · UMAP 2 ↑
Bulk transcriptomics{{ statBulk }}
UMAP 1 → · UMAP 2 ↑
Single-cell transcriptomics{{ statSc }}
UMAP 1 → · UMAP 2 ↑