Protein language models · 46 reference strains

Which genes can a bacterium live without?

Upload one genome. ProteomeFit calls every coding sequence essential or non-essential, then predicts how much each non-essential knockout changes growth, with a confidence score for every call.

Start a prediction
  1. 1
    Gene set
    GenBank CDS, CDS FASTA, or complete-genome FASTA (Prodigal)
  2. 2
    Embeddings
    ESM C 600M per protein, ProteomeLM-L across the whole proteome
  3. 3
    Essentiality
    Classifier probability, informed by orthologs in 46 Fitness Browser strains
  4. 4
    Fitness
    1,000-member ensemble regression for non-essential genes

Submit a genome

One strain per job. Every gene is scored in the context of the full proteome, so upload the complete gene set.

Input format

Good to know

  • A typical genome takes one to three minutes. When the server is busy, yours waits for earlier predictions to finish.
  • Keep this page open until the result appears. Nothing is stored on the server: leaving the page cancels the prediction, so download the CSV when it is done.
  • Each network address can submit 5 jobs per day.
  • Protein-only FASTA is not accepted. Codon features need nucleotide sequences.

Method

Stage 1 · Essentiality

An attention-pooled ESM C token representation, the raw sequence, the ProteomeLM context embedding and ortholog-derived context features from 46 reference strains feed a multilayer perceptron. Genes with probability ≥ 0.5 are called essential; their fitness columns are left blank.

Stage 2 · Fitness

The same inputs plus Biopython sequence and codon features feed 1,000 bootstrapped mini-MLPs. Their mean is the predicted fitness (z-scored within the strain), and their spread gives the uncertainty and the fitness confidence (conf_fit).

Output columns

seq_dnaCoding sequence (nucleotides, including the stop codon) used for the prediction.
callessential if conf_ess ≥ 0.5, otherwise non-essential.
fitness_zPredicted knockout fitness, z-scored within the strain (0 = strain average, negative = growth defect). Blank for essential genes.
fitness_sdStandard deviation across the 1,000 ensemble members.
pred_t|fitness_z| / fitness_sd.
conf_essConfidence of the essentiality classification: the predicted probability that the gene is essential (every gene).
conf_fitConfidence of the fitness prediction for non-essential genes: 2Φ(pred_t) − 1, from 0 (no evidence the fitness differs from the strain average) to 1. Blank for essential genes.