Mining Biarylitides

Reflecting work in the Crüsemann and Cryle Groups

Published here August 16, 2026

A Machine Learning-Based Genome Mining Approach Reveals Unprecedented Biarylitide Diversity

Leo Padva, Jemma Gullick, Friederike Biermann, Laura J. Coe, Yongwei Zhao, Sam Tucker, Lukas Zimmer, Julien Tailhades, Stefan Kehraus, Ralf B. Schittenhelm, James J. De Voss, Eric J.N. Helfrich, Max J. Cryle, and Max Crüsemann

JACS Au 2026, 6, 3952–3965. https://doi.org/10.1021/jacsau.6c00500

View Original Publication


Biarylitides are ribosomally synthesized and post-translationally modified peptides, RiPPs, whose precursor genes span a mere 18 base pairs, the shortest known coding sequences across the tree of life. That extreme miniaturization is also the central obstacle to discovering new members of the family: conventional BLAST searches miss distantly related cytochrome P450 cross-linking enzymes, PSI-BLAST accumulates false positives, and automatic gene annotation pipelines cannot detect an open reading frame this small. The result has been a systematic undercount of biarylitide biosynthetic gene clusters, BGCs, leaving the true breadth of precursor sequence diversity, cross-link chemistry, and co-encoded tailoring enzymes largely unknown. A broader map of the family would expose new P450 substrate tolerance, new building-block chemistry for antibiotic synthesis, and new avenues into organisms, including human-associated bacteria, whose biarylitide capacity has never been recognized.

Researchers in the Crüsemann and Cryle Groups at Goethe University Frankfurt and Monash University, Clayton, Australia, published in JACS Au, adapted AtropoFinder, a random-forest genome-mining pipeline originally built for atropopeptide BGCs, into a dedicated BiarylitideFinder. The classifier was trained on a curated set of 106 confirmed or strongly predicted biarylitide P450 sequences against nearly 10,000 negative examples, and a companion CoreFinder searched the 3 kb genomic neighborhood of each P450 hit for open reading frames of five to eight amino acids carrying aromatic residues at positions 3 and 5, the signature of a biarylitide core peptide. Applied to the full NCBI GenPept database, the workflow identified 277 unique BGCs, including 124 not detected in any prior study, and extended the known phylogenetic range of biarylitide producers to human oral microbiome bacteria in the genus Rothia. To validate the bioinformatic predictions, the team used in vitro cross-linking assays, deuterium-labeled substrates, and multi-dimensional NMR spectroscopy to characterize P450 enzymes associated with three previously uncharacterized core peptide motifs, confirming distinct cross-link regiochemistries for each and showing that substrate tolerance at the central pentapeptide position correlates with the type of biaryl bond formed.

The work demonstrates that machine-learning-assisted genome mining can recover biosynthetic diversity that sequence-similarity methods systematically miss when genes are this short and P450s this divergent. The 277-BGC landscape, with its newly mapped precursor motifs and extended producer phylogeny, provides a structured foundation for exploring biarylitide cross-link chemistry and its potential in biocatalytic production of cross-linked tripeptide antibiotic precursors. Full experimental details, NMR assignments, and the complete BGC table are in the original publication.

Mining Biarylitides

Figure 1. Structures and biosynthetic models of selected side-chain cross-linked peptide natural products. A| Structural diversity of biarylitides with cross-links shown in red, including the native examples of P450Blt, Micromonospora sp. MW-13, BytO, BytO2, SlyP, SgrB, and ShyB, as well as non-native examples from in vitro studies with P450Blt and ShyB. B| RiPP biosynthetic pathway for tryptorubin A, numbers indicate order of enzyme activity. C| RiPP biosynthetic pathway for the biarylitides, numbers indicate order of enzyme activity, example shown for P450Blt.