Sample Submission Guidelines Inquiry
Request a Quote

Reverse Metabolomics Service

Reverse metabolomics is a big-data discovery strategy that uses the MS/MS spectra of synthesized molecules as search queries against public metabolomics repositories, revealing which organs, disease states, and species carry those molecules. Where conventional metabolomics annotates the compounds in your samples, reverse metabolomics starts from the molecule and asks where it appears across hundreds of thousands of published datasets.

MS/MS query against 1.2 billion public spectra

Phenotype, organ, and species association mapping

Discovery of previously undescribed metabolites

Microbiome-derived metabolite investigation

Complementary to conventional annotation

Reverse metabolomics MS/MS data mining for metabolite discovery

Reverse Metabolomics — MS/MS Data Mining for Metabolite Discovery

Reverse metabolomics is a discovery framework that acquires MS/MS spectra for molecules of known structure and searches those spectra against public untargeted metabolomics repositories — using tools such as MASST and ReDU — to identify the organisms, organs, disease states, and sample types associated with each molecule.

Conventional untargeted metabolomics runs from sample to annotation: you collect samples, acquire LC-MS/MS data, and match features against known-metabolite databases. Reverse metabolomics inverts the flow. You begin with a molecule — often one synthesized for this purpose — and ask where its characteristic MS/MS fingerprint already exists across the global metabolomics archive. A molecule whose spectrum is found in hundreds of human datasets but was never annotated before is, in effect, a newly discovered biological molecule.

Reverse metabolomics delivers:

  • Phenotype association: which health or disease states, interventions, and sample types carry a given molecule
  • Organ and species distribution: where a molecule appears across human, rodent, and other datasets
  • New-metabolite discovery: molecules present in public data but previously unannotated in commercial databases
  • Microbiome attribution: which microbial producers and host-microbiome axes generate a molecule

Forward vs. Reverse Metabolomics — Two Complementary Discovery Paths

The two approaches answer different questions and are best used together:

Forward Metabolomics vs. Reverse Metabolomics

DimensionForward MetabolomicsReverse Metabolomics
Starting pointYour biological samplesA molecule of known structure
Data flowSample → acquisition → annotationSpectrum → public repository search → association
Primary question"What is in my sample?""Where does this molecule appear?"
OutputAnnotated feature list from your cohortPhenotype and organ map across global data
Discovery powerLimited to existing database entriesReveals previously unannotated molecules

For the experimental MS/MS data that reverse metabolomics queries, our untargeted metabolomics service provides high-resolution acquisition from your own samples — the two workflows read the same data from complementary directions.

Reverse Metabolomics Platform and Technical Parameters

Our reverse metabolomics workflow combines in-house MS/MS acquisition with public-repository search tooling:

ParameterSpecification
MS/MS acquisitionHigh-resolution LC-MS/MS for molecules of interest
In silico explorationPredicted MS/MS spectra for hypothesis screening before synthesis; custom synthesis support for candidate validation
Public spectrum searchMASST — searches MS/MS fingerprints against ~1.2 billion public spectra
Metadata filteringReDU — links matched datasets to organism, disease, biospecimen, and phenotype
Microbial strain attributionmicrobeMASST — links matched spectra to 60,000+ cultured microbial strains for species-level producer attribution
Network frameworkGNPS-based molecular networking for producer and co-occurrence analysis
Query formulationMassQL — mass-spectrometry query language for targeted structure classes
ConfidenceLevel 2–3 annotation with spectral match and retention alignment
DeliverablesAssociation tables, organ/phenotype heatmaps, dataset matches
Thermo Orbitrap Exploris 480-class LC-MS/MS system

Thermo Orbitrap Exploris 480-class

High-resolution LC-MS/MS for synthesized and reference molecules

Bioinformatics platform for public repository search

Repository Search Platform

MASST and ReDU queries against public data

Target Metabolite Classes for Reverse Discovery (Bile Acids, N-Acyl Amides, Lipids)

Reverse metabolomics has proven especially powerful for compound classes that are underrepresented in commercial databases but abundant in biological data:

Compound classExamples and notes
Bile acids and amidatesConjugated bile acids — microbially modified species linked to gut health and disease
N-acyl amidesFatty acid–amino acid conjugates with signaling roles in the microbiome-host axis
Fatty acid estersHydroxy fatty acid esters and related lipid classes
Microbiome metabolitesHost-microbiome derived molecules with limited database coverage
Xenobiotic derivativesDrug and environmental compound metabolites

Reverse Metabolomics Workflow — A Step-by-Step Guide

1

Molecule selection and spectra

We select molecules of interest and acquire their MS/MS spectra — from synthesized standards or reference compounds — as search queries.

2

Query formulation

Each spectrum is encoded as a Universal Spectrum Identifier (USI) and structured into MASST search queries.

3

Public repository search

MASST searches the MS/MS fingerprints against ~1.2 billion public spectra to find matching datasets.

4

Metadata linkage

ReDU links matched files to organism, disease state, biospecimen, phenotype, and other sample descriptors.

5

Association validation

Spectral matches pass quality thresholds before association — cosine similarity ≥ 0.7, a minimum of 4–6 matched fragment peaks, and mass tolerance ≤ 10–15 ppm. Validated associations are then checked across independent cohorts and, where applicable, confirmed by authentic standards.

6

Report delivery

Delivery of association tables, organ and phenotype heatmaps, dataset matches, and an interpretation summary.

Reverse metabolomics workflow from spectra to phenotype association

Input Requirements: Reference Compounds, Synthesized Molecules, and Spectra Files

Reverse metabolomics can work from either existing LC-MS/MS data or new samples for acquisition:

InputRequirements and notes
Existing MS/MS dataPreviously acquired LC-MS/MS files in standard format (mzML, mzXML) can be queried directly
New sample acquisitionBiofluids, tissues, or microbial cultures for fresh MS/MS acquisition
Reference compoundsStandards or synthesized molecules for which MS/MS spectra are to be generated
Minimum inputReverse search requires only a spectrum — no large sample cohort is mandatory
MetadataStudy context (organism, condition, intervention) improves association interpretation

Why Choose Our Reverse Metabolomics Service

  • Big-data discovery, not just annotation
    Reverse metabolomics asks where a molecule appears globally, revealing biological roles that sample-level annotation misses.
  • Public repository expertise
    We work with MASST, ReDU, and GNPS tooling to mine the largest public metabolomics archives.
  • New-molecule discovery
    Molecules present in public data but absent from commercial databases can be surfaced as novel biological findings.
  • Microbiome-host attribution
    Reverse search resolves which metabolites are microbially derived and where they associate with disease.
  • Complementary acquisition
    Pair reverse discovery with untargeted metabolomics for the full forward and reverse picture.

Reverse Metabolomics Data Deliverables and Association Evidence

You receive a complete, interpretation-ready reverse metabolomics result set:

  • Association tables linking each query molecule to matched datasets, organisms, and phenotypes
  • Organ and species distribution maps as heatmaps across human and rodent data
  • Producer and co-occurrence analysis from GNPS molecular networking
  • Dataset match reports listing the specific public studies where each molecule was found
  • Interpretation summary translating matches into biological hypothesis and follow-up actions
MS/MS spectrum matched to public repository datasets

Representative MS/MS spectrum of a query molecule with matched public repository datasets.

Organ and phenotype association heatmap from reverse metabolomics

Organ and phenotype association heatmap revealing where a molecule appears across public data.

Applications

  • Discovery biology — surface previously unannotated molecules and their phenotype associations from global data
  • Microbiome-host research — attribute metabolites to microbial producers and trace host-microbiome axes
  • Natural product research — locate known-structure molecules across datasets to prioritize bioactive candidates
  • Drug metabolism — search for drug and metabolite fingerprints across public data to map exposure and response
  • From discovery to validation — candidate biomarkers surfaced in public data can be carried into targeted MRM/PRM absolute quantification in your own cohort, closing the loop from repository discovery to validated measurement.

Case Study: First Identification of Cholestenoic Acid as an Endogenous Epigenetic Regulator

Cholestenoic acid as endogenous epigenetic regulator decreases hepatocyte lipid accumulation in vitro and in vivo

Wang, Y., Pandak, W. M., Hylemon, P. B., Min, H.-K., Min, J., Fuchs, M., et al. | American Journal of Physiology-Gastrointestinal and Liver Physiology, 2024, 326(2), G147–G162

DOI: 10.1152/ajpgi.00184.2023


Background

Hepatic lipid accumulation is a hallmark of metabolic liver disease, yet the endogenous small molecules that regulate hepatocyte lipid metabolism through epigenetic mechanisms remain incompletely characterized.

Challenge: Identify a previously uncharacterized endogenous bile acid species and define its regulatory role in hepatocyte lipid metabolism.


Analytical Approach

Untargeted lipidomics was performed at Creative Proteomics (New York) on HepG2 cells treated with cholestenoic acid, coupled with transcriptomic and epigenomic analysis. This study was, to the authors' knowledge, the first to identify the mitochondrial monohydroxy bile acid cholestenoic acid as an endogenous epigenetic regulator of lipid metabolism.


Key Findings

MetricFinding
New endogenous regulatorCholestenoic acid identified as an endogenous epigenetic regulator of lipid metabolism
Lipid accumulationCholestenoic acid decreased hepatocyte lipid accumulation in vitro and in vivo
Lipidomic readoutUntargeted lipidomics at Creative Proteomics profiled the metabolic response
MechanismEpigenome modification linked to global lipid metabolism regulation

What This Means for Your Reverse Metabolomics Study

  • New endogenous molecules drive biology. This study identified a previously uncharacterized bile acid as a functional regulator — the kind of molecule reverse metabolomics surfaces systematically by searching public data.
  • Bile acid biology is rich in unannotated species. The bile acid family is a prime reverse metabolomics target, where synthesized standards reveal species that conventional databases miss.
  • Lipidomics provides the discovery readout. Untargeted lipidomics generated the data layer that resolved the new regulator — the same acquisition we pair with reverse search.

Conclusion

This study shows how identifying a previously uncharacterized endogenous metabolite can reveal a regulatory mechanism. Our reverse metabolomics service extends this capability at scale — searching the MS/MS fingerprints of known-structure molecules across public data to surface new biological molecules and their phenotype associations.

What is reverse metabolomics?

Reverse metabolomics is a discovery strategy that acquires MS/MS spectra for molecules of known structure and searches those spectra against public untargeted metabolomics repositories to identify the organisms, organs, disease states, and sample types associated with each molecule.

How does reverse metabolomics differ from conventional metabolomics?

Conventional metabolomics runs sample-to-annotation — you acquire data from your samples and match features to databases. Reverse metabolomics runs molecule-to-phenotype — you start with a molecule's spectrum and ask where it appears across global public data.

What tools are used for reverse metabolomics?

MASST searches MS/MS fingerprints against roughly 1.2 billion public spectra; ReDU links matches to metadata such as organism, disease, and biospecimen; GNPS supports molecular networking and MassQL formulates targeted structure-class queries.

Can reverse metabolomics find new metabolites?

Yes. Molecules present in public data but absent from commercial databases can be surfaced as previously unannotated biological molecules. In published work, reverse metabolomics identified over a hundred bile acid species not previously described.

Do I need to synthesize compounds first?

Molecules of known structure — synthesized standards or reference compounds — provide the MS/MS queries. We work with you to select or source the molecules of interest; you do not need to have synthesized them yourself.

What sample or data input is required?

Reverse search requires only a spectrum. You can provide existing LC-MS/MS files, new samples for acquisition, or reference compounds for which we generate spectra.

Which metabolite classes work best?

Bile acids and amidates, N-acyl amides, fatty acid esters, microbiome metabolites, and xenobiotic derivatives are especially productive because they are abundant in biological data yet underrepresented in commercial databases.

How are associations validated?

Candidate associations are checked across independent public cohorts and, where applicable, confirmed by authentic standards to distinguish true biological signal from chance matches.

Can reverse metabolomics be combined with other services?

Yes. Pair reverse discovery with our untargeted metabolomics service for the full forward and reverse picture — acquisition from your samples plus association mapping across public data.

What deliverables do I receive?

Association tables, organ and species distribution heatmaps, producer and co-occurrence analysis, dataset match reports, and an interpretation summary with follow-up recommendations.

Publications

Bacteroides thetaiotaomicron enhances H2 metabolism in the gut

Davies, J., Mayer, M. J., Juge, N., et al.

Journal: Gut Microbes, 2024, 16(1)

Untargeted metabolomics of a gut commensal bacterium, resolving metabolic contributions relevant to the microbiome-host axis that reverse metabolomics queries at scale.

Anxiety-like behavior during protracted morphine withdrawal is driven by gut microbial dysbiosis and attenuated with probiotic treatment

Oppenheimer, M., Tao, J., Moidunny, S., et al.

Journal: Gut Microbes, 2025, 17(1)

Untargeted metabolomics linking gut microbial dysbiosis to behavioral phenotype, demonstrating the microbiome-derived metabolite associations that reverse search maps across public data.

A human iPSC-derived hepatocyte screen identifies compounds that inhibit lipid accumulation

Liu, J.-T., Doueiry, C., Jiang, Y.-l., et al.

Journal: Communications Biology, 2023, 6(1)

Untargeted metabolomics of hepatocytes resolving lipid-related metabolic changes, the bile acid and lipid biology where reverse metabolomics surfaces unannotated species.

For Research Use Only. Not for use in diagnostic procedures.
inquiry

Get Your Custom Quote

Connect with Creative Proteomics Contact Us Contact Us
return-top