AI Metabolomics Analysis — Why Deep Learning Changes the
Pipeline
Untargeted metabolomics is powerful, but its standard software pipeline leaves value on the table at four
well-known bottlenecks. Automated peak pickers over-detect noise, producing false-positive features that
corrupt downstream statistics. Spectral libraries cover only a fraction of the metabolome, so most detected
features stay unannotated. Classical statistics handle linear relationships but struggle with the nonlinear,
high-dimensional structure of real cohorts. And single-omics correlation analysis cannot connect a
metabolite change to its upstream genes or downstream phenotype.
Artificial intelligence (AI) changes each of these. Convolutional neural networks classify true peaks from
instrument noise at a level conventional peak-picking thresholds cannot reach. In-silico fragmentation and
structure prediction identify metabolites that exist in no spectral library. Random forest, support-vector
machines and gradient boosting select the metabolite combinations that actually separate cohorts. Knowledge
graphs and graph neural networks link metabolite changes into mechanism networks across omics layers.
The result is a metabolomics analysis pipeline that returns cleaner feature tables, deeper annotation,
interpretable predictive models, and a mechanism narrative — not just a list of changed features.
AI Signal Denoising and Deep-Learning Peak Curation
Raw LC-MS data is noisy. Conventional peak-picking algorithms over-detect: they call chromatographic noise,
low-intensity irreproducible signals, and misdefined peak boundaries as features, inflating feature tables
with false positives. These errors then propagate through alignment, statistics and annotation.
- Deep-learning peak curation — CNN-based classifiers trained on large sets of extracted
ion chromatograms distinguish true metabolite signals from instrument noise, with accuracy in the high 90s
on benchmark evaluation.
- Baseline correction and alignment — AI-assisted retention-time alignment and drift
correction keep features comparable across batches and injection order.
- Cleaner feature tables — the output is a curated feature matrix where noise peaks are
removed before statistics, so downstream models see real biology rather than artifacts.
- Reproducibility — fixed, versioned preprocessing means the same raw files give the same
curated table, which is what reviewers and collaborators need.
A clean feature table is the foundation of every later step — annotation, modeling and integration all
inherit the quality of the matrix they start from. For studies that also need spatial resolution, our spatial metabolomics services apply the same curation
workflow to mass spectrometry imaging datasets.
AI Metabolite Identification and In-Silico Structure
Annotation
The deepest bottleneck in untargeted metabolomics is the dark metabolome: in a typical study, most detected
features are absent from standard spectral libraries. Library matching alone leaves them unannotated. AI
in-silico metabolite annotation methods identify these unknowns directly from the mass spectrum.
- Molecular formula from isotope patterns and fragmentation — computational fragmentation
trees explain the observed MS/MS spectrum and infer the molecular formula and fragment relationships.
- Structural fingerprint prediction — machine learning predicts the structural features
of an unknown molecule from its spectrum, then searches that fingerprint against molecular structure
databases such as PubChem.
- Retention-time and CCS prediction — predicted retention time and, where ion mobility is
used, collision cross-section (CCS) provide independent orthogonal evidence that supports or refutes
candidate structures.
- Confidence-scored annotation — every annotation carries a confidence level, so you know
which identifications are high-confidence versus putative.
These approaches, in the class of SIRIUS and CSI:FingerID, achieve identification rates well above
library-only workflows and have been applied to resolve unknowns in complex biological samples. For
discovery-scale coverage, our untargeted metabolomics
service provides the acquisition depth, and AI annotation extends how much of that coverage is
actually identified.
| Stage |
What AI Adds |
| Molecular formula |
Isotope-pattern and fragmentation-tree inference, beyond nominal-mass matching |
| Structure |
Predicted fingerprint searched against molecular structure databases |
| Orthogonal evidence |
Retention-time and CCS prediction cross-validation |
| Confidence |
Per-feature confidence levels separating high-confidence from putative annotations |
AI Biomarker Discovery and Machine-Learning Cohort Mining
With thousands of features and complex cohorts, classical statistics can miss the structure in the data or
overfit it. Machine learning metabolomics is designed for exactly this setting — high-dimensional,
nonlinear, cohort-scale data.
- Feature selection — random forest, support-vector machines and gradient boosting
(including XGBoost) identify the metabolite combinations that best separate phenotypes, ranking features
by contribution.
- Cohort stratification — trained classifiers stratify samples into groups and support
predictive modeling and sample-level classification for research cohorts.
- Candidate biomarker panels — the top-ranked metabolites are assembled into
multi-metabolite panels with ROC analysis and cross-validation.
- Explainable AI — feature-importance and SHAP-style analysis show which metabolites
drive each model, so the model is interpretable, not a black box.
Models are validated with proper train/test splits and cross-validation, with results reported
transparently. For the complementary statistics and figure work, our metabolomics data analysis service covers classical
multivariate analysis; and where candidates move to absolute quantification, our targeted metabolomics service validates them with
isotope-labeled standards.
Every candidate panel is stress-tested against overfitting: models are evaluated on held-out test sets,
cross-validation and permutation tests, and only features that survive feature-importance and SHAP-style
analysis are carried forward. Top-ranked candidates then receive orthogonal validation — isotope-dilution
targeted LC-MS/MS absolute quantification against isotope-labeled internal standards — so what the model
identifies is confirmed as a real, quantifiable metabolite at the laboratory level, closing the loop between
computational discovery and experimental confirmation.
AI Multi-Omics Integration and Mechanism Networks
A metabolite change rarely stands alone. Connecting it to upstream gene expression, protein abundance,
microbial composition, and downstream phenotype is where mechanism lives — and where single-omics
correlation analysis runs out.
- Knowledge-graph integration — curated biological knowledge from literature and public
databases is organized into graphs that connect metabolites, genes, proteins and microbes.
- Graph-based learning — graph neural network methods learn relationships across omics
layers that correlation analysis cannot detect, from host-microbiome co-metabolism to metabolic regulation
networks.
- Mechanism networks — the output is a network of metabolite–protein–gene–microbe
connections with the supporting evidence, ready for interpretation and publication.
- Cohort-scale integration — transcriptomics, proteomics, microbiome and metabolomics
from the same or paired cohorts are integrated rather than analyzed in isolation.
The result is a coherent mechanism story — for example, how a microbial metabolite links a gut community to
a host pathway — delivered as a publication-grade network. Our integrative metabolome and transcriptome
analysis service covers the matched-data workflows this integration builds on.
Why Choose Our AI Metabolomics Service
We built this service around the four stages where standard metabolomics software loses value, and applied
AI where it measurably helps — as a working method, not a marketing label.
- AI at the Right Stages
Deep learning for peak curation and annotation, machine learning for mining and integration; each applied where it changes the outcome.
- Dark-Metabolome Depth
In-silico identification extends annotation beyond spectral libraries with confidence scoring.
- Cleaner Data First
Deep-learning denoising removes noise features before any statistics are run.
- Explainable Models
Feature importance and SHAP-style analysis keep machine-learning models interpretable.
- Reproducible Pipelines
Versioned preprocessing and model training, so results are audit-ready.
- Analysis-Ready Outputs
Curated feature matrices, confidence-scored annotations and explainable models delivered in analysis-ready formats.
Available AI Bioinformatics Services — Annotation,
Biomarker and Network Modules
Each AI capability is available as a standalone, purchasable service module — scope the one that matches
your study, or combine them into an end-to-end pipeline.
- In-Silico Unknown Metabolite Annotation Service — annotates features absent from
spectral libraries (the dark metabolome) using fragmentation trees, structural fingerprint prediction and
retention-time/CCS scoring, with per-feature confidence levels. For studies where spectral-library
coverage is the annotation bottleneck.
- Machine-Learning Candidate Biomarker Discovery — feature selection, cohort
stratification and multi-metabolite biomarker panels with cross-validation, ROC analysis and explainable
feature importance. For complex, high-dimensional research cohorts.
- AI-Driven Multi-Omics and Co-Metabolism Network Analysis — knowledge-graph and
graph-based integration across metabolomics, transcriptomics, proteomics and microbiome data, including
host-microbiome co-metabolism networks. For studies connecting metabolites to genes, proteins and
microbial communities.
Each module can be purchased independently or combined into a single project. Contact us to scope which
module fits your study design.
AI-Augmented Metabolomics Workflow — From Raw Data to
Mechanism
Every project runs through a defined, reproducible pipeline where each AI step is applied at the stage it
helps most.
Sample Requirements for AI Metabolomics Analysis
| Sample Type |
Minimum Amount |
Preparation |
Storage and Shipping |
| Plasma / serum |
≥ 100 µL |
Collect in EDTA or heparin tube; centrifuge; aliquot; avoid hemolysis |
−80°C; dry ice |
| Tissue |
≥ 30 mg |
Snap-freeze immediately; record wet weight; avoid thawing |
−80°C; dry ice |
| Cells / cell pellets |
≥ 1×106 cells |
Quench; wash with cold buffer; snap-freeze |
−80°C; dry ice |
| Feces / cecal content |
≥ 100 mg |
Collect fresh; snap-freeze immediately |
−80°C; dry ice |
| Urine |
≥ 1 mL |
Centrifuge; aliquot; avoid repeated freeze-thaw |
−80°C; dry ice |
Raw data files can also be provided directly for AI analysis when the acquisition is complete and
acquisition metadata is documented. We confirm input requirements and study design with you up front, so
volumes, replicates and metadata are agreed before analysis begins.
Deliverables — Curated Features, Annotated Tables and
Explainable Models
Each project delivers an audit-ready AI analysis package:
- Curated feature matrix — post-AI-denoising, with QC annotations
- Annotated metabolite table — confidence levels and supporting evidence
- Machine-learning models — with performance metrics (ROC, cross-validation)
- Candidate biomarker panels — and cohort stratification results
- Multi-omics integration networks — with evidence annotations
- Full QC report — preprocessing, model validation, reproducibility
- Written biological interpretation — linking findings to mechanism hypotheses
Every model output, annotation and network edge is traceable to its source data. Where cohort-scale
statistics or custom visualization are needed, our metabolomics data analysis team extends the core package.
End-to-End Deep Learning Turns Raw LC-MS Data
into Disease-Specific Metabolic Profiles
Background
Conventional untargeted metabolomics runs through a stepwise pipeline — peak picking, alignment,
annotation, then statistics. Each step accumulates its own errors, and features that cannot be matched to a
library are simply dropped. A published study set out to remove these bottlenecks by training a
deep-learning model to go directly from raw LC-MS data to a biological answer.
Challenge:
Train an end-to-end model that classifies samples and reveals disease-specific metabolic profiles directly
from raw LC-MS data, without the error accumulation of stepwise feature extraction.
Analytical Approach
The study introduced an end-to-end deep-learning method that takes raw LC-MS data (retention time, m/z,
intensity) as input and bypasses conventional peak extraction and metabolite identification entirely. An
ensemble of 18 convolutional sub-models (DenseNet121 architecture) was trained on 859 human serum samples
spanning three cohorts from separate hospitals — 210 healthy samples, 323 benign lung nodule samples and 326
lung adenocarcinoma samples — and tested on an independent dataset.
Key Findings (from the published study)
| Finding |
Evidence |
| End-to-end classification from raw data |
Deep-learning model classified samples directly from raw LC-MS data, bypassing stepwise feature
extraction |
| High performance on independent test set |
AUC of 0.99 on the independent testing dataset |
| Early-stage detection accuracy |
96.1% accuracy in detecting early-stage lung adenocarcinoma |
| Key-metabolite and network outputs |
Heatmaps of important metabolic signals and disease-related metabolite–protein networks |
| Cross-cancer generalization |
Applied to lipid metabolomics of 928 cell lines, revealing metabolites and proteins associated
with 23 cancer types |
What this means for your metabolomics program:
- End-to-end AI removes the error accumulation of stepwise pipelines, giving you predictions and insight
directly from raw data.
- Features that would be dropped for lack of a library match are used by the model, so no signal is
wasted.
- The same approach supports biomarker discovery, cohort stratification and mechanism outputs from a
single workflow.
- Inter-hospital batch effects are handled by the model architecture, improving cross-cohort
comparability.
- Decision-grade classification and network outputs support go/no-go and mechanism hypotheses in research
programs.
Conclusion
This published example shows how end-to-end deep learning transforms raw metabolomics data into
classification, biomarker and network insight — the same AI-first philosophy we apply across peak curation,
annotation, mining and integration in our service.
Reference
- Deng, Y., Yao, Y., Wang, Y., Yu, T., Cai, W., Zhou, D., Yin, F., Liu, W., Liu, Y., Xie, C., Guan, J.,
Hu, Y., Huang, P., Li, W.
An end-to-end
deep learning method for mass spectrometry data analysis to reveal disease-specific metabolic
profiles. Nature Communications 15: 7136 (2024).
An end-to-end deep learning method for mass spectrometry data analysis to reveal disease-specific metabolic profiles
Deng, Y., Yao, Y., Wang, Y., Yu, T., Cai, W., Zhou, D., Yin, F., Liu, W., Liu, Y., Xie, C., Guan, J., Hu,
Y., Huang, P., Li, W.
Journal: Nature Communications
Year: 2024
DOI:
https://doi.org/10.1038/s41467-024-51433-3
SIRIUS 4: a rapid tool for turning tandem mass spectra into metabolite structure information
Dührkop, K., Fleischauer, M., Ludwig, M., Aksenov, A.A., Melnik, A.V., Meusel, M., Dorrestein, P.C.,
Rousu, J., Böcker, S.
Journal: Nature Methods
Year: 2019
DOI:
https://doi.org/10.1038/s41592-019-0344-8
Searching molecular structure databases with tandem mass spectra using CSI:FingerID
Dührkop, K., Shen, H., Meusel, M., Rousu, J., Böcker, S.
Journal: Proceedings of the National Academy of Sciences
Year: 2015
DOI:
https://doi.org/10.1073/pnas.1509788112