An integrated workflow for robust alignment and simplified quantitative analysis of NMR spectrometry data

Vu, Trung N.; Valkenborg, Dirk; Smets, Koen; Verwaest, Kim A.; Dommisse, Roger; Lemière, Filip; Verschoren, Alain; Goethals, Bart; Laukens, Kris

doi:10.1186/1471-2105-12-405

Title

An integrated workflow for robust alignment and simplified quantitative analysis of NMR spectrometry data

Author

Vu, Trung N.

Valkenborg, Dirk

Smets, Koen

Verwaest, Kim A.

Dommisse, Roger

Lemière, Filip

Verschoren, Alain

Goethals, Bart

Laukens, Kris

Abstract

Background Nuclear magnetic resonance spectroscopy (NMR) is a powerful technique to reveal and compare quantitative metabolic profiles of biological tissues. However, chemical and physical sample variations make the analysis of the data challenging, and typically require the application of a number of preprocessing steps prior to data interpretation. For example, noise reduction, normalization, baseline correction, peak picking, spectrum alignment and statistical analysis are indispensable components in any NMR analysis pipeline. Results We introduce a novel suite of informatics tools for the quantitative analysis of NMR metabolomic profile data. The core of the processing cascade is a novel peak alignment algorithm, called hierarchical Cluster-based Peak Alignment (CluPA). The algorithm aligns a target spectrum to the reference spectrum in a top-down fashion by building a hierarchical cluster tree from peak lists of reference and target spectra and then dividing the spectra into smaller segments based on the most distant clusters of the tree. To reduce the computational time to estimate the spectral misalignment, the method makes use of Fast Fourier Transformation (FFT) cross-correlation. Since the method returns a high-quality alignment, we can propose a simple methodology to study the variability of the NMR spectra. For each aligned NMR data point the ratio of the between-group and within-group sum of squares (BW-ratio) is calculated to quantify the difference in variability between and within predefined groups of NMR spectra. This differential analysis is related to the calculation of the F-statistic or a one-way ANOVA, but without distributional assumptions. Statistical inference based on the BW-ratio is achieved by bootstrapping the null distribution from the experimental data. Conclusions The workflow performance was evaluated using a previously published dataset. Correlation maps, spectral and grey scale plots show clear improvements in comparison to other methods, and the down-to-earth quantitative analysis works well for the CluPA-aligned spectra. The whole workflow is embedded into a modular and statistically sound framework that is implemented as an R package called "speaq" ("spectrum alignment and quantitation"), which is freely available from http://code.google.com/p/speaq/.

Language

English

Source (journal)

BMC bioinformatics. - London

Publication

London : 2011

ISSN

1471-2105

DOI

10.1186/1471-2105-12-405

Volume/pages

12 (2011) , 14 p.

Article Reference

405

ISI

000297044500001

Medium

E-only publicatie

Full text (Publisher's DOI)

https://doi.org/10.1186/1471-2105-12-405

Full text (open access)

https://repository.uantwerpen.be/docman/irua/1366a4/f13682a5.pdf

Faculty/Department				Faculty of Sciences. Chemistry Faculty of Sciences. Mathematics and Computer Science

Research group				ADReM Data Lab (ADReM)
Project info				Intelligent analysis and data-mining of mass spectrometry-based proteome data.
Publication type				A1 Journal article

Subject				Mathematics Chemistry Biology Engineering sciences. Technology Computer. Automation

Affiliation				Publications with a UAntwerp address

Web of Science

View record in Web of Science®

View citing articles in Web of Science®

Identifier

Creation

24.10.2011

Last edited

15.11.2022

To cite this reference

https://hdl.handle.net/10067/918740151162165141