Unravelling associations between unassigned mass spectrometry peaks with frequent itemset mining techniques

Vu, Trung Nghia; Mrzic, Aida; Valkenborg, Dirk; Maes, Evelyne; Lemière, Filip; Goethals, Bart; Laukens, Kris

doi:10.1186/S12953-014-0054-1

Title

Unravelling associations between unassigned mass spectrometry peaks with frequent itemset mining techniques

Author

Vu, Trung Nghia

Mrzic, Aida

Valkenborg, Dirk

Maes, Evelyne

Lemière, Filip

Goethals, Bart

Laukens, Kris

Abstract

BackgroundMass spectrometry-based proteomics experiments generate spectra that are rich in information. Often only a fraction of this information is used for peptide/protein identification, whereas a significant proportion of the peaks in a spectrum remain unexplained. In this paper we explore how a specific class of data mining techniques termed ?frequent itemset mining? can be employed to discover patterns in the unassigned data, and how such patterns can help us interpret the origin of the unexpected/unexplained peaks.ResultsFirst a model is proposed that describes the origin of the observed peaks in a mass spectrum. For this purpose we use the classical correlative database search algorithm. Peaks that support a positive identification of the spectrum are termed explained peaks. Next, frequent itemset mining techniques are introduced to infer which unexplained peaks are associated in a spectrum. The method is validated on two types of experimental proteomic data. First, peptide mass fingerprint data is analyzed to explain the unassigned peaks in a full scan mass spectrum. Interestingly, a large numbers of experimental spectra reveals several highly frequent unexplained masses, and pattern mining on these frequent masses demonstrates that subsets of these peaks frequently co-occur. Further evaluation shows that several of these co-occurring peaks indeed have a known common origin, and other patterns are promising hypothesis generators for further analysis. Second, the proposed methodology is validated on tandem mass spectrometral data using a public spectral library, where associations within the mass differences of unassigned peaks and peptide modifications are explored. The investigation of the found patterns illustrates that meaningful patterns can be discovered that can be explained by features of the employed technology and found modifications.ConclusionsThis simple approach offers opportunities to monitor accumulating unexplained mass spectrometry data for emerging new patterns, with possible applications for the development of mass exclusion lists, for the refinement of quality control strategies and for a further interpretation of unexplained spectral peaks in mass spectrometry and tandem mass spectrometry.

Language

English

Source (journal)

Proteome science. - London

Publication

London : 2014

ISSN

1477-5956

DOI

10.1186/S12953-014-0054-1

Volume/pages

12 (2014) , 8 p.

Article Reference

54

ISI

000348393600001

Pubmed ID

25429250

Medium

E-only publicatie

Full text (Publisher's DOI)

https://doi.org/10.1186/S12953-014-0054-1

Full text (open access)

https://repository.uantwerpen.be/docman/irua/b53af0/d003ed4b.pdf

Faculty/Department				Faculty of Sciences. Biology Faculty of Sciences. Chemistry Faculty of Sciences. Mathematics and Computer Science

Research group				ADReM Data Lab (ADReM)
Publication type				A1 Journal article

Subject				Chemistry Biology Computer. Automation

Affiliation				Publications with a UAntwerp address

Web of Science

View record in Web of Science®

View citing articles in Web of Science®

Identifier

Creation

19.11.2014

Last edited

04.03.2024

To cite this reference

https://hdl.handle.net/10067/1204160151162165141