Summarizing categorical data by clustering attributes

Mampaey, Michael; Vreeken, Jilles

doi:10.1007/S10618-011-0246-6

Title

Summarizing categorical data by clustering attributes

Author

Mampaey, Michael

Vreeken, Jilles

Abstract

For a book, its title and abstract provide a good first impression of what to expect from it. For a database, obtaining a good first impression is typically not so straightforward. While low-order statistics only provide very limited insight, downright mining the data rapidly provides too much detail for such a quick glance. In this paper we propose a middle ground, and introduce a parameter-free method for constructing high-quality descriptive summaries of binary and categorical data. Our approach builds a summary by clustering attributes that strongly correlate, and uses the Minimum Description Length principle to identify the best clustering-without requiring a distance measure between attributes. Besides providing a practical overview of which attributes interact most strongly, these summaries can also be used as surrogates for the data, and can easily be queried. Extensive experimentation shows that our method discovers high-quality results: correlated attributes are correctly grouped, which is verified both objectively and subjectively. Our models can also be employed as surrogates for the data; as an example of this we show that we can quickly and accurately query the estimated supports of frequent generalized itemsets.

Language

English

Source (journal)

Data mining and knowledge discovery. - Boston, Mass., 1997, currens

Publication

Boston, Mass. : 2013

ISSN

1384-5810 [print]

1573-756X [online]

DOI

10.1007/S10618-011-0246-6

Volume/pages

26 :1 (2013) , p. 130-173

ISI

000313116400005

Full text (Publisher's DOI)

https://doi.org/10.1007/S10618-011-0246-6

Full text (publisher's version - intranet only)

https://repository.uantwerpen.be/docman/iruaauth/c8cc45/3013631.pdf

Faculty/Department				Faculty of Sciences. Mathematics and Computer Science

Research group				ADReM Data Lab (ADReM)

Publication type				A1 Journal article

Subject				Computer. Automation

Affiliation				Publications with a UAntwerp address

Web of Science

View record in Web of Science®

View citing articles in Web of Science®

Identifier

Creation

28.02.2013

Last edited

09.10.2023

To cite this reference

https://hdl.handle.net/10067/1059290151162165141