HARMONY: Efficiently mining the best rules for classification

Jianyong Wang, George Karypis

Research output: Contribution to conferencePaper

103 Scopus citations

Abstract

Many studies have shown that rule-based classifiers perform well in classifying categorical and sparse high-dimensional databases. However, a fundamental limitation with many rule-based classifiers is that they find the rules by employing various heuristic methods to prune the search space, and select the rules based on the sequential database covering paradigm. As a result, the final set of rules that they use may not be the globally best rules for some instances in the training database. To make matters worse, these algorithms fail to fully exploit some more effective search space pruning methods in order to scale to large databases. In this paper we present a new classifier, HARMONY, which directly mines the final set of classification rules. HARMONY uses an instance-centric rule-generation approach and it can assure for each training instance, one of the highest-confidence rules covering this instance is included in the final rule set, which helps in improving the overall accuracy of the classifier. By introducing several novel search strategies and pruning methods into the rule discovery process, HARMONY also has high efficiency and good scalability. Our thorough performance study with some large text and categorical databases has shown that HARMONY outperforms many well-known classifiers in terms of both accuracy and computational efficiency, and scales well w.r.t. the database size.

Original languageEnglish (US)
Pages205-216
Number of pages12
DOIs
StatePublished - 2005
Event5th SIAM International Conference on Data Mining, SDM 2005 - Newport Beach, CA, United States
Duration: Apr 21 2005Apr 23 2005

Other

Other5th SIAM International Conference on Data Mining, SDM 2005
CountryUnited States
CityNewport Beach, CA
Period4/21/054/23/05

Fingerprint Dive into the research topics of 'HARMONY: Efficiently mining the best rules for classification'. Together they form a unique fingerprint.

  • Cite this

    Wang, J., & Karypis, G. (2005). HARMONY: Efficiently mining the best rules for classification. 205-216. Paper presented at 5th SIAM International Conference on Data Mining, SDM 2005, Newport Beach, CA, United States. https://doi.org/10.1137/1.9781611972757.19