Abstract
This paper introduces a Gaussian mixture model designed to explore the implications of long-tailedness on classification perfor mance. Our study reveals that simple under-specified classifiers are inherently limited in reducing generalization error within this framework, a limitation overcome by well-specified and even more complex over-specified classifiers capable of fitting rare examples. Through theoretical analysis and empirical evaluation on both synthetic and real datasets, we demonstrate how the performance gap between under-specified and well/over-specified classifiers varies with changes in the tail length of the distribu tion. Our results validate the necessity of fitting rare instances in training datasets to achieve enhanced generalization capabilities in models faced with long-tail distributed data.
| Original language | English (US) |
|---|---|
| Article number | e70037 |
| Journal | Statistical Analysis and Data Mining |
| Volume | 18 |
| Issue number | 4 |
| DOIs | |
| State | Published - 2025 |
Bibliographical note
Publisher Copyright:© 2025, John Wiley and Sons Inc. All rights reserved.
Keywords
- class imbalance
- gaussian mixtures
- long-tail distributions
Fingerprint
Dive into the research topics of 'Classifier Performance on Long-Tail Distributions'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS