Skip to main navigation Skip to search Skip to main content

Classifier Performance on Long-Tail Distributions

  • Artur Pak
  • , Maxat Tezekbayev
  • , Arman Bolatov
  • , Igor Melnykov
  • , Kaisar Dauletbek
  • , Zhenisbek Assylbekov

Research output: Contribution to journalArticlepeer-review

Abstract

This paper introduces a Gaussian mixture model designed to explore the implications of long-tailedness on classification perfor mance. Our study reveals that simple under-specified classifiers are inherently limited in reducing generalization error within this framework, a limitation overcome by well-specified and even more complex over-specified classifiers capable of fitting rare examples. Through theoretical analysis and empirical evaluation on both synthetic and real datasets, we demonstrate how the performance gap between under-specified and well/over-specified classifiers varies with changes in the tail length of the distribu tion. Our results validate the necessity of fitting rare instances in training datasets to achieve enhanced generalization capabilities in models faced with long-tail distributed data.

Original languageEnglish (US)
Article numbere70037
JournalStatistical Analysis and Data Mining
Volume18
Issue number4
DOIs
StatePublished - 2025

Bibliographical note

Publisher Copyright:
© 2025, John Wiley and Sons Inc. All rights reserved.

Keywords

  • class imbalance
  • gaussian mixtures
  • long-tail distributions

Fingerprint

Dive into the research topics of 'Classifier Performance on Long-Tail Distributions'. Together they form a unique fingerprint.

Cite this