Skip to main navigation Skip to search Skip to main content

Learning From Crowdsourced Noisy Labels: A signal processing perspective

Research output: Contribution to journalArticlepeer-review

Abstract

One of the primary catalysts fueling advances in artificial intelligence (AI ) and machine learning (ML) is the availability of massive, curated datasets. A commonly used technique to curate such massive datasets is crowdsourcing, in which data are dispatched to multiple annotators. The annotator-produced labels are then fused to serve downstream learning and inference tasks. This annotation process often creates noisy labels for various reasons, such as the limited expertise or unreliability of annotators, among others. Therefore, a core objective in crowdsourcing is to develop methods thateffectively mitigate the negative impact of such label noise on learning tasks. This feature article introduces advances in learning from noisy crowdsourced labels. The focus is on key crowdsourcing models and their methodological treatments, from classical statistical models to recent deep learning-based approaches, emphasizing analytical insights and algorithmic developments. In particular, this article reviews the connections between signal processing (SP) theory and methods, such as identifiability of tensor and nonnegative matrix factorization and novel, principled solutions of longstanding challenges in crowdsourcing—showing how SP perspectives drive the advancements of this field. Furthermore, this article touches upon emerging topics that are critical for developing cutting-edge AI/ML systems, such as crowdsourcing in reinforcement learning with human feedback (RLHF) and direct preference optimization (DPO), which are key techniques for fine-tuning large language models (LLMs).

Original languageEnglish (US)
Pages (from-to)84-106
Number of pages23
JournalIEEE Signal Processing Magazine
Volume42
Issue number3
DOIs
StatePublished - 2025

Bibliographical note

Publisher Copyright:
© 1991-2012 IEEE.

Fingerprint

Dive into the research topics of 'Learning From Crowdsourced Noisy Labels: A signal processing perspective'. Together they form a unique fingerprint.

Cite this