Skip to Main Content (Press Enter)

Logo UNIMI
  • ×
  • Home
  • Persone
  • Attività
  • Ambiti
  • Strutture
  • Pubblicazioni
  • Terza Missione

Expertise & Skills
Logo UNIMI

|

Expertise & Skills

unimi.it
  • ×
  • Home
  • Persone
  • Attività
  • Ambiti
  • Strutture
  • Pubblicazioni
  • Terza Missione
  1. Pubblicazioni

EHLD: A General Data Framework for Harmful Language Detection and Evaluation

Articolo
Data di Pubblicazione:
2026
Citazione:
EHLD: A General Data Framework for Harmful Language Detection and Evaluation / F. Mohammadi, P.C.. - In: DIGITAL LIBRARY PERSPECTIVES. - ISSN 2059-5824. - (2026). [Epub ahead of print] [10.1108/DLP-01-2026-0028]
Abstract:
Purpose- Research into harmful language detection is hindered by fragmented datasets and incompatible label schemas, which significantly limit evaluation. This paper introduces EHLD: a general, extensible data framework supporting the integration, enrichment and evaluation of diverse harmful language resources, with a particular focus on Italian.

Design/Methodology/Approach- EHLD defines a unified schema with essential detection attributes (e.g., harmful_or_not and harm_subcategory), optional contextual dimensions (e.g., target, category, intersectionality), evaluation-oriented text complexity features, and metadata for provenance. We instantiate EHLD by integrating three data sources: curated datasets from the literature, large-scale LLM-based annotation from the Mappa dell’Intolleranza, and LLM generated data for inclusive language detection. We further introduce label_confidence to explicitly encode label reliability. Model selection follows an iterative cycle combining distributional analysis, linguistic diversity assessment, and performance evaluation.

Findings- The resulting dataset contains 237,956 instances with 18 features and exhibits substantial linguistic variety (Self-BLEU drops from 0.759 at 1-gram to 0.107 at 4-grams). After iterative, data-centric refinement, GPT-5-nano is selected as the reference model and achieves an F1-score of 0.813 on binary harmful language detection, outperforming reported baselines such as BERT-based multitask models and other LLMs on comparable Italian settings.

Research limitations/implications- Due to differences in datasets, domains, label definitions and evaluation protocols, comparisons across studies are not fully controlled. The framework partly relies on automatically generated labels, which, despite confidence-aware handling, may introduce residual noise.

Practical implications- By combining standardised labels, provenance tracking and complexity-oriented evaluation features, EHLD supports the construction of reproducible datasets and richer model auditing, enabling a more transparent deployment of harmful language detection systems.

Originality/value- EHLD contributes a reusable, confidence-aware integration framework that connects different harmful language resources and allows iterative, data-driven model selection and evaluation, with a focus on Italian.
Tipologia IRIS:
01 - Articolo su periodico
Keywords:
Harmful language Detection; Data Integration; Data Evaluation;
Elenco autori:
F. Mohammadi, P. Ceravolo, S. Maghool, M. Tamborini, M.E. D'Amico
Autori di Ateneo:
CERAVOLO PAOLO ( autore )
D'AMICO MARIA ELISA ( autore )
MOHAMMADI FATEMEH ( autore )
TAMBORINI MARTA ANNAMARIA ( autore )
Link alla scheda completa:
https://air.unimi.it/handle/2434/1235218
Link al Full Text:
https://air.unimi.it/retrieve/handle/2434/1235218/3303733/Attached%20standard%20file_.PDF
Progetto:
MUSA - Multilayered Urban Sustainability Actiona
  • Aree Di Ricerca

Aree Di Ricerca

Settori


Settore INFO-01/A - Informatica
  • Informazioni
  • Assistenza
  • Accessibilità
  • Privacy
  • Utilizzo dei cookie
  • Note legali

Realizzato con VIVO | Progettato da Cineca | 26.7.0.0