A4 Refereed article in a conference publication

Local Language Models for Context-Aware Adaptive Anonymization of Sensitive Text




AuthorsAdeseye, Aisvarya; Isoaho, Jouni; Virtanen, Seppo; Tahir, Mohammad

EditorsFerens, Ken; Deligiannidis, Leonidas; Arabnia, Hamid R.; de la Fuente, David; Olivas, José A.

Conference nameWorld Congress in Computer Science, Computer Engineering, and Applied Computing

Publication year2026

Journal: Communications in Computer and Information Science

Book title Applied Cognitive Computing and Artificial Intelligence

Volume2933

First page 190

Last page204

ISBN978-3-032-22204-6

eISBN978-3-032-22205-3

ISSN1865-0929

eISSN1865-0937

DOIhttps://doi.org/10.1007/978-3-032-22205-3_14

Publication's open availability at the time of reportingNo Open Access

Publication channel's open availability Partially Open Access publication channel

Web address https://doi.org/10.1007/978-3-032-22205-3_14


Abstract

Qualitative research often contains personal, contextual, and organizational details that pose privacy risks if not handled appropriately. Manual anonymization is time-consuming, inconsistent, and frequently omits critical identifiers. Existing automated tools tend to rely on pattern matching or fixed rules, which fail to capture context and may alter the meaning of the data. This study uses local Large Language Models (LLMs) to build a reliable, repeatable, and context-aware anonymization process for detecting and anonymizing sensitive data in qualitative transcripts. We introduce a Structured Framework for Adaptive Anonymizer (SFAA) that includes three steps: detection, classification, and adaptive anonymization. The SFAA incorporates four anonymization strategies: rule-based substitution, context-aware rewriting, generalization, and suppression. These strategies are applied based on the identifier type and the risk level. The identifiers handled by the SFAA are guided by major international privacy and research ethics standards, including the GDPR, HIPAA and National Standards TCPS 2, and OECD guidelines. This study followed a dual-method evaluation that combined manual and LLM-assisted processing. Two case studies were used to support the evaluation. The first includes 82 face-to-face interviews on gamification in organizations. The second involves 93 machine-led interviews using an AI-powered interviewer to test LLM awareness and workplace privacy. Two local models, LLaMA and Phi were used to evaluate the performance of the proposed framework. The results indicate that the LLMs found more sensitive data than a human reviewer. Phi outperformed LLaMA in finding sensitive data, but made slightly more errors. Phi was able to find over 91% of the sensitive data and 94.8% kept the same sentiment as the original text, which means it was very accurate, hence, it does not affect the analysis of the qualitative data.



Last updated on 02/06/2026 07:53:56 AM