A4 Refereed article in a conference publication
Local Language Models for Context-Aware Adaptive Anonymization of Sensitive Text
Authors: Adeseye, Aisvarya; Isoaho, Jouni; Virtanen, Seppo; Tahir, Mohammad
Editors: Ferens, Ken; Deligiannidis, Leonidas; Arabnia, Hamid R.; de la Fuente, David; Olivas, José A.
Conference name: World Congress in Computer Science, Computer Engineering, and Applied Computing
Publication year: 2026
Journal: Communications in Computer and Information Science
Book title : Applied Cognitive Computing and Artificial Intelligence
Volume: 2933
First page : 190
Last page: 204
ISBN: 978-3-032-22204-6
eISBN: 978-3-032-22205-3
ISSN: 1865-0929
eISSN: 1865-0937
DOI: https://doi.org/10.1007/978-3-032-22205-3_14
Publication's open availability at the time of reporting: No Open Access
Publication channel's open availability : Partially Open Access publication channel
Web address : https://doi.org/10.1007/978-3-032-22205-3_14
Qualitative research often contains personal, contextual, and organizational details that pose privacy risks if not handled appropriately. Manual anonymization is time-consuming, inconsistent, and frequently omits critical identifiers. Existing automated tools tend to rely on pattern matching or fixed rules, which fail to capture context and may alter the meaning of the data. This study uses local Large Language Models (LLMs) to build a reliable, repeatable, and context-aware anonymization process for detecting and anonymizing sensitive data in qualitative transcripts. We introduce a Structured Framework for Adaptive Anonymizer (SFAA) that includes three steps: detection, classification, and adaptive anonymization. The SFAA incorporates four anonymization strategies: rule-based substitution, context-aware rewriting, generalization, and suppression. These strategies are applied based on the identifier type and the risk level. The identifiers handled by the SFAA are guided by major international privacy and research ethics standards, including the GDPR, HIPAA and National Standards TCPS 2, and OECD guidelines. This study followed a dual-method evaluation that combined manual and LLM-assisted processing. Two case studies were used to support the evaluation. The first includes 82 face-to-face interviews on gamification in organizations. The second involves 93 machine-led interviews using an AI-powered interviewer to test LLM awareness and workplace privacy. Two local models, LLaMA and Phi were used to evaluate the performance of the proposed framework. The results indicate that the LLMs found more sensitive data than a human reviewer. Phi outperformed LLaMA in finding sensitive data, but made slightly more errors. Phi was able to find over 91% of the sensitive data and 94.8% kept the same sentiment as the original text, which means it was very accurate, hence, it does not affect the analysis of the qualitative data.