A4 Refereed article in a conference publication

Evaluating LLM Proficiency in Analyzing Privacy Aspects of Network Traffic




AuthorsPuhtila, Panu; Heino, Timi; Terho, Hanna-Kaisa; Rajapaksha, Sammani; Harshani, S.; Koivunen, Lauri; Mäkilä, Tuomas

EditorsBabic, Snjezana; Car, Zeljka; Cicin-Sain, Marina; Ergovic, Pavle; Galinac Grbac, Tihana; Gros, Stjepan; Jovic, Alan; Jurekovic, Darko; Katulic, Tihomir; Koricic, Marko; Kralj, Nenad; Mornar, Vedran; Petrovic, Juraj; Skala, Karolj; Skvorc, Dejan; Sruk, Vlado; Tijan, Edvard; Valacich, Joe; Vrcek, Neven; Vrdoljak, Boris

Conference nameMIPRO ICT and Electronics Convention

Publication year2026

Journal: International Convention on Information and Communication Technology, Electronics and Microelectronics

Book title 2026 49th MIPRO ICT and Electronics Convention (MIPRO)

Volume49

First page 795

Last page800

ISBN979-8-3315-6310-3

eISBN979-8-3315-6309-7

ISSN1847-3938

eISSN1847-3946

DOIhttps://doi.org/10.1109/MIPRO70003.2026.11592022

Publication's open availability at the time of reportingNo Open Access

Publication channel's open availability No Open Access publication channel

Web address https://ieeexplore.ieee.org/document/11592022


Abstract

Important aspect of privacy research is the analysis of HTTP Archive (HAR) files, which record the network traffic happening out of a given domain. Large language models (LLMs) have demonstrated capabilities that could be useful in this kind of work, but research into the validity of the privacy analysis by the LLMs is lacking. Such investigation is necessary if we want to use these technologies reliably in research. We measure and compare the efficiency of four LLMs in analysing the HAR files; GPT-4o, o1-preview, LLaMA3.3B70 and Claude Sonnet 3.5. We evaluate LLM ability to detect third parties present in the website, analysis of cookies and of User-Agent strings. We experiment on whether the use of Retrieval Augmented Generation (RAG) improves results, compared to default LLM. Results indicate that all of the studied LLMs are incapable of absolutely correct analysis, with or without RAG assistance, although the flaws in their output are relatively small. In general o1preview attained best results, for example in analysing third party URLs it was correct 88% of the time. In contrast, LLaMA3.3B70 was the worst, achieving only 34% of correct answers in this task. Results in other tasks followed similar pattern, while RAG gave inconclusive results.



Last updated on 03/08/2026 09:17:26 AM