Evaluating LLM Proficiency in Analyzing Privacy Aspects of Network Traffic
: Puhtila, Panu; Heino, Timi; Terho, Hanna-Kaisa; Rajapaksha, Sammani; Harshani, S.; Koivunen, Lauri; Mäkilä, Tuomas
: Babic, Snjezana; Car, Zeljka; Cicin-Sain, Marina; Ergovic, Pavle; Galinac Grbac, Tihana; Gros, Stjepan; Jovic, Alan; Jurekovic, Darko; Katulic, Tihomir; Koricic, Marko; Kralj, Nenad; Mornar, Vedran; Petrovic, Juraj; Skala, Karolj; Skvorc, Dejan; Sruk, Vlado; Tijan, Edvard; Valacich, Joe; Vrcek, Neven; Vrdoljak, Boris
: MIPRO ICT and Electronics Convention
: 2026
International Convention on Information and Communication Technology, Electronics and Microelectronics
: 2026 49th MIPRO ICT and Electronics Convention (MIPRO)
: 49
: 795
: 800
: 979-8-3315-6310-3
: 979-8-3315-6309-7
: 1847-3938
: 1847-3946
DOI: https://doi.org/10.1109/MIPRO70003.2026.11592022
: https://ieeexplore.ieee.org/document/11592022
Important aspect of privacy research is the analysis of HTTP Archive (HAR) files, which record the network traffic happening out of a given domain. Large language models (LLMs) have demonstrated capabilities that could be useful in this kind of work, but research into the validity of the privacy analysis by the LLMs is lacking. Such investigation is necessary if we want to use these technologies reliably in research. We measure and compare the efficiency of four LLMs in analysing the HAR files; GPT-4o, o1-preview, LLaMA3.3B70 and Claude Sonnet 3.5. We evaluate LLM ability to detect third parties present in the website, analysis of cookies and of User-Agent strings. We experiment on whether the use of Retrieval Augmented Generation (RAG) improves results, compared to default LLM. Results indicate that all of the studied LLMs are incapable of absolutely correct analysis, with or without RAG assistance, although the flaws in their output are relatively small. In general o1preview attained best results, for example in analysing third party URLs it was correct 88% of the time. In contrast, LLaMA3.3B70 was the worst, achieving only 34% of correct answers in this task. Results in other tasks followed similar pattern, while RAG gave inconclusive results.