Proceedings of International Conference on Applied Innovation in IT  ·  2026/06/12  ·  Vol. 14  ·  Issue 4  ·  pp. 49–60
Hybrid Lexical and Forensic Linguistic Features for AI-Generated Text Detection Using Machine Learning
Amira N. Abd Al-Jaleel and Husam J. Mohammed
The rapid and widespread use of large language models (LLMs) has ushered in the rapid increase in the quantity of AI-produced text everywhere on the Internet, which has necessitated significant concerns about the field of digital forensics, cybersecurity, and academic integrity. The correct distinction between text produced by human authors and text produced by artificial intelligence can be described as a difficult task, which is largely conditioned by the growing fluency of the generative models as well as by the little generalizability of the existing detection tools. The model in question uses a combination of TF-IDF in unigrams and bigrams with 21 linguistically interpretable features. The features are created to match the text and capture it in the stylistic, syntactic, semantic, and discourse coherence levels. They were conducted on the large-scale Human vs. LLM dataset that contains 788,000 text samples of various sources of humans and several advanced and modern language models generated texts. It was tested on the performance of several classification algorithms, such as logistic regression, linear support vector machine, random forest, and LightGBM, but not limited to these products. The experimental outcomes proved that retaining natural linguistic patterns and applying criminal linguistic characteristics played an important role in enhancing detection performance. The LightGBM classifier achieved the best results, with an accuracy rate of 93.4% and an F1 score of 0.93 % for human texts and 0.94 % for AI-generated texts. These findings, moving away from hidden deep learning models, help forensic experts because the new method makes digital investigations more reliable and makes AI more responsible. This transparent tool ensures that textual evidence is ready for court; consequently, study contributes to better safety in the digital world.
LLM Forensics Cybersecurity Linguistic Features Hybrid Framework Feature Fusion
References
  1. R. Sousa-Silva, “Fighting Cyber-malice: A Forensic Linguistics Approach to Detecting AI-generated Malicious Texts,” [Online]. Available: https://www.weforum.org/agenda/2024/01/.
  2. A. Kayabas, A. E. Topcu, Y. I. Alzoubi, and M. Yildiz, “A Deep Learning Approach to Classify AI-Generated and Human-Written Texts,” Applied Sciences (Switzerland), vol. 15, no. 10, May 2025, [Online]. Available: https://doi.org/10.3390/app15105541.
  3. B. Boutadjine, F. Harrag, and K. Shaalan, “Human vs. Machine: A Comparative Study on the Detection of AI-Generated Content,” ACM Transactions on Asian and Low-Resource Language Information Processing, vol. 24, no. 2, Feb. 2025, [Online]. Available: https://doi.org/10.1145/3708889.
  4. L. Chang, “Detecting AI-Generated Text: A Comparative Study of Machine Learning Algorithms,” J. Syst. Cybern. Inf., vol. 23, pp. 36-41, Dec. 2025, [Online]. Available: https://doi.org/10.54808/JSCI.23.06.36.
  5. G. P. Georgiou, “Differentiating between Human-Written and AI-Generated Texts Using Linguistic Features Automatically Extracted from an Online Computational Tool.”
  6. S. Mitrović, D. Andreoletti, and O. Ayoub, “ChatGPT or Human? Detect and Explain. Explaining Decisions of Machine Learning Model for Detecting Short ChatGPT-generated Text,” Jan. 2023, [Online]. Available: http://arxiv.org/abs/2301.13852.
  7. C. Opara, “Distinguishing AI-Generated and Human-Written Text Through Psycholinguistic Analysis,” May 2025, [Online]. Available: http://arxiv.org/abs/2505.01800.
  8. K. Gryka, K. Gradon, M. Kozlowski, M. Kutyla, and A. Janicki, “Detection of AI-Generated Emails - A Case Study,” in ACM International Conference Proceeding Series, Association for Computing Machinery, Jul. 2024, [Online]. Available: https://doi.org/10.1145/3664476.3670465.
  9. M. S. Ibrahim, J. A. Eleiwy, H. M. Muhi-Aldeen, Y. Al-Yasiri, and A. A. Nafea, “Human to Chatbot Text Classification Using Multi-Source AI Chatbots and Machine Learning Models,” Journal of Intelligent Systems and Internet of Things, vol. 16, no. 1, pp. 152-165, 2025, [Online]. Available: https://doi.org/10.54216/JISIoT.160113.
  10. S. Parimanoharan and R. D. Nawarathna, “Assessing Classical Machine Learning and Transformer-based Approaches for Detecting AI-Generated Research Text,” [Online]. Available: https://orcid.org/0000-0000-0000-0000.
  11. R. Ardeshirifar, “Comparing Hand-Crafted and Deep Learning Approaches for Detecting AI-Generated Text: Performance, Generalization, and Linguistic Insights,” AI and Ethics, vol. 5, no. 4, pp. 4197-4209, Aug. 2025, [Online]. Available: https://doi.org/10.1007/s43681-025-00699-4.
  12. S. Fariello, G. Fenza, F. Forte, M. Gallo, and M. Marotta, “Distinguishing Human From Machine: A Review of Advances and Challenges in AI-Generated Text Detection,” International Journal of Interactive Multimedia and Artificial Intelligence, vol. 9, no. 3, pp. 6-18, 2025, [Online]. Available: https://doi.org/10.9781/ijimai.2024.12.002.
  13. K. Teja Repaka, M. Anirudh Bondugula, and S. Sashaank Adibhatla, “Benchmarking Distributed Machine Learning Systems with Large Language Models on Human vs. LLM Text Corpus.”
  14. G. A. Godghase, R. Agrawal, T. Obili, and M. Stamp, “Distinguishing Chatbot from Human,” Aug. 2024, [Online]. Available: http://arxiv.org/abs/2408.04647.
  15. A. M. Elkhatat, K. Elsaid, and S. Almeer, “Evaluating the Efficacy of AI Content Detection Tools in Differentiating Between Human and AI-Generated Text,” International Journal for Educational Integrity, vol. 19, no. 1, Dec. 2023, [Online]. Available: https://doi.org/10.1007/s40979-023-00140-5.
  16. R. Khera, A. F. Pedroso, V. K. Keloth, H. Xu, G. S. Silva, and L. H. Schwamm, “Scientific Writing in the Era of Large Language Models: A Computational Analysis of AI-Versus Human-Created Content,” Stroke, vol. 56, no. 10, pp. 3078-3083, Oct. 2025, [Online]. Available: https://doi.org/10.1161/STROKEAHA.125.051913.
  17. D. Valiaiev, “Detection of Machine-Generated Text: Literature Survey,” 2023, [Online]. Available: https://doi.org/0000001.0000001.
  18. C. Maddugoda, “A Comprehensive Review: Detection Techniques for Human-Generated and AI-Generated Texts.”
  19. P. Ranade, A. Piplai, S. Mittal, A. Joshi, and T. Finin, “Generating Fake Cyber Threat Intelligence Using Transformer-Based Models,” Jun. 2021, [Online]. Available: http://arxiv.org/abs/2102.04351.
  20. T. Kumarage, J. Garland, A. Bhattacharjee, K. Trapeznikov, S. Ruston, and H. Liu, “Stylometric Detection of AI-Generated Text in Twitter Timelines,” Mar. 2023, [Online]. Available: http://arxiv.org/abs/2303.03697.
  21. Z. Zeng, L. Sha, Y. Li, K. Yang, G. Gašević, and G. Chen, “Towards Automatic Boundary Detection for Human-AI Collaborative Hybrid Essay in Education,” 2024, [Online]. Available: https://www.kaggle.com/c/asap-aes.
  22. Z. Grinberg, “Human vs. LLM Text Corpus,” Kaggle, [Online]. Available: https://www.kaggle.com/datasets/starblasters8/human-vs-llm-text-corpus.
  23. G. P. Georgiou, “Differentiating Between Human-Written and AI-Generated Texts Using Automatically Extracted Linguistic Features,” Information (Switzerland), vol. 16, no. 11, Nov. 2025, [Online]. Available: https://doi.org/10.3390/info16110979.


Proceedings of the International Conference on Applied Innovations in IT by Anhalt University of Applied Sciences is licensed under CC BY-SA 4.0
 ·  This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License

ICAIIT 2026
International Conference on Applied Innovation in IT
Navigation
Publisher
ISSN2199-8876
Location Anhalt University of Applied Sciences
Phone +49 (0) 3496 67 5611
Address Building 01, Room 425
Bernburger Str. 55
D-06366 Köthen, Germany
Open Access License

All works are licensed under the Creative Commons Attribution-ShareAlike 4.0 International License (CC BY-SA 4.0), unless otherwise noted.

Published by ICAIIT in cooperation with Anhalt University of Applied Sciences.

© 2026 ICAIIT — International Conference on Applied Innovations in IT. Anhalt University of Applied Sciences, Köthen, Germany.
Visitors: site traffic counter