Lightweight Detection of AI-Generated Product Reviews Using Stylometric Features and Logistic Regression

Authors

  • Karthik Juluri Independent Researcher Author

Keywords:

AI-Generated Text Detection, Product Reviews, Stylometry, Logistic Regression, Explainable Artificial Intelligence, Lightweight Machine Learning

Abstract

Generative artificial intelligence can produce fluent and persuasive product reviews at negligible marginal cost, creating new opportunities for inauthentic promotion, reputation attacks, and manipulation of electronic-commerce signals. Many machine-generated-text detectors use transformer classifiers, model likelihoods, or repeated perturbations, which can impose substantial computational, access, and interpretability costs. This paper proposes a lightweight detector based on 20 inexpensive stylometric features, standardization, and L2-regularized logistic regression. The framework extracts lexical-diversity, sentence-structure, punctuation, repetition, readability, and personal-style indicators; returns a probability rather than a categorical accusation; and explains each score through signed feature contributions. Because no real review corpus is claimed, feasibility is evaluated through a fully reproducible simulation of 4,000 partially overlapping review-level feature vectors, evenly divided between simulated human and AI classes. A fixed-seed 70/15/15 split, validation-only regularization selection, 5% development-label noise, and a clean held-out test set are used. The executed simulation obtains 0.762 accuracy, 0.759 F1-score, 0.762 balanced accuracy, 0.523 Matthews correlation coefficient, and 0.850 ROC-AUC. Repeated-bigram ratio is the strongest positive simulated association, whereas repeated-word and contraction frequencies are strong negative associations. The results demonstrate conceptual feasibility, not real-world generalizability. The intended role is an auditable, CPU-friendly first-stage screening mechanism that supports calibrated referral and human review rather than definitive authorship adjudication.

References

M. Ott, Y. Choi, C. Cardie, and J. T. Hancock, “Finding deceptive opinion spam by any stretch of the imagination,” in Proceedings of the 49th annual meeting of the association for computational linguistics: Human language technologies, Association for Computational Linguistics, 2011, pp. 309–319. Available: https://aclanthology.org/P11-1032/

M. Crawford, T. M. Khoshgoftaar, J. D. Prusa, A. N. Richter, and H. A. Najada, “Survey of review spam detection using machine learning techniques,” Journal of Big Data, vol. 2, no. 1, p. 23, 2015, doi: 10.1186/s40537-015-0029-9.

D. I. Adelani, H. Mai, F. Fang, H. H. Nguyen, J. Yamagishi, and I. Echizen, “Generating sentiment-preserving fake online reviews using neural language models and their human- and machine-based detection,” in Advanced information networking and applications, in Advances in intelligent systems and computing, vol. 1151. 2020, pp. 1341–1354. doi: 10.1007/978-3-030-44041-1_114.

G. Jawahar, M. Abdul-Mageed, and V. S. L. Lakshmanan, “Automatic detection of machine generated text: A critical survey,” in Proceedings of the 28th international conference on computational linguistics, 2020, pp. 2296–2309. doi: 10.18653/v1/2020.coling-main.208.

A. Uchendu, Z. Ma, T. Le, R. Zhang, and D. Lee, “TURINGBENCH: A benchmark environment for turing test in the age of neural text generation,” in Findings of the association for computational linguistics: EMNLP 2021, 2021, pp. 2001–2016. doi: 10.18653/v1/2021.findings-emnlp.172.

Y. Sari, M. Stevenson, and A. Vlachos, “Topic or style? Exploring the most useful features for authorship attribution,” in Proceedings of the 27th international conference on computational linguistics, Association for Computational Linguistics, 2018, pp. 343–353. Available: https://aclanthology.org/C18-1029/

S. Gehrmann, H. Strobelt, and A. M. Rush, “GLTR: Statistical detection and visualization of generated text,” in Proceedings of the 57th annual meeting of the association for computational linguistics: System demonstrations, 2019, pp. 111–116. doi: 10.18653/v1/P19-3019.

D. Ippolito, D. Duckworth, C. Callison-Burch, and D. Eck, “Automatic detection of generated text is easiest when humans are fooled,” in Proceedings of the 58th annual meeting of the association for computational linguistics, 2020, pp. 1808–1822. doi: 10.18653/v1/2020.acl-main.164.

L. Fröhling and A. Zubiaga, “Feature-based detection of automated language models: Tackling GPT-2, GPT-3 and Grover,” PeerJ Computer Science, vol. 7, p. e443, 2021, doi: 10.7717/peerj-cs.443.

E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn, “DetectGPT: Zero-shot machine-generated text detection using probability curvature,” in Proceedings of the 40th international conference on machine learning, in Proceedings of machine learning research, vol. 202. 2023, pp. 24950–24962. Available: https://proceedings.mlr.press/v202/mitchell23a.html

J. Su, T. Zhuo, D. Wang, and P. Nakov, “DetectLLM: Leveraging log rank information for zero-shot detection of machine-generated text,” in Findings of the association for computational linguistics: EMNLP 2023, 2023, pp. 12395–12412. doi: 10.18653/v1/2023.findings-emnlp.827.

J. Kirchenbauer, J. Geiping, Y. Wen, J. Katz, I. Miers, and T. Goldstein, “A watermark for large language models,” in Proceedings of the 40th international conference on machine learning, in Proceedings of machine learning research, vol. 202. 2023, pp. 17061–17084. Available: https://proceedings.mlr.press/v202/kirchenbauer23a.html

B. Guo et al., “How close is ChatGPT to human experts? Comparison corpus, evaluation, and detection,” arXiv preprint arXiv:2301.07597, 2023, doi: 10.48550/arXiv.2301.07597.

V. Verma, E. Fleisig, N. Tomlin, and D. Klein, “Ghostbuster: Detecting text ghostwritten by large language models,” in Proceedings of the 2024 conference of the north american chapter of the association for computational linguistics: Human language technologies, 2024, pp. 1702–1717. doi: 10.18653/v1/2024.naacl-long.95.

R. Mohawesh et al., “Fake reviews detection: A survey,” IEEE Access, vol. 9, pp. 65771–65802, 2021, doi: 10.1109/ACCESS.2021.3075573.

J. Hay, B.-L. Doan, F. Popineau, and O. A. Elhara, “Representation learning of writing style,” in Proceedings of the sixth workshop on noisy user-generated text, 2020, pp. 232–243. doi: 10.18653/v1/2020.wnut-1.30.

C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, pp. 206–215, 2019, doi: 10.1038/s42256-019-0048-x.

D. Chicco and G. Jurman, “The advantages of the matthews correlation coefficient (MCC) over F1 score and accuracy in binary classification evaluation,” BMC Genomics, vol. 21, no. 1, p. 6, 2020, doi: 10.1186/s12864-019-6413-7.

A. Pagnoni, M. Graciarena, and Y. Tsvetkov, “Threat scenarios and best practices to detect neural fake news,” in Proceedings of the 29th international conference on computational linguistics, 2022, pp. 1233–1249. Available: https://aclanthology.org/2022.coling-1.106/

D. Macko et al., “MULTITuDE: Large-scale multilingual machine-generated text detection benchmark,” in Proceedings of the 2023 conference on empirical methods in natural language processing, 2023, pp. 9960–9987. doi: 10.18653/v1/2023.emnlp-main.616.

F. Pedregosa et al., “Scikit-learn: Machine learning in python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011, Available: https://www.jmlr.org/papers/v12/pedregosa11a.html

L. Dugan et al., “RAID: A shared benchmark for robust evaluation of machine-generated text detectors,” in Proceedings of the 62nd annual meeting of the association for computational linguistics, 2024, pp. 12463–12492. doi: 10.18653/v1/2024.acl-long.674.

K. Krishna, Y. Song, M. Karpinska, J. Wieting, and M. Iyyer, “Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense,” in Advances in neural information processing systems, 2023. Available: https://arxiv.org/abs/2303.13408

V. S. Sadasivan, A. Kumar, S. Balasubramanian, W. Wang, and S. Feizi, “Can AI-generated text be reliably detected? Stress testing AI text detectors under various attacks,” Transactions on Machine Learning Research, 2025, Available: https://openreview.net/forum?id=OOgsAZdFOt

W. Liang, M. Yuksekgonul, Y. Mao, E. Wu, and J. Zou, “GPT detectors are biased against non-native english writers,” Patterns, vol. 4, no. 7, p. 100779, 2023, doi: 10.1016/j.patter.2023.100779.

D. Weber-Wulff et al., “Testing of detection tools for AI-generated text,” International Journal for Educational Integrity, vol. 19, no. 1, p. 26, 2023, doi: 10.1007/s40979-023-00146-z.

Downloads

Published

2026-09-30

Issue

Section

Articles

How to Cite

Lightweight Detection of AI-Generated Product Reviews Using Stylometric Features and Logistic Regression. (2026). Journal of Advanced Intelligent Computing and Informatics, 1(1), 1-12. https://landing.wrunion.org/ojs/index.php/jaici/article/view/22