Data Quality-Aware Detection of Fraudulent Job Advertisements Using Text and Missing-Metadata Patterns

Authors

  • Karthik Juluri Independent Researcher Author
  • Krishna Kant Pandey Department of Mechatronics, Manipal University Jaipur, Jaipur, Rajasthan Author

Keywords:

Fraudulent Job Advertisements, Data Quality, Missing Metadata, Natural Language Processing, TF-IDF, Imbalanced Classification, Logistic Regression

Abstract

Fraudulent job advertisements can expose applicants to financial loss, identity theft, and wasted effort. Most automated detectors emphasize wording, although an advertisement’s missing or incomplete metadata may provide complementary evidence about how it was created. This paper proposes a data quality-aware decision-support framework that joins term frequency–inverse document frequency (TF–IDF) text features with explicit missingness indicators, aggregate completeness measures, and simple text-quality variables. A reproducible CPU-only simulation uses the open Employment Scam Aegean Dataset and a fixed stratified holdout. After removing 330 repeated predictor-identical records, the text-only logistic model obtained precision 0.673, recall 0.830, F1-score 0.743, balanced accuracy 0.905, and precision–recall area 0.877. Fusion increased recall to 0.889, balanced accuracy to 0.925, and precision–recall area to 0.886, but reduced precision to 0.535 and F1-score to 0.668. Metadata quality alone was considerably weaker. Thus, missingness provided a modest recall-oriented signal rather than a universal improvement. The practical contribution is a transparent ablation protocol for deciding when completeness patterns warrant inclusion in human-reviewed recruitment-fraud triage.

References

S. Vidros, C. Kolias, G. Kambourakis, and L. Akoglu, “Automatic detection of online recruitment frauds: Characteristics, methods, and a public dataset,” Future Internet, vol. 9, no. 1, p. 6, 2017, doi: 10.3390/fi9010006.

B. Alghamdi and F. Alharby, “An intelligent model for online recruitment fraud detection,” Journal of Information Security, vol. 10, no. 3, pp. 155–176, 2019, doi: 10.4236/jis.2019.103009.

S. Lal, R. Jiaswal, N. Sardana, A. Verma, A. Kaur, and R. Mourya, “ORFDetector: Ensemble learning based online recruitment fraud detection,” in 2019 Twelfth International Conference on Contemporary Computing (IC3), 2019, pp. 1–5, doi: 10.1109/IC3.2019.8844879.

J. Li, Y. Li, H. Han, and X. Lu, “Exploratory methods for imbalanced data classification in online recruitment fraud detection: A comparative analysis,” in Proceedings of the 2021 4th International Conference on Computing and Big Data, 2021, pp. 75–81, doi: 10.1145/3507524.3507537.

S. Mahbub, E. Pardede, and A. S. M. Kayes, “Online recruitment fraud detection: A study on contextual features in Australian job industries,” IEEE Access, vol. 10, pp. 82776–82787, 2022, doi: 10.1109/ACCESS.2022.3197225.

K. Nanath and L. Olney, “An investigation of crowdsourcing methods in enhancing the machine learning approach for detecting online recruitment fraud,” International Journal of Information Management Data Insights, vol. 3, no. 1, p. 100167, 2023, doi: 10.1016/j.jjimei.2023.100167.

S. Seaman, J. Galati, D. Jackson, and J. Carlin, “What is meant by ‘missing at random’?” Statistical Science, vol. 28, no. 2, pp. 257–268, 2013, doi: 10.1214/13-STS415.

T. Emmanuel, T. Maupong, D. Mpoeleng, T. Semong, B. Mphago, and O. Tabona, “A survey on missing data in machine learning,” Journal of Big Data, vol. 8, p. 140, 2021, doi: 10.1186/s40537-021-00516-9.

M. Saar-Tsechansky and F. Provost, “Handling missing values when applying classification models,” Journal of Machine Learning Research, vol. 8, pp. 1623–1657, 2007.

R. Mitra, S. F. McGough, T. Chakraborti, C. Holmes, R. Copping, N. Hagenbuch, S. Biedermann, J. Noonan, B. Lehmann, A. Shenvi, X. V. Doan, D. Leslie, G. Bianconi, R. Sanchez-Garcia, A. Davies, M. Mackintosh, E.-R. Andrinopoulou, A. Basiri, C. Harbron, and B. D. MacArthur, “Learning from data with structured missingness,” Nature Machine Intelligence, vol. 5, no. 1, pp. 13–23, 2023, doi: 10.1038/s42256-022-00596-z; arXiv:2304.01429.

G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Information Processing & Management, vol. 24, no. 5, pp. 513–523, 1988, doi: 10.1016/0306-4573(88)90021-0.

C. Rudin, “Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead,” Nature Machine Intelligence, vol. 1, no. 5, pp. 206–215, 2019, doi: 10.1038/s42256-019-0048-x; arXiv:1811.10154.

F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay, “Scikit-learn: Machine learning in Python,” Journal of Machine Learning Research, vol. 12, pp. 2825–2830, 2011.

P. Branco, L. Torgo, and R. P. Ribeiro, “A survey of predictive modeling on imbalanced domains,” ACM Computing Surveys, vol. 49, no. 2, pp. 31:1–31:50, 2016, doi: 10.1145/2907070; arXiv:1505.01658.

T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLOS ONE, vol. 10, no. 3, p. e0118432, 2015, doi: 10.1371/journal.pone.0118432.

K. H. Brodersen, C. S. Ong, K. E. Stephan, and J. M. Buhmann, “The balanced accuracy and its posterior distribution,” in 2010 20th International Conference on Pattern Recognition, 2010, pp. 3121–3124, doi: 10.1109/ICPR.2010.764.

T. Gebru, J. Morgenstern, B. Vecchione, J. W. Vaughan, H. Wallach, H. Daume III, and K. Crawford, “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021, doi: 10.1145/3458723; arXiv:1803.09010.

E. M. Bender and B. Friedman, “Data statements for natural language processing: Toward mitigating system bias and enabling better science,” Transactions of the Association for Computational Linguistics, vol. 6, pp. 587–604, 2018, doi: 10.1162/tacl_a_00041; ACL Anthology: Q18-1041.

N. Mehrabi, F. Morstatter, N. Saxena, K. Lerman, and A. Galstyan, “A survey on bias and fairness in machine learning,” ACM Computing Surveys, vol. 54, no. 6, pp. 115:1–115:35, 2021, doi: 10.1145/3457607; arXiv:1908.09635.

J. Lu, A. Liu, F. Dong, F. Gu, J. Gama, and G. Zhang, “Learning under concept drift: A review,” IEEE Transactions on Knowledge and Data Engineering, vol. 31, no. 12, pp. 2346–2363, 2019, doi: 10.1109/TKDE.2018.2876857; arXiv:2004.05785.

Downloads

Published

2026-09-30

Issue

Section

Articles