📘 : Nexus Global Research Journal of Multidisciplinary (NGRJM) Volume 2, Issue 6, 2026 (Page : 204-211)

ABSTRACT:

Crime prediction in developing nations faces a fundamental challenge: the absence of structured, incident-level data prevents the application of modern machine learning methods that have transformed predictive policing in data-rich environments. This paper addresses this challenge directly by proposing and evaluating a synthetic dataset methodology for criminological research in Nigeria. A dataset of 3,150 synthetic crime incidents spanning Yenagoa, Bayelsa State (2017 to 2023) was constructed from aggregate official statistics and validated against National Bureau of Statistics records. A hybrid ensemble model comprising XGBoost, Random Forest, and Gradient Boosting trained on 44 engineered demographic, spatial, and temporal features achieved an overall accuracy of 43.0 percent and balanced accuracy of 31.0 percent, exceeding random expectation but falling substantially below operational thresholds reported in data-rich settings. Complete classification failure on Fraud and Deception categories, combined with 226 mutual misclassifications between Property and Violent Crimes, empirically demonstrates that demographic and spatial features lack the discriminative power necessary for typological crime classification. Feature importance analysis revealed a diffuse distribution in which the top ten predictors explained only 36.4 percent of predictive variance. These findings quantify Nigeria’s data infrastructure gap, establish a reproducible synthetic data methodology for data-scarce criminological research, and directly motivate contextual feature augmentation as a priority for future work.

Keywords: Crime prediction, synthetic data, predictive policing, machine learning, data scarcity, feature engineering