Machine Learning for the Detection of Fake News on Social Media: A Critical Narrative Review of Methods, Evaluation Validity and Deployment Readiness
Mary Oluwakemi Abioye *
Department of Data Science, University of West England, Bristol, England.
*Author to whom correspondence should be addressed.
Abstract
Automated detection of false and misleading news circulating on social media has become one of the most heavily researched applications of machine learning in the computational social sciences, yet the practical value of the resulting systems remains contested. Reported classification accuracies on standard benchmarks frequently approach ceiling, while independent assessments of the same model families under temporal, topical and adversarial shift record substantially weaker performance. This review examines the disjunction between benchmark success and operational capability, and asks what the accumulated evidence genuinely supports. The literature was identified through structured searching of openly accessible scholarly indexes and metadata registries, supplemented by backward and forward citation tracking from recent reviews and by targeted retrieval of methodological and institutional sources. Evidence was appraised for construct validity of labels, evaluation design, transparency, replication and external validity, and was synthesised thematically rather than catalogued study by study. Four findings emerge with reasonable confidence. Label provenance, rather than model architecture, is the dominant determinant of what a classifier learns, and source-level labelling propagates publisher-specific stylistic signals that inflate apparent accuracy. Social-context and propagation models achieve stronger discrimination than content-only models but forfeit the early-detection window and depend on platform data whose availability has narrowed. Multimodal and evidence-retrieval systems address genuine failure modes of text-only classification, although gains are reported on heterogeneous benchmarks that resist direct comparison. Large language models function simultaneously as a generative threat that decouples writing style from veracity and as a source of reasoning and rationale that improves small detectors, with the second role better evidenced than autonomous zero-shot verification. Evidence remains concentrated in English-language, politically framed, text-dominant corpora drawn from a small number of platforms, and almost no study measures downstream effects on audiences or on fact-checking workflows. Progress now depends less on architectural novelty than on label construction, temporally honest evaluation and outcome measurement beyond the confusion matrix.
Keywords: Misinformation detection, benchmark validity, domain generalisation, natural language processing, large language models, automated fact-checking, social media analytics