Multimodal Temporal Graph Neural Networks for Detecting and Disrupting Dark-Web Cybercrime Communities: A Vendor-Disjoint Inductive Evaluation Finding No Measurable Benefit from the Graph
Oleksandr Shchur, Maximilian Mumme, Aleksandar Bojchevski, Stephan Günnemann · arXiv (Cornell University) · 2018
For: Multimodal Temporal Graph Neural Networks for Detecting and Disrupting Dark-Web Cybercrime Communities: A Vendor-Disjoint Inductive Evaluation Finding No Measurable Benefit from the Graph This record presents an academic research paper together with its complete reproducible experimental artefacts. The study evaluates MM-TGNN, a multimodal temporal graph neural network designed for detecting and disrupting dark-web cybercrime communities, under a strict vendor-disjoint inductive protocol designed to distinguish genuine architectural benefit from the artefacts that transductive evaluation and label leakage can introduce. Research Question Does the graph structure and temporal machinery of MM-TGNN contribute measurable benefit over feature-only classification on the dark-web community-detection task, once label leakage and transductive evaluation are controlled for? Methodology We evaluate MM-TGNN and four baseline GNN architectures (Static GCN, GraphSAGE, Static MM-GCN, EvolveGCN-O) alongside two reference baselines — a weighted-prior chance floor and a graph-free Sentence-BERT logistic regression ceiling — using ten-fold cross-validation (five-fold × two repeats, vendor-disjoint, stratified by market) on a real Gwern-archive-derived corpus (1,350 nodes, 2,408 edges, 7 markets, 150 vendors). A sanitized-text control verifies that the SBERT signal is legitimate rather than leaked. Paired Wilcoxon signed-rank tests assess significance. Key Findings The graph-free SBERT-LogReg ceiling (macro-F1 = 0.1447 ± 0.0434) is the highest-scoring method and does not significantly outperform MM-TGNN (0.1130 ± 0.0505; p = 0.128), indicating a failure to reject the null: the graph provides no measurable benefit over feature-only classification on this task. MM-TGNN significantly outperforms the chance floor (p = 0.012) and two of the three static GNNs (p < 0.05), suggesting the GNNs learn real signal above chance. Static message passing appears to actively degrade performance below the graph-free ceiling on this sparse, homophily-poor graph. Contribution A disciplined negative result about the named architecture, accompanied by a three-part evaluation protocol (vendor-disjoint inductive splits, graph-free baselines, chance floors) that the dark-web GNN subfield may adopt to distinguish genuine architectural benefit from evaluation artefacts. Archive Contents · MMTGNN_InductiveEvaluation_AcademicPaper_2026-08-13.docx — full paper (Word). · main.pdf — typeset PDF of the paper. · run_final_experiment.py — complete reproducible Python implementation (812 lines; requires torch, torch-geometric, transformers, sentence-transformers, scikit-learn, scipy, networkx, matplotlib, pandas). · arXiv_submission.zip — arXiv-ready LaTeX source package (main.tex, refs.bib, figures, PDF). Dataset Munhouiani drug-listings CSV, derived from the Gwern Darknet Market Archives, publicly available at https://github.com/munhouiani/Drug-Listings-Dataset. Limitations The image modality (ResNet-50) is untested because the dataset contains no product images; the evaluation uses a single dataset; the temporal credit assignment is truncated (GRU weights detached between snapshots); and the repeated-CV structure violates the independence assumption of the paired Wilcoxon test, making the positive significances anti-conservative. The central non-significant claim (p = 0.128) is made more, not less, defensible by this anti-conservatism.