2026-09-11 | | Total: 8
Graph Neural Networks (GNNs) are powerful models for handling attributed graphs in tasks such as classification, link prediction, and community detection, as they enable the aggregation of information from both structural and semantic sources. However, progress in community detection is hindered by the lack of high-quality datasets, since ground-truth community labels are often unavailable and most algorithms proposed in recent literature rely on the same benchmark datasets for model training and evaluation. To address this issue, attributed random graph generators are commonly employed to create synthetic graphs for assessing the strengths and limitations of GNN-based models. Nevertheless, most existing generators rely heavily on power-law degree distributions, despite recent evidence indicating that scale-free networks are rare, particularly in social network contexts. Moreover, state-of-the-art attributed graph generators provide limited flexibility, as they do not allow users to construct communities with varying densities, degree distributions, and sub-community structures. To overcome these limitations, we introduce the Synthetic Community-Aware Attributed Graph Generator (SynCo), a graph generation algorithm that allows users to control the node degree distribution and sub-community structure. We evaluate SynCo across three different tasks: graph mimicking, hyperparameter evaluation, and node clustering tuning. The results show that our model outperforms state-of-the-art approaches in synthetic graph generation and data augmentation, while preserving the original distributions of duplicated and augmented datasets, as confirmed by statistical tests well know in literature. We also demonstrate the ability of SynCo to generate nodes in large scale, up to 2.1 million nodes.
In a mechanistic model of the dawn chorus, Kaye showed that heterogeneous activation thresholds and a shared feedback signal determined by the population's active fraction can produce abrupt collective activation and hysteresis. We extend this mechanism to a network of interacting agents. Each node has a continuous activation level, and a nonnegative row-stochastic matrix determines how node activities contribute to individual feedback. We prove that sufficiently weak feedback yields a unique globally attracting equilibrium. For all feedback strengths, the homogeneous dynamics exactly reproduce Kaye's scalar equation; consequently, network topology does not alter the folds or cusp of the homogeneous branch, and no heterogeneous mode becomes unstable before the homogeneous mode. For equitable partitions, the network admits an exact quotient system in which nodes within a block receive the same aggregate input from every block. When blocks are uncoupled, the quotient reduces to independent copies of Kaye's scalar equation. We show that every assignment of stable scalar equilibria to blocks persists under sufficiently weak interblock coupling, producing a combinatorial family of stable quotient equilibria that lift to stable full-network equilibria and remain under small perturbations that break exact equitability. For two symmetrically coupled blocks, branches with unequal block activities terminate at a pair of symmetry-related cusp bifurcations. Near the onset of bistability, we derive scaling laws for the interblock coupling at which these bifurcations occur, the activity difference between the blocks at bifurcation, and the corresponding shift of the external stimulus from the scalar cusp. Numerical continuation confirms the scaling laws for gamma, logistic, and normal threshold distributions.
We develop a quantum approach to spectral feature extraction from the density of states (DOS) of a problem-dependent Hamiltonian, and apply it to machine learning on signed graphs. We propose to embed a signed graph as an Ising model instance with positive and negative interactions, and use the standardized moments of the Ising DOS as features for learning. We show that these moments count signed closed walks, are switching-invariant, and are size-free by construction. As a benchmark, we target learning the frustration index, an NP-hard measure of structural balance that can be labeled exactly at moderate size. At zero field, the models can be sampled classically, allowing the quantum extraction procedure to be certified against exact ground truth. We propose DOS-QPE, a phase estimation on a purified maximally mixed probe, which samples the spectral density with orders of magnitude fewer shots than Hadamard test-based trace sampling and feeds the resulting features directly into classically trained models. On $1.4\times10^5$ labeled graphs the exact DOS determines the frustration index, and five moments recover it with a mean error of 0.4, well below one sign flip. Beyond zero field, the underlying trace-estimation problem is DQC1-complete, providing access to spectral features for which no efficient classical sampling method is known. Our work opens routes towards quantum applications in social network balance analysis, spin-glass studies, correlation clustering, and protein-interaction networks.
Where a mobile signal is available shapes who can work, learn, bank, seek health care and respond to crises in the digital age, yet no globally consistent, sub-national record of mobile network coverage exists. We present such a record: annual 1km maps of the probability of 2G, 3G and 4G coverage for 214 countries and territories for the years 1999 to 2030. The maps are produced by three independent models: a calibrated machine-learning model, a techno-economic simulator of network build-out, and a spatial deep-learning model. The three estimates are then combined, per country and technology and in proportion to their measured accuracy, into a single best estimate with per-pixel 90% uncertainty bands; all four layers are released as part of the dataset. Because mobile roll-out closely follows a country's socio-economic conditions (population distribution, electrification, physical infrastructure), the models are grounded in existing geospatial data and tuned on 2,409 quality-screened operator-reported coverage maps, which are available up to 2020. For 2021--2024 the maps are predicted from recent geospatial data alone; for 2025--2030 they are extrapolated from demographic and infrastructure projections. On countries held out during training, the machine-learning model attains AUC 0.89--0.92. Baseline comparisons and the combined product's external validation are reported in Technical Validation. The dataset supports mapping the global digital divide, linking connectivity to household-survey outcomes, and humanitarian and infrastructure planning.
Social media platforms have become central to shaping political discourse, serving as arenas where narratives form and evolve, influencing public opinion. Identifying and analyzing these narratives, particularly when they compete across different political ideologies, is crucial for understanding the dynamics of modern political communication. This paper presents an unsupervised framework for identifying and characterizing competing narratives in political discourse on social media, focusing on German politicians' tweets. The framework employs a multi-stage pipeline that integrates natural language processing techniques such as topic modeling, event detection, and event linking. By forming data into coherent stories and uncovering the distinct perspectives of user communities, the system is able to detect the key competing narratives, highlighting the divergent framings and conflicts surrounding trending political topics. Two case studies on polarizing political issues demonstrate the efficacy of the methodology, showcasing its ability to uncover and analyze divergent viewpoints. The findings contribute to the broader understanding of how narratives propagate within the digital public sphere and offer insights for policymakers, social media platforms, and researchers interested in monitoring political discourse.
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
We consider a fully connected gossip network of $n$ nodes that track a binary continuous-time Markov source through a strategic sender transmitting updates under a communication budget at a rate that depends on the source state. The receivers exchange packets through gossip and decide whether to follow the sender. We model this interaction as a Stackelberg game and analyze it through a stochastic hybrid systems (SHS) framework. We prove that the sender's budget constraint binds at every interior optimum, reducing its problem to a one-dimensional search on the budget line. When the sender pushes its preferred state at the higher rate, the receivers gossip at the highest available rate. Gossip has no direction of its own and works against the asymmetry in the sender's policy rather than reinforcing it. We prove that an optimistic Stackelberg equilibrium exists, and that it is unique and explicitly characterized whenever a policy on the strategic half of that line is feasible at the gossip cap. Monte Carlo simulations agree with the analytical recursion.
Informational active matter research has shown how measurement-informed decisions produce collective order, so far in systems that reach consensus. We identify a collective information engine that orders by differentiation instead, and construct a minimal instance using anti-coordination games, the simplest game class in which differentiated role information has value. Agents infer their role from a noisy social signal, and role-following action feeds back into that signal, shaping the incentive to follow roles. The schemas prescribing roles are themselves dynamic: resources accrued through coordinated role-play combine with population variability to shape and reinforce the schemas that generated them. The model thereby operationalizes Sewell's duality of schemas and resources, the influential but hitherto qualitative resolution of sociology's structure--agency debate. The engine ignites when a loop gain $Λ$---the product of identity persistence, cognitive capacity, and social feedback channel fidelity---exceeds one. For a schema repertoire, roles emerge in a bifurcation cascade whose functional form is fixed by the repertoire's eigenvalue spectrum, ranging from monitorable logarithmic sequences to deceptive spike-then-avalanche onsets. Resource accumulation turns schema competition into a replicator dynamics that selects the cascade type endogenously. Subcritical behavioral covariance reveals that type before onset, enabling early detection, and feedback channel parameters bias which type is selected. For agents interacting on a platform, adaptive platform design then becomes a control lever to throttle emergent coordination, e.g. for mitigating risks from distributional AI. Joining game theory, collective dynamics, and information engines, we open a route to an information thermodynamics of agent.