2026-08-18 | | Total: 142
This paper investigates a movable-element simultaneously transmitting and reflecting reconfigurable intelligent surface (ME-STARS) assisted rate-splitting multiple access (RSMA) system under imperfect channel state information (CSI). Unlike conventional STARS with fixed element positions, the elements of ME-STARS can be repositioned within a predefined region, providing additional spatial degrees of freedom for improving the cascaded transmitter--STARS--user channels. To exploit this flexibility while accounting for CSI uncertainty, we formulate a robust sum-rate maximization problem that jointly optimizes the transmit beamforming, common-rate allocation, reflection and transmission coefficients, and ME-STARS element positions, subject to transmit-power, user-rate, minimum inter-element spacing, and movement-region constraints. The resulting problem is highly non-convex due to the strong coupling among the design variables and the position-dependent channels. To address this challenge, an iterative optimization framework is developed in which the transmit beamforming, STARS coefficients, and element positions are successively optimized through tractable convex reformulations. In particular, the element positions are updated sequentially using a majorization--minimization (MM) framework, where quadratic surrogate functions are constructed from the first- and second-order derivatives of the position-dependent channels while preserving the minimum inter-element spacing constraint. Simulation results demonstrate that the proposed ME-STARS design consistently outperforms the considered benchmark schemes. Moreover, the performance gains remain significant under increasing CSI uncertainty, highlighting the effectiveness of element repositioning for robust RSMA transmission.
Time-critical interactive systems increasingly require ultra-low-latency device identification for multiple users, yet prevailing approaches such as passwords, QR codes, and RFID/NFC are constrained by human input, frame-based sensing, or near-contact range. This paper presents ECO-ID, an event-camera-based optical system for multi-user, ultra-low-latency identification over visible light communication (VLC). Leveraging microsecond-resolution, asynchronous observations of brightness transitions, ECO-ID employs a spatiotemporal coding design: disjoint LED subsets provide spatial separation among users, while user-specific timing delays encode identities without inter-user synchronization. The optical channel and event-driven sensing reduce full-scene capture relative to frame cameras and limit the RF attack surface, while enabling rapid token verification with freshness and replay protection. We implement a prototype and demonstrate that ECO-ID can practically achieve approximately 99.8\% localization and 98.7\% identification with 0.64 ms mean latency, while theoretically supporting identification at the scale of tens of concurrent users. Overall, ECO-ID provides a fast, privacy-conscious, and security-aware alternative for scalable multi-user identification in time-critical interactive environments.
Lung nodule segmentation in computed tomography is essential for extracting clinically relevant information for lung cancer assessment and treatment planning. Foundation models have shown notable segmentation capabilities, but state-of-the-art approaches often depend on input prompts, such as points or boxes, making their performance sensitive to prompt quality and placement. Understanding the limitations and constraints of prompt-based foundation models is therefore essential for designing reliable medical image segmentation solutions. In this work, we investigate how prompt quality affects foundation models performance for lung nodule segmentation. We further propose a synthetic prompt-generation model to test if the dependence on manually provided prompts can be mitigated by generating synthetic prompts that can also improve segmentation performance. Perturbation experiments show that bounding box prompts generally outperform point prompts, while latest specialized medical imaging models achieve better performance than general purpose ones. The proposed approach obtains a Dice coefficient of 0.85, suggesting that synthetic prompt generation as a promising strategy for lung nodule segmentation with foundation models.
We investigate power-efficient multiuser integrated sensing and communication (ISAC) assisted by an element-grouping extremely large-scale intelligent reflecting surface (EG-XL-IRS). The grouping pattern is designed using slowly varying statistical channel state information (S-CSI), so that both IRS-related channel acquisition and online passive beamforming operate in the group domain rather than the element domain. We reveal a fundamental gain-rank tradeoff induced by element grouping: phase-consistent grouping can coherently enhance selected deterministic propagation components, while excessive concentration on a common deterministic mode can reduce the effective spatial rank of the multiuser channel and, for extended targets, the diversity of desired-scatterer responses. Motivated by this observation, we develop a task-adaptive rank-aware grouping strategy that balances weak-user enhancement and target-scatterer illumination while preserving task-relevant spatial dimensions. For each candidate grouping pattern, the transmit covariances and group-wise reflection phases are jointly optimized under communication and sensing quality-of-service constraints, followed by physical phase recovery and feasibility verification. Numerical results show that the proposed design substantially reduces the required transmit power compared with representative grouping benchmarks under the same grouping dimension and online optimization budget.
Design structure matrices (DSMs) are used to comprehensively represent complex systems. They visualize and describe the dependencies between various variables, processes, states, and events. As such they are used in several system engineering approaches, such as requirement and interface management, fault detection, and supervisory control. Currently, a DSM is typically built from knowledge of experts. This may lead to an incomplete or imbalanced DSMs. For instance, elements and links might be missing or superfluous. In this article, we propose a novel method to acquire the DSM using state-of-the-art network identification methods. This demonstrates a proof-of-principle of identifying DSMs from data as an additional tool to the standard heuristic approach. In the future, we plan to embed DSMs in system design and supervisory controllers. We apply this technique to identify the DSM of a fusion reactor modelled by a five-chamber plasma model describing the transport in a tokamak.
Individually measuring head-related transfer functions (HRTFs) at scale remains a central challenge for personalised spatial audio, motivating growing interest in synthetic HRTFs. We evaluated the numerical, computational, and behavioural validity of synthetic HRTFs, generated through the boundary element method simulation using Mesh2HRTF, against measured and KEMAR HRTFs using the Extended SONICOM dataset. Across 200 subjects, synthetic HRTFs deviated less from measured than KEMAR in interaural time and level differences, but residual errors, together with elevated spectral distortion, concentrated at low, rear elevations. This is consistent with the omission of torso geometry from the synthesis pipeline. Two computational models revealed a corresponding pattern of predicted localisation errors, with synthetic HRTFs positioned between measured and KEMAR. In a virtual reality localisation task (N = 20), synthetic HRTFs matched measured on every polar metric, while KEMAR was significantly worse. However, behavioural error clustered around the front-back midline regardless of condition, not at the low elevations implicated numerically or by the models. A separate spatial release from masking task (N = 18) showed no effect of HRTF type. Together, these results indicate that high-resolution synthetic HRTFs preserve behavioural localisation performance, despite discrepancies between the numerical/model-predicted bias and the spatial pattern of behavioural error.
This paper presents a real-time orthogonal frequency-division multiplexing (OFDM) radar embedded in the OpenAirInterface (OAI) 5G base-station process. The radar removes communication symbols by regularized element-wise division and performs range-Doppler processing and ordered-statistic constant-false-alarm-rate detection online without modifying the 5G waveform. The implemented system provides 2.57 m nominal range resolution and 0.28 m/s velocity resolution. Hardware measurements identify and mitigate several implementation-specific limitations, most notably a deterministic carrier-dependent transmit-receive phase rotation on a Universal Software Radio Peripheral (USRP) X300. Selecting a tuning-grid-aligned carrier improves mean-removal clutter suppression from -16.4 dB to 38.0 dB and reduces coherent-integration loss from 19.81 dB to 0.27 dB. The measured processing gain closely agrees with its predicted value. Instrumented worker timing confirms real-time operation, with a conservative 58.4 percent utilization bound and no dropped soundings. A custom E2 service model, E2SM-RADAR, exports detections and a compact slow-time product to a near-real-time RAN Intelligent Controller. Live end-to-end operation demonstrates reliable delivery and supports controller-side tracking, micro-Doppler analysis, and classification. With a commercial user equipment connected on the same carrier, measurements show no measurable difference in downlink throughput estimate with sensing enabled, while the radar sensing bandwidth follows the scheduler allocation.
An accurate estimation of the state of health (SOH) underpins a safe and optimized use of the battery system. Although compelling, data-driven SOH estimation models typically require large amounts of high-quality labeled cycling data, while in practice such labels are often sparse in both quantity and coverage. Therefore, in this work, we propose a degradation-aligned self-supervised learning (SSL) framework based on a convolutional neural network-gated recurrent unit (CNN-GRU) model, which learns aging-consistent representations from unlabeled data through a cycle-order ranking objective as the pretext task for pretraining, thereby enabling robust SOH estimation after fine-tuning on sparsely labeled data. Test results showcase that the proposed ranking-based SSL approach proves to endow the pretrained model with degradation-aligned information from unlabeled data, and after fine-tuning the model can carry out accurate, robust SOH estimation, even when only an extremely limited amount of 1% of unevenly distributed labeled training data is available, where the MAE of 1.718% and RMSE of 2.329% can be achieved on the test cell. In addition, in-depth analyses are presented regarding the influences of label distribution of battery degradation data. We believe this work could shed new light on SOH estimation of lithium-ion batteries under label sparsity in real-world applications.
Oscillation is a critical issue that power systems have long faced. Especially over the past two decades, with the large-scale inte-gration of renewable energy into the grid, oscillation problems have posed a serious threat to the secure operation of power systems. However, the current literature has not fully explained the oscillation mechanism of renewable energy integrated power systems (REIPSs). In this paper, the underlying mechanism of stimulated oscillations is explored, with novel analytical methods proposed. Firstly, it is explained from both mathematical formu-las and physical interpretations that for an oscillation mode characterized by a pair of complex conjugate poles, the oscilla-tion risk under disturbance depends on the relative positional relationship between the corresponding poles and all other poles and zeros on the complex plane, rather than their standalone locations, i.e., the stability perceived by classical theory. Then the underlying mechanism of high amplitude oscillations induced by closely-located poles under even slight disturbance is clarified. On this basis, a theoretical framework for stimulated oscilla-tions applicable to REIPSs, covering its definition, mechanism, and methods, is proposed. Finally, this paper discusses the rela-tionship between the stimulated oscillation theory proposed herein and the classical stability-based theory, revealing that the research findings surpass rather than negate the classical theo-ries.
Objective assessment of learning remains a fundamental challenge in education. Electroencephalography (EEG) provides a direct, non-invasive window into the neural correlates of knowledge acquisition, including cognitive familiarity. This study benchmarks fifteen machine learning (ML) and deep learning (DL) models for EEG-based familiarity prediction across two cognitive domains: faces (factual knowledge) and mathematical equations (conceptual knowledge). Using continuous EEG data from 23 participants, we extract spectral features (Power Spectral Density) across six frequency bands. We show that while standard stratified cross-validation yields artificially high classification performance (up to 0.9853 F1-score using CNN) due to temporal leakage across neighboring epochs, a rigorous trial-independent validation (Group K-Fold) drops the peak performance to 0.6038 F1-score (using CNN), which is still statistically significant above the 25% chance level. This highlights the critical necessity of trial-independent evaluation to avoid overestimating model generalizability. Furthermore, feature importance and SHAP analysis reveal that temporal and frontal Gamma and Beta oscillations are the most critical biomarkers for familiarity. This work establishes a realistic benchmark for EEG-based cognitive monitoring in educational technologies.
Near-field antenna measurements underpin the characterization of electrically large apertures, yet the fidelity of the Near-Field to Far-Field (NF-FF) transformation depends on the reconstruction algorithm's assumptions and robustness to real-world imperfections, including those from drone-based scanning platforms. Classical FFT-based modal expansion is efficient on uniformly sampled canonical grids but fails when phase-coherent acquisition cannot be maintained. We address this via a phaseless NF-FF algorithm reconstructing the far field from amplitude-only data through iterative phase retrieval. When sampling becomes sparse or irregular, even amplitude-based methods break down, motivating the $\textbf{Adaptive Sparse Inverse Radiation Estimator (ASPIRE)}$, a full-complex inverse source framework that solves a Method-of-Moments problem over RWG basis functions via cascaded rSVD and regularized shrinkage. Mutual coupling between basis functions is explicitly resolved, improving reconstruction fidelity beyond coupling-agnostic inverse-source formulations. The solver is accelerated via a Multilevel Fast Multipole Method engine with Numba just-in-time compilation, achieving a $1.2\times$ reduction in matrix-vector product time and up to $15\times$ lower memory usage relative to dense evaluation at N=100K. Across frequency bands and positioning/truncation error scenarios, the pipeline sustains algorithmic stability and achieves sub-degree beamwidth reconstruction error. These results establish an error-aware framework for algorithm selection across fixed and drone-based near-field measurement platforms.
This paper outlines a sonification design to support fault detection in the transmission of I2S transport signals. I2S is a protocol for communicating real-time digital audio between integrated circuits that, while in wide and general use, does not include built-in error detection. Moreover, given the nature of the protocol transmission faults affecting timing, framing and alignment can be difficult to identify using conventional visual methods. The proposed design addresses this with an approach informed by Audification, wherein oversampling controls temporal rescaling to render protocol structure (SCK and WS) and payload data (SD) across separate stereo channels. A preliminary computational feasibility study was carried out to measure feature-space separability of I2S faults in the generated auditory representations as opposed to listener performance. It evaluates the design across several payload types and error conditions including jitter, bit-slip, and word-length errors. Class separability was assessed through clustering analyses of extracted features. The evaluation results show that while oversampling produces systematic changes in feature values, it does not meaningfully improve separability between error classes. However, a modest but consistent improvement in separability is observed as a function of the joint representation of structural and payload information across channels. The findings suggest that feature-space separability in sonified communication protocol data may be dependent on the integration of complementary information streams, rather than on signal scaling alone.
Terahertz time-domain spectroscopy (THz-TDS) based on air-plasma generation and balanced air-biased coherent detection offers gap-free broadband coverage, but individual continuous-scan traces are strongly affected by pulse-to-pulse fluctuations and electronic noise. Reaching a useful signal-to-noise ratio therefore requires averaging multiple traces, which directly increases measurement time. We propose a learned denoising approach that recovers high-quality THz waveforms from as few as one complete continuous delay sweep, referred to here as a single-scan trace. A compact one-dimensional residual U-Net is trained using two complementary strategies: a reference-supervised baseline that maps individual noisy traces to long-average reference waveforms, and a Noise2Noise approach that learns from pairs of independently acquired noisy traces without requiring a clean training target. Averaging the predictions of both models reduces systematic bias and yields a trace-reduction factor of approximately $5.4\times$ at $K=1$, meaning that one denoised trace achieves the reconstruction accuracy of averaging approximately five raw traces. The Noise2Noise model alone achieves $4.9\times$, outperforming both the reference-supervised baseline ($4.6\times$) and classical Wiener filtering ($3.2\times$). These results show that self-supervised learning from repeated noisy measurements can support faster continuous-scan THz-TDS without hardware modification.
This paper proposes a two-layer model predictive control (MPC) framework for the real-time operation of data centers integrated with on-site photovoltaic generation, battery energy storage, waste heat recovery, and district heating. The upper layer employs scenario-based stochastic optimization to jointly optimize intraday market participation, workload scheduling, and energy management under uncertainty. The lower layer adopts an adaptive tube-based MPC strategy that compensates short-term disturbances while tracking the dispatch references given by the upper layer. The framework further integrates multi-horizon forecasting to support real-time decision making. Microservice-based simulation studies under representative clear-sky and overcast operating conditions demonstrate that the proposed framework accurately tracks dispatch plans despite fast photovoltaic and workload fluctuations. Compared with single-layer control strategies, the adaptive lower-layer controller substantially reduces real-time dispatch deviations and the associated imbalance costs. In addition, the proposed framework naturally adapts to seasonal operating conditions and responds to carbon-aware operating signals, offering a practical approach for economically efficient, sustainable, and grid-supportive operation of future data centers.
Multi-step rollouts are essential for model-based reinforcement learning (RL) and predictive control, yet learned dynamics models often become unstable when recursively applied, leading to divergence and unreliable policy updates. This paper proposes a model-agnostic hybrid dynamics framework that blends a provably contracting nominal model with a flexible excursion model through an uncertainty-guided switching law. The switching signal is derived from calibrated epistemic uncertainty and activates only when the system leaves the nominal region, ensuring that each model operates within its reliability regime. Under clearly stated smoothness and boundedness assumptions, we show that the resulting hybrid predictor yields globally bounded recursive multi-step rollouts: trajectories remain Lyapunov-stable in the nominal region and exhibit at most affine growth during excursions. To illustrate the theory in practice, we instantiate the hybrid dynamics framework within a model-based RL scheme that uses real one-step transitions for value learning and hybrid rollouts for policy improvement. Experiments on a nonlinear Duffing oscillator demonstrate stable long-horizon prediction and improved cost-effort trade-offs relative to a stabilizing baseline.
The stringent energy-efficiency requirements of future Integrated Sensing and Communications (ISAC) systems are fundamentally challenged. Unlike conventional communication systems, ISAC transmitters must radiate significantly higher power to ensure reliable target detection, forcing the High-Power Amplifier (HPA) to operate closer to saturation, where nonlinear distortions become unavoidable. Consequently, the robustness of every candidate ISAC waveform to HPA nonlinearities must be carefully assessed. In this context, this paper investigates the robustness of Affine Filter Bank Modulation (AFBM), a recently proposed waveform that combines the delay-Doppler resilience of affine modulation with reduced Peak-to-Average Power Ratio (PAPR) and improved spectral containment. We develop a statistical characterization of the Ambiguity Function (AF) of the amplified AFBM waveform, deriving approximate expressions for its mean, variance, and Rician-distributed magnitude. Furthermore, a low-complexity Gaussian belief propagation receiver accounting for HPA nonlinearities is proposed for communication detection. Simulation results validate the analytical framework and demonstrate that AFBM preserves favorable sensing characteristics and robust Bit Error Rate (BER) performance even under severe nonlinear amplification.
Learning-based Model Predictive Control (MPC) using Gaussian processes (GPs) is an effective approach for safe control in the presence of model mismatch. High-probability safety guarantees typically require uncertainty bounds that hold uniformly over the entire state--input domain, but existing bounds are available only for full GP regression. Since exact GP inference scales poorly with the number of data points, its deployment is impractical in large-data regimes. We close this gap by developing a scalable GP framework that admits the derivation of uniform uncertainty bounds. We formalize a deterministic trigonometric feature Gaussian process (DTF-GP), a finite-dimensional kernel approximation based on discretized trigonometric features that reduces GP regression to Bayesian linear regression in feature space. We derive a high-probability uniform uncertainty bound for the proposed DTF-GP and provide its closed-form solution for the squared-exponential kernel case. Finally, we integrate the DTF-GP into a learning-based MPC scheme and demonstrate that it provides high-probability safety guarantees and exploration performance comparable to a full GP while improving computational efficiency in large-data regimes.
Direct Device-to-Satellite (D2S) communications promise global connectivity to unmodified user equipment (UE), extending coverage beyond terrestrial networks. Realizing this promise is fundamentally challenging: severe path loss and limited UE transmit power push uplink SNRs far below terrestrial norms, while suitable spectrum remains scarce. Together, these constraints impose a spectral-efficiency (SE) bottleneck, and under such conditions the efficiency of the UE power amplifier becomes critical, jointly governing transmit power and battery life. To improve UE-side power efficiency, 3GPP has adopted Discrete Fourier Transform-spread OFDM (DFT-s-OFDM) as an optional uplink waveform, exploiting its substantially lower Peak-to-Average Power Ratio (PAPR) relative to OFDM. To break the SE bottleneck, we show that aggressive non-orthogonal transmission, in which the number of concurrent users exceeds the number of receive antennas by more than 2x, can unlock substantial capacity gains that remain entirely unexploited. Realising these gains, however, requires receiver architectures that, to the best of our knowledge, have not yet been developed. DFT-s-OFDM intensifies the difficulty: the DFT spreading couples signal components across subcarriers, inflating the effective dimensionality of the detection problem. We address both challenges with a novel receiver design that jointly exploits the SE gains of aggressive non-orthogonal transmission and the power-efficiency benefits of DFT-s-OFDM. Simulations under realistic channel-estimation errors and high-mobility Doppler show that the proposed scheme achieves 2x the SE of baseline, surpasses recent nonlinear MIMO receivers by 40% at 15% of their complexity, and reduces PAPR by up to 6 dB relative to DFT-s-OFDM MIMO and 11 dB relative to OFDM.
Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.
We study an Estimated Time of Arrival (ETA)-based traffic-coordination framework for Urban Air Mobility corridors with merging at constrained waypoints (CWPs), where approved ETAs at CWPs serve as Required Times of Arrival (RTAs). Vehicle operators submit ETA plans at the merging point for approval by corridor-management authorities before corridor entry. Corridor entry is then scheduled by enforcing pairwise ETA gaps that maintain inter-vehicle separation on shared corridor sections. We develop two trajectory bounds to compute sufficient ETA gaps: a worst-case bound based on prescribed speed limits, and a stochastic bound based on probabilistic position envelopes under acceleration uncertainty. Using these bounds, we formulate sufficient ETA-gap computation and first-come, first-served corridor entrance scheduling. Simulations show that ETA coordination improves safety over an unscheduled baseline. The worst-case bound provides stronger robustness under higher disturbance levels, whereas the stochastic bound allows higher throughput under mild disturbances while relying on probabilistic modeling assumptions.
Binaural speech enhancement for hearing aids aims to reduce noise while preserving the interaural cues needed for spatial localization. Although deep neural network-based methods achieve strong noise reduction, they often distort the rela- tionship between the left and right signals. In this paper, we propose two novel binaural cue preservation losses. First, a binaural reconstruction error loss that directly penalizes masking-induced distortion in the relationship between the left and right spectra, providing a more direct measure of the binaural consistency than conventional separate interaural level differences (ILD) and interaural phase differences (IPD) errors as in prior work. Second, a binaural cue loss that jointly models ILD and IPD to better preserve the binaural structure. Experimental results show that both proposed losses maintain strong noise reduction performance and reduce masking- induced distortion compared to the state-of-the-art baseline cue loss, while the second proposed joint binaural cue loss also outperforms the baseline in ILD preservation.
This paper investigates the joint design of beamforming and radar receive filters in a multiuser bistatic integrated sensing and communications (ISAC) system, aiming to maximize the minimum radar signal-to-interference-plus-noise ratio (SINR) under communications SINR and transmit power constraints. We consider two scenarios: transmitted signals are either known or unknown at the radar receiver. We develop tractable solutions to the resulting non-convex optimization problems in both cases. For the known-signal case, we derive closed-form radar receive filters and iteratively design beamforming using fractional programming (FP) and successive convex approximation (SCA). For the unknown case, we adopt an alternating optimization (AO) approach to jointly design the beamforming and receive filters. Numerical results demonstrate that, while both approaches achieve comparable performance under per-slot optimization, knowledge of the transmitted symbols provides significant gains in multi-slot processing via coherent integration. Moreover, the proposed ISAC designs perform close to the radar-only benchmark under moderate communication requirements.
Differentiable nonlinear model predictive control (NMPC) provides a principled way to embed optimal control structure into end-to-end learning paradigms, but its practical use is often limited by the computational and memory costs of both forward optimization and backward sensitivity propagation. This brief proposes PANDA, a matrix-free solver for differentiable NMPC. In the forward pass, PANDA combines proximal-gradient iterations with quasi-Newton acceleration and introduces an adaptive stepsize enlargement mechanism to mitigate the conservativeness of monotone stepsize reduction. The resulting stepsize behavior and its effect on local convergence are theoretically analyzed. In the backward pass, PANDA performs implicit differentiation from the residual equation and computes adjoint sensitivities using Krylov-subspace iterative methods together with automatic-differentiation-based Matrix-Vector product operators, thereby avoiding explicit Hessian and Jacobian construction. The method is evaluated on a nonconvex trailer NMPC problem embedded in an imitation learning task. The results show that PANDA achieves much faster forward and backward computation and lower memory overhead than representative differentiable optimization solvers, while maintaining effective imitation learning performance.
This letter investigates reliability constrained hybrid beamforming for transceiver separated multistatic integrated sensing and communication in vehicular networks. A target position Cramer Rao bound minimization problem is formulated under outage probability, transmit-power, and analog constant modulus constraints. To handle the constrained non convex problem, we develop a proportional-integral Lagrangian proximal policy optimization algorithm. Simulation results show that the proposed algorithm keeps the average outage probability at or below the reliability threshold, around 8%-10%, improves constraint satisfaction, and achieves stable sensing performance.
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.