2026-09-28 | | Total: 90
Stacked intelligent metasurfaces process the transmitted field layer by layer, and fluid antennas make the position of every radiator a design variable. Combined, they pack radiators at sub-wavelength spacings within and across layers, where mutual coupling governs the physics that current models omit. This paper develops a coupling-consistent model of a fluid transmit layer illuminating a stack of passive fluid beyond-diagonal layers. The multiport impedance matrix of any port constellation is shown to be the impedance kernel sampled at the three-dimensional separations, in one closed spherical-Hankel form whose in-plane restriction is intra-layer coupling and whose axial restriction replaces the scalar Rayleigh--Sommerfeld propagator. Only the transmitter is driven: the layer currents are induced, and the familiar cascade is the single-pass limit of one matrix inversion. In the wavenumber domain the light circle still separates radiation from reaction, now with a plane-wave propagator attached, every Fourier mode sees a transmission line loaded by the layers, and an efficiency identity shows that what the stack costs in efficiency is set by its total induced-current norm, not by its layer count. Because the load is never assumed layer-diagonal, interconnections may join atoms on different sheets, and vertical networks repeated at every site are spectrally local. Alternating optimization designs precoders, positions, and loads with closed-form gradients. Numerically, fluid ports on a $λ/8$ grid perform within a few per cent of continuously movable ones, while a design based on the conventional cascade model can do worse than open-circuiting the stack.
Many voice interaction applications require exact control over both the content and delivery of responses, typically using pre-recorded lines. Recent full-duplex models respond with low latency but cannot guarantee exact content or reproduce a specific recorded performance, while cascaded systems can be constrained to predefined responses at the cost of additional latency. We propose RePlay, a spoken dialogue system adapted from PersonaPlex that handles multi-turn conversations by retrieving and playing pre-recorded lines. Using probing, we identify the layer and frame at which the upcoming response becomes recoverable, and use this hidden state as the retrieval query. RePlay retains only the layers up to that point and replaces text and speech generation with lightweight turn-taking and retrieval heads. In simulated multi-turn interviews, RePlay reaches a median latency of 383 ms, 3 to 7 times lower than ASR-LLM cascades of comparable dialogue quality, at the cost of lower exact-line accuracy. In a user study, participants preferred RePlay in 63% of ratings versus 12% for a fast cascade with a small LLM (p = 0.008), and showed a non-significant preference (46% vs. 21%) over a slower cascade with a stronger LLM.
This study evaluates a mathematical model of syllable production using an EMA dataset from a recently published study, which proved highly compatible with the model's architecture. The data set consists of regularly structured French phrases of equal duration and reduced phonetic and syllabic complexity. A dedicated procedure was used to transform EMA recordings into Maeda parameters. The Model-generated trajectories were then realigned with these transformed data using Canonical Time Warping (CTW). Statistical validation was conducted via a permutation test, comparing alignment quality against a surrogate condition with phonetically incongruent model outputs. The results show significantly better alignment for phonetically congruent pairings, revealing a strong structural correspondence between the model and the articulatory data. These findings provide experimental support for the representational validity of the model.
In daily life, people hear speech, footsteps, and music around them. We can often recognize these sounds and judge where they come from. Each sound source can be shown on a separate acoustic map, a rectangular image covering $360^{\circ}$ horizontally and $180^{\circ}$ vertically. The map shows the directions occupied by the source as a region and the sound energy within that region. A class label identifies the sound. Predicting these labeled acoustic maps from audio is called semantic acoustic imaging. Such maps could help robots perceive their surroundings and allow augmented reality displays to show sound regions and classes over the real world. Existing models can recognize sound classes and estimate a direction for each source. However, a direction alone does not describe the source region or its energy. Acoustic imaging must also distinguish sound sources in nearby directions, while the number of active sources and the regions they occupy can change over time. We therefore propose the Semantic Acoustic Imaging Detector (SAID), which predicts a separate labeled acoustic map for each active source from audio. First, we pretrain Audio2Sph, SAID's audio encoder, through sound energy estimation across directions without class labels. Then, we train the complete SAID model to predict source regions, energy, and classes together. We also develop a pipeline that generates simulated recordings for pretraining and supports fine-tuning on real recordings. On the official DCASE2026 Task 3 Track A evaluation set, our submitted system ranks first with 0.1080 macro-averaged mean average precision (Macro mAP) and 0.3962 Macro Pearson $r$. Demos and code are provided at https://github.com/IN03X/SAID.
Multimodal inference in safety-critical applications, such as autonomous driving and robot navigation, requires heterogeneous sensor observations to reach an edge server within a task-prescribed temporal window. Referenced to the event or state being inferred, this window ends at the latest time an observation remains useful for the current decision. Since sensor readiness times, data volumes, wireless channel conditions, and task relevance vary across sensors, maximising network throughput does not necessarily minimise the fused prediction error when the window closes. We cast multimodal uplink scheduling as sequential wireless evidence acquisition and, for a linear minimum mean-square error (LMMSE) fusion model, derive a conditional evidence gain metric from the reduction in residual-error volume. Defined through second-order statistics, this metric applies beyond Gaussian models and coincides with conditional mutual information when the target and prediction errors are jointly Gaussian. We then develop MIRA, a greedy task-driven maximum information-rate allocation policy that combines conditional evidence gain with each sensor's channel state information and remaining data volume, while updating sensor relevance as evidence is acquired. Experiments on synthetic classification, human activity recognition, and vehicle-trajectory regression show that MIRA outperforms both relevance-only and channel-only scheduling. Relative to the former, MIRA improves classification accuracy by up to 70% and reduces regression mean-square error (MSE) by up to 5.5%. Relative to the latter, it requires 58% less acquisition time to attain an 80% target accuracy, while achieving up to 95% higher accuracy and 9% lower regression MSE. These gains are achieved without maximising received-data volume.
This paper presents mathematically principled filters for multiple-model systems with states of different dimensionality, specifically the variable-dimension interacting multiple model (VD-IMM) filter and the variable-dimension generalised pseudo-Bayesian filter of order 2 (VD-GPB2). To do so, we first provide a Bayesian modelling of a variable dimensional dynamic system, and its measurements. Then, for variable dimensional linear Gaussian dynamic and measurement models, the VD-IMM filter is derived by assuming a posterior that has a Gaussian density for each mode, and then performing a Kullback-Leibler Divergence (KLD) minimisation after each prediction step to keep the Gaussian density form for each mode. Subsequently, the Bayesian update step is performed, keeping a Gaussian density for each mode. The VD-GPB2 filter considers the same model as the VD-IMM filter and also assumes that the posterior for each mode is Gaussian. In contrast, the VD-GPB2 filter propagates a Gaussian mixture for each mode in the prediction step. In the update step, the VD-GPB2 filter performs a KLD minimisation to have a Gaussian density for each mode. Simulation results show the benefits of the VD-IMM and VD-GPB2 filters compared to previous alternatives.
Congenital heart disease (CHD) is the most common birth defect, yet a large fraction of cases remain undetected on prenatal ultrasound, in part because current artificial-intelligence methods assume that the key diagnostic frames have already been isolated from a study, by a clinician or by a view classifier. We remove that assumption and address CHD screening directly at the level of the whole ultrasound study. We propose a two-stage framework that first learns transferable frame representations by self-supervised masked-autoencoder pre-training on unlabeled fetal ultrasound, then identifies cardiac frames with a disease-robust module and aggregates them with a transformer-based multiple instance learning (MIL) model that produces a case-level diagnosis from study-level labels alone. The model further returns its highest-scoring frames for clinician review, and a hierarchical head separates critical from non-critical CHD. On the internal test set of our multi-source development cohort (FUSE), the proposed cardiac-gated MIL model reaches an area under the curve (AUC) of 0.985 with a specificity of 0.990, outperforming the reproduced NATMED ensemble (AUC 0.861, specificity 0.600) and the FetalCLIP foundation model (AUC 0.867, specificity 0.710). On an independent external cohort, all models initially perform near chance, but label-free CORAL adaptation raises the proposed model from an AUC of 0.513 to 0.944, whereas whole-study and view-dependent baselines do not recover. These results indicate that whole-study MIL with disease-robust cardiac-frame identification is an accurate and deployable route to prenatal CHD screening.
High-frequency industrial cyber-physical systems (CPS) stream continuous sensor telemetry at sub-10-second intervals. Multi-horizon forecasting over hours is required for predictive maintenance and operational control. However, deploying machine learning models over extended horizons (H = 2,160 steps at dt = 5 s) encounters the Dual-Timescale Forecasting Dilemma: the trade-off between short-horizon kinetic momentum and long-horizon diurnal thermodynamic equilibrium. Autoregressive rollouts suffer from compounding error drift, while direct multi-output projectors exhibit variance explosion and noise extrapolation. In this work, we prove that autoregressive mutual information decays exponentially toward zero over extended lead times, and that the optimal minimum-variance estimator converges asymptotically to the periodic diurnal baseline (Lemma 1 and Theorem 1). Guided by these bounds, we propose the Dynamic Asymptotic Decomposition (DAD) framework. The architecture couples a continuous sigmoidal authority transition schedule \(α(h)\) with monotonic L2 regularization scaling \(λ(h)\), transferring prediction authority from kinetic autoregression to diurnal equilibrium while enforcing physical non-negativity. Evaluated across a 30-day industrial telemetry stream (N = 518,400 steps) over 16 walk-forward test windows, DAD achieves an overall Mean Absolute Error (MAE) of 0.1764, representing a 78.3% error reduction over Deep LSTM (0.8143), 60.2% over PatchTST (0.4431), and 20.8% over DLinear (0.2227). Computational profiling demonstrates edge execution in 0.72 ms with 0.203 MFLOPs and 72k parameters on a single CPU core (over 3400x faster than Transformer baselines), validated by an automated dual-stage CI/CD quality assurance harness and physical fail-safe architecture.
Speech-based screening is a promising, non-invasive approach for detecting Alzheimer's disease and related cognitive risks. However, models trained on a single domain often generalize poorly to unseen languages, tasks, or recording protocols. This paper investigates this deployment gap using a leave-one-corpus-out evaluation across four distinct datasets. Among 70 interpretable speech and language features, 59 exhibit direction conflicts between healthy control and cognitive risk groups across corpora, with pause, silence, and speech rate showing high protocol sensitivity. Furthermore, while the XLM-R text baseline achieves strong average performance, its Area Under the ROC Curve (AUC) drops to 0.520 on the weakest held-out domain. A standard GroupDRO baseline reaches a 0.766 mean speaker AUC and a 0.504 worst-domain AUC under the same protocol. To address this, we propose a fusion method that integrates XLM-R text baseline scores with evidence anchors selected during training. Balanced fusion achieves a 0.785 mean speaker AUC, while anchor-heavy fusion raises the worst-case speaker AUC to 0.615. This work highlights the need to audit feature transferability and report worst-case domain robustness in cognitive speech screening.
Digital twins are increasingly used in diabetes research, but reproducing individual glucose dynamics requires accurate identification of glucoregulatory model parameters. Traditional sensitivity analysis can identify influential parameters, yet a ranking based on limited conditions may miss parameters that matter during specific disturbances or for particular individuals. We therefore examine both the magnitude and timing of parameter influence across dynamic input-output conditions and assess whether a common ranking holds across participants. We analyze the Hovorka glucoregulatory model using data from 192 participants receiving automated insulin delivery therapy in the Type 1 Diabetes and Exercise Initiative dataset. We extend Sobol sensitivity analysis to time series and rank parameter influence under four conditions: full-day profiles, isolated meal disturbances, insulin bolus injections, and postprandial responses. We combine the condition-specific results into a global ranking and use it to select parameters for participant-specific identification. Compared with population parameters, identification restricted to the sensitivity-derived subset reduces the root mean square error of 60-minute glucose predictions by 60%, to approximately 31 mg/dL. These findings suggest that a global ranking can capture parameter influence across individuals and dynamic conditions. By narrowing the parameters requiring identification, this approach reduces computational cost and could accelerate the development of personalized diabetes digital twins.
Reliable high-dimensional variable selection requires scalable error-controlling methods. The Terminating-Random Experiments (T-Rex) selector estimates the false discovery rate (FDR) by aggregating early-terminated forward-selection paths in which predictors compete with synthetic dummies. We address two remaining challenges: i) predictor dependence can bias dummy-predictor competition; ii) computation is wasted on recomputing terms shared across experiments. Building on memory-efficient virtual dummies, which sequentially sample projections from their exact conditional law, we estimate the conditional Gaussian law of inactive predictors and map draws onto the remaining sphere radius. A shared lazy Gram cache computes the response product and each requested Gram column once across experiments. Simulations across three covariance structures show that uniform spherical dummies can exceed the target FDR, whereas the proposed method empirically controls FDR. Caching yields more than a sixfold speedup, and the method remains feasible with 100000 predictors, where competing FDR-controlling methods become computationally impractical.
This article develops a physics-informed deep learning framework for single-snapshot 3D phase-only positioning in narrowband phase-coherent distributed multiple-input multiple-output (D-MIMO) networks under two-ray propagation. Unlike prior phase-coherent D-MIMO localization methods that largely assume line-of-sight (LoS)-only channels, we consider a LoS path and a specular ground reflection. Since the narrowband single-snapshot observation cannot resolve these components in delay, their coherent superposition induces structured perturbations in the carrier phase measurements. To address this, we develop a Gaussian process (GP)-based model that learns the quasi-periodic phase distortion caused by ground reflection from carrier phase measurements collected at a small number of training locations and uses it to generate high-quality synthetic samples. Building on these samples, we propose the Phase-Only Positioning Transformer (POPT), an encoder-only transformer that captures inter-antenna point (AP) phase relationships without solving the highly non-convex maximum-likelihood (ML) problem underlying model-based estimators. We also derive the fundamental position error bound (PEB) and develop maximum-likelihood estimation (MLE) for the considered multipath phase-only D-MIMO positioning problem. Numerical results show that, with only 50 GP training locations, the proposed method achieves near-PEB accuracy. The analytical MLE is highly sensitive to relative permittivity mismatch, whereas the proposed method mitigates this dependence by learning the phase perturbation directly from measurement data. Compared with MLE, the proposed approach also reduces floating-point operation (FLOP) complexity and inference time by about 1.7 and 3.6 orders of magnitude, respectively.
Large-scale Digital Surface Models (DSMs) can be produced cost-effectively from satellite images via stereo-photogrammetry. However, the resulting 3D maps are often contaminated by noise, outliers, and voids. On the other hand, aerial LiDAR provides high-accuracy elevation measurements at a substantially higher cost. In this work, we study diffusion models conditioned both on photogrammetric DSMs and Pléiades imagery to refine vertically co-registered DSMs. We introduce a modified Stable Diffusion 3 architecture with a pruned text stream and a patch-wise normalization strategy, enabling stable training on LiDAR data and transfer from natural images to elevation maps. Experiments in French cities demonstrate that multimodal conditioning improves elevation accuracy, reducing Dense Urban RMSE from 6.00 to 3.45 m in the in-context cities and from 4.16 to 2.77 m in the held-out city of Bordeaux.
Automated EEG artifact removal may improve downstream analysis but can also alter predictive information. We benchmarked nine automated artifact removal methods against a common no-artifact-removal baseline for cross-dataset EEG age prediction. We introduce Signal Quality Index (SQI)-guided GEDAI, which leverages local signal-quality assessment to restrict correction to the channel--epoch pairs requiring intervention. Three deep neural architectures were trained on TUEG and evaluated without target-domain fitting on ds005385, LEMON, and TDBRAIN. Across this setting, GEDAI and SQI-guided GEDAI were the only methods with consistent gains over baseline in age prediction performance across all datasets and architectures ($Δ$MAE $=-0.77/-0.64$ years, $ΔR^2=+0.083/+0.072$, respectively). The remaining methods were neutral or detrimental on average ($Δ$MAE $=+0.22\pm0.16$ years, $ΔR^2=-0.023\pm0.016$ across methods). The two GEDAI-based methods achieved closely matched performance, while SQI guidance reduced the median modification ratio from $74.78\%$ to $42.80\%$. These findings show that curation benefits are method-dependent and establish SQI guidance as a more selective operating point, leaving more of the original EEG unchanged and limiting the potential loss of neural activity while retaining most of GEDAI's predictive benefit.
Neural receivers outperform conventional 5G NR processing chains, but their compute and memory demands hinder real-time deployment. For a standard-compliant multi-user MIMO neural receiver, the 4-bit number format, not merely the bit width, determines whether compression preserves the gain over classical receivers. We apply weight and activation quantization-aware training (QAT) and, separately, 50% magnitude pruning, comparing INT8/INT4 and FP8 (E4M3)/FP4 (E2M1) weights with INT8 post-ReLU activations. Trained on 3GPP UMi channels and evaluated on TDL-B and TDL-C, 8-bit weight-activation models remain within 0.05 dB of FP32 at 10% and 1% block error rate (BLER). At 4 bits, uniform INT4 loses 3.3-3.7 dB and falls below LS-LMMSE, whereas FP4 more than halves this loss (1.3-1.4 dB) and still outperforms it by about 0.5 dB, even after pruning. FP4's denser near-zero grid matches the trained weight distribution, and FP4 avoids the residual-path over-pruning seen with INT4. An analytic cost model projects 66x fewer bit-operations and 8.8x less weight storage for pruned 4-bit-weight inference.
This paper presents a data-driven framework for joint observer design and sparse sensor scheduling for unknown nonlinear networked systems. The nonlinear dynamics are approximated online as a piecewise sequence of locally linearized discrete-time systems, resulting in a switched linear representation recursively identified through Subspace State-Space System Identification (4SID). To ensure consistency across regime transitions, an Orthogonal Procrustes alignment is introduced to promote coordinate consistency across consecutive regime transitions and mitigate artificial discontinuities caused by arbitrary state-space coordinate changes. Based on the identified local realizations, a dual-rate predictor-corrector observer is designed. The observer gain is computed through a convex optimization problem that jointly addresses estimation accuracy, sensor sparsity, and stability requirements. In particular, an L_{2,1}-norm regularization term promotes column sparsity in the gain matrix, enabling the automatic selection of informative measurement channels, while a spectral-norm constraint guarantees Schur stability of the estimation error dynamics. The proposed framework is first validated on a traffic network simulated in Aimsun Next. Results obtained on an 18-link network show that the observer accurately reconstructs macroscopic traffic states while significantly reducing the number of active sensors required for real-time estimation.
It is necessary to characterize chemical reaction networks (CRNs) accurately to facilitate engineering of both synthetic and naturally occurring biological systems. Estimation of multiple unknown parameters and states in CRNs requires first, an identifiability analysis, and then adoption of a suitable estimation technique. In this paper, we took an example of a reduced order gene expression system, performed parameter sensitivity analysis for multiple unknown parameters, and used a modified Rao-Blackwellised particle filter (RBPF) to estimate parameters and states. The proposed framework estimates parameters using a particle filter and states using an extended Kalman filter (EKF) with process noise covariance updated recursively based on chemical Langevin equation (CLE). We compared the accuracy of parameter estimation, error in state estimation, and whiteness of the innovation sequence for the proposed filter with fixed choices of noise covariance. We found that the RBPF with updated noise covariance finds a balance between these three criteria, demonstrating its suitability for joint state and parameter estimation for stochastic CRNs.
Renewable energy power plants are increasingly expected to provide ancillary services to the grid. Yet, the variability of sources such as solar photovoltaics requires additional operational flexibility to deliver these services. In this context, hybrid energy storage systems (HESSs) offer an appropriate solution, because different technologies with complementary characteristics can be applied and leveraged to share the power demand and improve the overall performance. Yet, a proper power sharing requires a detailed model of the storage elements. This aspect has barely been studied in the literature, where most models consider constant efficiencies. Moreover, most formulations are based on optimisation problems that are computationally too demanding to be solved in real time. In this paper, a controller is proposed to minimise the losses of a HESS considering detailed power-dependent efficiency curves of each storage technology. The power sharing method is based on the analytical verification of the Karush-Kuhn-Tucker (KKT) conditions, which makes it computationally efficient and suitable for real-world deployment. The proposed controller is applied to a test case consisting of a 10 MW PV power plant with a HESS based on a lithium-ion battery and a redox-flow battery, each with 2.5 MW power and 5 MWh capacity. The main contributions are verified via numerical simulations performed in MATLAB/Simulink, whereas a real-time implementation deployed in OPAL-RT demonstrates the viability of the algorithm for real-time applications.
Quantum Machine Learning is a novel field of research aimed at devising machine learning approaches exploiting principles of quantum mechanics, such as superposition, entanglement and interference. In this context, we present a scalable hybrid Quantum Diffusion Model, and evaluate its use for medical image analysis. Specifically, our method is based on a Discrete-Time Quantum Walk algorithm, executed on a real quantum device, to model the forward dynamics of the diffusion model. For the backward step of the diffusion model, we devise and evaluate a classical learning model, which is used to reversely denoise the data. In contrast with other existing attempts at applying quantum machine learning for image analysis tasks, severely limited by the size of existing quantum devices, our method allows to process real-world large size medical data. In particular, we present results on grayscale and RGB images, as well as 3D volumes of moderate sizes. We benchmark our results by reproducing an alternative classical counterpart model, based on diffusion models on discrete state spaces. By doing so, we compare the generation capabilities of both models in terms of three distinct state-of-the-art metrics in the field of image generation, showing the competitive, promising results of our approach.
Conventional tractometry averages white matter microstructure into a single profile, masking intra-bundle spatial heterogeneity and diluting localized alterations. We propose an automatic sub-bundle tractometry framework using Fréchet-based hierarchical clustering to decompose bundles into geometrically coherent sub-bundles, resolving complex fanning configurations without empirical thresholds. Multi-compartment metrics (fractional anisotropy, FA; isotropic free-water fraction, IFW) are projected onto these clusters. We evaluated this framework in 140 participants across early-life (ELD) and late-life (LLD) depression cohorts, alongside age-matched healthy controls. We assessed performance based on microstructural homogeneity, sensitivity to age, and clinical group differences. Proposed sub-bundle metrics were significantly more homogeneous than classical full bundle profiles (p < 0.001). Our approach showed enhanced sensitivity, detecting more age-correlated bundles in elderly controls (11 vs. 5) and more clinical group differences in LLD (FA: 6 vs. 4; IFW: 23 vs. 17) with larger effect sizes. In ELD, it improved the detection of subtle FA differences between remitted and resistant individuals (2.9% vs. 1.8%). Gains were particularly pronounced in LLD, reflecting the framework's ability to resolve fanning tracts susceptible to brain aging. By preserving spatial heterogeneity, this automatic decomposition provides a robust foundation for large-scale clinical tractometry.
We propose a self-supervised approach for learning room impulse response (RIR) representations from single-channel noisy-reverberant speech. It consists of first training on reverberant data, then on noisy-reverberant data, and finally with a teacher-student approach, where the student learns to replicate the teacher's embeddings when given a noisy version of the reverberant input. We assess their representational capabilities by estimating acoustic room parameters from them. Conditioning a discriminative speech enhancement model on the derived embeddings yields consistent gains across all evaluated metrics, including downstream word error rate, for both reverberant and noisy-reverberant speech.
Virtual power plants (VPPs) are an effective solution for increasing the penetration of renewable energy sources (RES) in power systems. Moreover, VPPs can provide ancillary services if generators and loads are closely connected and properly coordinated. This aspect has been addressed in several studies, yet they mostly focus on the dispatch optimisation level, leaving real-time operation aspects unresolved. To close this gap, a central controller for VPPs based on a rolling-horizon optimisation is proposed in this work. Its main objective is to maximise revenues while delivering frequency restoration services. A day-ahead optimisation firstly defines the power setpoint and the upward and downward reserves of the VPP. Then, the rolling-horizon optimiser, executed every minute, distributes the power and reserve references between the VPP units to achieve their close tracking while taking into account resource availability (wind speed, irradiance, etc.), operational constraints and battery degradation. The effectiveness of the controller is tested using a two-area interconnected power system under different RES generation and demand profiles. The simulations are performed in MATLAB/Simulink and include detailed dynamic models of the VPP elements, demonstrating the applicability of the algorithm to real systems. The day-ahead and the rolling-horizon optimisation problems are modelled using YALMIP and solved using Gurobi. The obtained results demonstrate the VPP supporting grid frequency restoration by optimally allocating its resources, even under the presence of forecast uncertainty.
Efficient acquisition of correlated continuous-time signals is critical in applications such as distributed sensor networks and array processing. However, existing correlation models often fail to capture the shared structure of physical signals measured in close proximity. In this paper, we model each signal in an ensemble as the sum of a hidden common lowpass component and a signal-specific innovation highpass component with disjoint spectral supports, where the common bandwidth is unknown. We first derive theoretical identifiability conditions guaranteeing unique decomposition and reconstruction from subsampled observations. Further, we propose a practical recovery algorithm and a joint reconstruction framework that leverages parametric structured dictionaries parameterized by the unknown common bandwidth. Numerical experiments on an ensemble of four signals demonstrate exact reconstruction with a 25% aggregate sampling rate reduction under identifiability conditions, and robust recovery with up to a 74% rate reduction when all channels are sampled below the Nyquist rate, significantly reducing data acquisition and hardware overhead.
Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the interaction between the training objective and the representation's information capacity: supervised representations only disentangle speaker identity when their capacity is sufficiently constrained.
Meta-adaptive filtering (Meta-AF) provides a data-driven alternative to hand-crafted adaptive filter updates by employing a learned optimizer throughout online adaptation. However, when Meta-AF is used for active noise control (ANC), time-varying acoustic paths remain a major challenge. In particular, the physics-informed optimizer features are constructed using a secondary path estimate and become mismatched when the physical path changes, leading to inaccurate filter updates and degraded noise reduction. To address this problem, this paper proposes a Coupled Meta-Adaptive Filtering Active Noise Control (CoMeta-AF-ANC) method, which applies meta-learning to jointly learn the control filter adaptation and acoustic path tracking within a closed-loop framework. The proposed Meta-Gated Joint Path Identifier (MG-JPI) simultaneously tracks the primary and secondary paths from the available ANC signals without auxiliary noise, while the updated secondary path estimate is fed back to reconstruct the Meta-AF controller features. A delayless dual-rate realization performs learned adaptation at the frame rate while generating the control signal at the sampling rate in the time domain. Evaluation using measured headrest acoustic paths shows that CoMeta-AF-ANC outperforms representative ANC algorithms in tracking time-varying acoustic paths, while maintaining higher stability across unseen head movement scenarios. It also generalizes well to real-world noises not encountered during training.