2026-09-22 | | Total: 20
Standard valuation methods, including discounted cash flow, the income approach standard IDW S 1 of the Institute of Public Auditors in Germany, and market multiples, compress milestone probabilities, continuation options, and risk shifts into opaque aggregate parameters; none provides a structured protocol for decomposing AI integration into auditable option-level assumptions. We propose an industry-agnostic taxonomy separating AI Integrators from AI Providers. AI Integrators are further classified by their Integration Depth Level, ranging from no integration to AI at the core of the product or process. A milestone-gated real-options overlay decomposes milestone state value into five components, and an Analytic Hierarchy Process-based Success Readiness Index derives per-option probabilities from structured pairwise comparisons for scenario analysis. Applied to an AI-native energy software-as-a-service firm, the framework yields a coherent valuation band traceable to identifiable option-level assumptions. Risk concentrates in later-stage continuation options, matching the structural prediction for AI Providers. The protocol applies across the firm lifecycle, including mergers and acquisitions due diligence. The case is a single-firm demonstration of protocol coherence, not empirical validation; multi-case testing against realised post-exit valuations is left to future research.
The sooner we receive information, and the more accurate it is, the better planning decisions we can make. Every day, prediction markets let anyone bet on tomorrow's high temperature in cities around the world, creating a market-implied forecast built on dispersed information. We use the past five years of market data from the Kalshi exchange for seven American cities to extract, hour by hour, the market-implied forecast. We use this forecast as a measuring instrument to see how much information about the temperature the market makes public before the public forecasting system does. We race it against the leading American and European weather forecasts. In six of the seven cities we study, the market beats the most accurate single public forecast, the National Blend of Models (NBM). Aggregating every city-day, at the end of the market's first hour of trading it beats the best single public product by about 10 percent in root-mean-square error, and holds its lead through the day, overnight, and into the target day. Looking at how the forecasts move over time, we find the National Blend travels four times further toward the market between its postings than the market travels toward the NBM. The market does not react to new weather forecast updates; instead, the forecast slowly publishes information that the market had already shared publicly.
We study optimal stopping under dynamic risk measures with simultaneous ambiguity in the probability model and the discount rate. We introduce a paired ambiguity framework combining Girsanov model uncertainty with cash subadditive risk evaluation and characterize the stopping value by an upper reflected backward stochastic differential equation (BSDE). We establish structural properties of the resulting stopping operator and study quadratic drivers associated with entropic risk measures, obtaining explicit stopping rules in several benchmark cases. We then develop a deep learning scheme for the reflected quadratic BSDE. The convergence analysis uses discrete reflection and truncation to reduce the quadratic problem to a globally Lipschitz system and combines reflected BSDE discretization estimates with neural network approximation errors. Numerical experiments for American options illustrate the effects of discount rate and entropic ambiguity on stopping values and exercise decisions.
Organizations invest heavily in internal mechanisms such as incentive systems, governance structures, performance measurement, analytics, and repeated reorganizations. Yet realized performance gains are often weak or short-lived, particularly in large and mature organizations. Standard explanations emphasize allocative inefficiency arising from incentive misalignment, information problems, or bounded rationality. While relevant, these explanations struggle to account for environments in which effort is high, optimization capabilities are sophisticated, and aggregate performance remains far below apparent potential. This paper develops a theory of structural inefficiency inside organizations. The central mechanism is the presence of internal externalities: unpriced cross-effects of organizational mechanisms acting simultaneously on multiple performance objectives. When such effects are mixed in sign, organizational effort is partially canceled internally, creating a structural ceiling on realized performance. The paper introduces a simple representation of organizational primitives and key performance indicators and derives a realized-performance law in which internal conflict enters as a multiplicative discount factor. Under weak and realistic conditions, structural inefficiency dominates allocative inefficiency, explaining why optimization often fails before organizational coherence is restored.
Intraday electricity markets enable participants to adjust energy positions close to delivery, with price forecasts necessary to support trading and the scheduling of flexible electricity resources as well as increasingly responsive consumers. Forecasting model specifications have progressed from using macro-features, such as renewable generation and load, to the micro-features of continuous orderbooks. A recent advanced deep learning model, OrderFusion, explicitly models micro-level buy-sell orderbook interactions. However, despite its superior comparative forecasting performance, it considers only one delivery product at a time, ignoring neighboring-product information. Moreover, when forecasting aggregated price indices such as the liquid German ID3, ID2, and ID1 products, it omits the information contained in price trajectories. In contrast, pretrained time-series foundation models have shown success in financial, renewable-energy, and day-ahead electricity price forecasting. However, their performance on intraday orderbook data remains an open research question of considerable practical importance. In this paper, we propose OrderFusion+, an open-source deep learning model that combines historical orders from the target and neighboring delivery products to forecast probabilistic buy-sell price trajectories. We benchmark OrderFusion+ against forecasting baselines and pretrained foundation models, and investigate dynamic market conditions through the designed dynamic masking mechanism, revealing insights into market efficiency. The implementation and forecasts can be found at: https://runyao-yu.com/OrderFusion/
We characterize the worst- and best-case values of a mean-variance functional over a 2-Wasserstein ball. Using quantile representations and the geometry of attainable means and standard deviations, we reduce both infinite-dimensional problems to scalar equations and construct the extremal laws as location-scale transformations of the reference distribution. Their values depend on the reference law only through its first two moments. We then derive dual representations and formulate proportional risk sharing under heterogeneous beliefs as a finite-dimensional optimization problem. Under homogeneous beliefs, we show that the classical proportional allocation remains optimal for every ambiguity radius and $α$-maxmin weight. Finally, we study a coupled distortion-variance functional and characterize its worst-case quantile through a convex-envelope construction, allowing the extremal law to change in shape as well as location and scale.
Day-ahead electricity price forecasts support trading and storage decisions, but for battery arbitrage predicting intraday price spreads is more relevant than predicting individual hourly prices. Here we show that a temporal hierarchy forecasting (THieF) framework that jointly reconciles forecasts of hourly electricity prices and all intraday price spreads consistently improves performance across two major European electricity markets and three different forecasting architectures. Using five years of out-of-sample data from Germany and Spain, we obtain accuracy improvements of up to 19.7% and profit gains of up to 10.4% relative to unreconciled hourly price forecasts. The gains persist even for a highly accurate pretrained TabPFN foundation model. Our results demonstrate that exploiting coherent relationships between economically relevant forecasting targets can improve both predictive accuracy and decision value, and that better statistical forecasts do not necessarily imply better economic decisions.
Modeling the dynamics of option implied volatility surface (IVS) is crucial for pricing, hedging, and risk-managing option portfolios. We develop a universal conditional diffusion model that learns to jointly generate next-day IVS increments and the underlying stock's returns. The model is trained on pooled data from 50 stocks and evaluated on a test set comprising 50 in-sample stocks and 50 out-of-sample stocks excluded from training. The training objective is primarily the minimization of MSE, with the variants that jointly impose surface smoothness, and that penalize the presence of static-arbitrage, which are all economically meaningful constraints. Across in-sample and out-of-sample stocks, our diffusion model outperforms the VolGAN benchmark in reducing arbitrage violations, improving stock risk prediction, and aligning explained variance ratios by the first three principal components. The fact that the model extrapolates well to stocks that are excluded from training indicates that the learned dynamics can be shared across stocks. These findings support a universal conditional diffusion as a credible framework for multi-stock implied volatility surface scenario generation.
We study the design of priority pricing systems with heterogeneous agents in environments in which improving quality for some agents reduces the average quality that can be provided. Contrary to the equity-efficiency tradeoff emphasized in public debates, we show that under economically natural conditions priority pricing can Pareto-improve on an equal-allocation benchmark. Three priority tiers suffice for such an improvement, combining higher quality for a fee, lower quality with compensation, and an intermediate tier at the benchmark quality; two tiers are never enough. Our results provide a framework for overcoming equity-efficiency tensions in applications such as lane pricing, waiting-line design, public provision, and insurance.
People often face environments where multiple models compete to explain the same observations. This paper examines how people update beliefs in such settings and how preferences over payoff-relevant states shape model selection and belief updating. This paper first develops a framework where preference-driven bias distorts the perceived model, affecting Bayesian and best-fit updating differently. In a laboratory experiment, most participants are classified as Bayesian updaters, who average across models, while a substantial minority are classified as best-fit updaters, who select the model that best fits the observed signal. Within-participant comparisons between the symmetric payoff and asymmetric payoff conditions indicate that asymmetric payoffs shift reported beliefs toward the preferred state, particularly among participants classified as best-fit updaters. Relative to symmetric payoffs, asymmetric payoffs increase the reported belief of the preferred state by about 8 percentage points among best-fit updaters, while the estimated effect among Bayesian updaters is close to zero. These findings help us better understand model-based learning and have implications for domains such as political polarization and financial investment, where competing narratives and strong preferences often coexist.
This work presents an operator-based visual analytics pipeline for exploring synthetic systemic risk dynamics in financial networks. The framework is formulated as a composition of mathematical operators that sequentially transform synthetic financial observations into dynamic scientific visualizations. The pipeline consists of six operators: latent risk mapping, probabilistic score generation, financial network construction, distance-based contagion dynamics, visual encoding, and perspective projection. Together, these operators provide a modular computational structure linking nonlinear risk surfaces, time-dependent probabilistic states, network topology, and shock propagation. A reproducible implementation demonstrates the architecture through controlled synthetic experiments. Nonlinear latent risk representations are converted into probabilistic scores, embedded in a weighted financial network, and propagated via shortest-path contagion. The resulting states are then mapped into dynamic visual representations. The experiments illustrate how visualization can be treated as an explicit stage of the analytical process rather than a post-processing step. The proposed formulation does not aim to introduce new predictive models or contagion mechanisms. Instead, it offers a transparent and modular pipeline that connects generative risk modeling, network dynamics, and scientific visualization within a unified operator-based architecture. This approach supports reproducibility and facilitates the exploration and communication of complex systemic-risk processes in synthetic settings.
We study affine stochastic Volterra equations on the cone of symmetric positive semidefinite matrices. For scalar kernels acting entrywise on the matrix dynamics, we establish weak existence by exploiting stochastic invariance results for Volterra equations on convex domains and derive a conditional Fourier--Laplace transform formula characterized by matrix-valued Riccati--Volterra equations. As an application, we extend the Gibson--Schwartz commodity model by replacing its variance-covariance structure with a Volterra--Wishart process. The resulting model allows for memory in the variances and for stochastic instantaneous correlation, while retaining affine tractability. Its joint Fourier--Laplace transform admits an exponential-affine representation governed by a matrix Riccati--Volterra equation. While the existence theory considered here excludes kernels that are singular at the origin, shifted fractional kernels remain admissible and provide a tractable specification with power-law memory.
Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrated deployment framework for evaluating whether those signals remain useful in financial decision systems. Computer science research has developed strong methods for time-series forecasting, text classification, multimodal stock prediction, graph-based market modeling, and machine-learning operations, yet these streams do not provide a domain-specific protocol that jointly tests financial language-model outputs under event-time observability, probability calibration, execution timing, transaction costs, liquidity constraints, capacity limits, operational diagnostics, and statistical inference. We introduce MFAST, a Market-Friction-Aware Sentiment-to-Trading framework that converts timestamped financial text into auditable, reproducible, and market-feasible trading decisions. The application is news-based trading, where firm-specific text must be linked to securities before portfolio decisions can be evaluated. The framework links Refinitiv News Analytics to Center for Research in Security Prices (CRSP) equity data, restricts the primary out-of-sample evaluation to post-release news outside disclosed foundation-model data-freshness periods, and adds a public replication arm using open financial text and public price data. Results show that decoder-only language models outperform encoder baselines and dictionary sentiment in classification, calibration, return prediction, and net portfolio performance, while operational diagnostics reveal trade-offs among accuracy, latency, memory, throughput, and inference cost. The paper shows that credible evaluation of financial language models requires an end-to-end engineering approach combining language understanding, temporal discipline, market-friction-aware deployment, and reproducible validation.
We introduce leaky-integrator reconstruction, a training-free method that cures the error accumulation of recursive differenced forecasting. Our first contribution is diagnostic: predicting one-step changes and integrating them by cumulative summation, the standard remedy for non-stationarity, is a discrete integrator with a pole on the unit circle, and we show this makes recursive rollout of a nonlinear model diverge, its 336-step error reaching several times that of a well-behaved forecaster (normalised MAE 1.6-3.8 versus about 0.8) across every neural architecture tested. Our second, central contribution is the fix: move the pole inside the unit circle with a leaky integrator H(z) = 1/(1 - gamma z^-1), gamma < 1, which provably bounds the accumulated error variance. Applied at reconstruction time with a single fixed gamma=0.9 (no retraining, a two-line change to any deployed one-step or foundation-model forecaster), it shrinks error at every horizon, the mean gain over seven diverging architectures and twenty datasets growing from ~3% at H=24 to 23% at H=96, 37% at H=192 and 51% (43-74% across those architectures) at H=336 (78% with an oracle pole). Crucially, it is provably inert where no pathology exists (stable or joint predictors already at the irreducible rate), making it a safe, general default.
Mitigating \emph{drawdown}, the decline in wealth from its running peak, presents a canonical problem in path-dependent risk control. In this paper, we develop a finite-horizon control framework that enforces a prescribed maximum percentage drawdown limit in multi-asset stochastic systems. Our first result is an exact robust-invariance theorem characterizing every control action that preserves a prescribed drawdown limit against all supported returns. We show that every robustly safe control admits a \emph{drawdown-modulated} form: the product of the current drawdown \emph{cushion} and a feasible \emph{normalized direction}. This yields a complete parameterization of robustly drawdown-safe policies. Additionally, under stagewise-independent returns, we show that optimizing over all robustly safe causal policies reduces to a one-dimensional Bellman recursion and yields an optimal robustly safe state-feedback policy. Finally, we characterize the linear time-invariant (LTI) gains satisfying a prescribed drawdown limit and prove that optimal drawdown modulation achieves no lower expected return under the same limit. Strict expected-return improvement holds for horizons of at least two stages whenever the LTI policy has positive expected one-stage net return.
Pairs trading exploits mean reversion in the relationship between related assets. We adapt this idea to political betting markets by modelling the combined implied probability of the two major-party nominees with a latent Ornstein-Uhlenbeck process whose mean-reversion level varies over time and whose observations contain additive noise. Model parameters are estimated from regularly sampled odds data using a state-space likelihood, with consecutive repeated values represented by a single retained observation and the elapsed number of sampling intervals preserved in the continuous-time transition. Parametric-bootstrap upper prediction bounds identify signal times at which the combined implied probability is likely to decline, and a no-intercept Bradley-Terry-type model selects the candidate-specific odds quote. The candidate-selection model is trained on 2020 U.S. presidential-election data and evaluated out of sample on 2024 data. The 2024 analysis produced 130 signals, empirical one-step coverage of 95.1%, a mean synthetic odds-price return of 1.86%, and an unannualized per-trade Sharpe-type ratio of 1.12. These returns are frictionless descriptive quantities rather than executable betting-exchange profits. The results support the integrated framework as a proof of concept for two-candidate electoral markets.
We examine log-optimal portfolio allocation when the long-run price of an asset follows a power-law trajectory, $P(t)=At^α$, and its instantaneous return variance decays as $σ^2(t)=σ_0^2 t^{-2γ}$. Under a zero risk-free-rate benchmark, the continuous-time Kelly fraction scales as $K^\ast(t)=(α/σ_0^2)t^{2γ-1}$. Exact temporal invariance therefore occurs when $γ=1/2$, whereas deviations from this value produce systematic age dependence in the allocation. We propose a scaling hypothesis connecting growth in network participation, effective market liquidity, and declining volatility. Under a specified set of scaling assumptions, this model predicts the benchmark exponent $γ=1/2$. Using historical daily Bitcoin prices, we estimate the power-law price exponent and examine the sensitivity of the volatility exponent to the length of the rolling window. For windows of four to nine years, the estimated volatility exponents have an arithmetic mean of 0.53 and a cross-window standard deviation of approximately 0.03. Because these estimates are obtained from overlapping observations and the same underlying price history, this spread is interpreted as a measure of model sensitivity rather than a formal confidence interval. Finally, we show that a time-dependent multiplicative contribution to return variance generally breaks exact Kelly invariance. We illustrate this result using a scenario in which Bitcoin transaction-fee variability affects the effective variance process. The results identify the conditions under which log-optimal allocation can remain stable under non-stationary power-law asset dynamics and clarify the assumptions required when applying this result to Bitcoin.
AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.
Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older applicants. This negative shift occurs in eight of the ten models that we evaluate. Post-trained models have much more correlated decisions than base models which is likely driven by human capital traits like skills or college major. However, greater consensus among models increases global systemic exclusion rates from 5.6% to 17.3% and exacerbates demographic inequalities, with intersectional systemic exclusion rates ranging from 12.2% to 21.7% for post-trained models. We find that this inequality is primarily driven by age-based discrimination that is exacerbated in post-training. These results indicate that while post-training techniques may improve models' abilities to select the best applicants, they may raise systemic inequality risks for those at the margin by uniformly introducing new biases.
People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.