2026-09-28 | | Total: 12
Many models, such as fixed-effect models for panel or network data, are hard to estimate because they feature nuisance parameters that are both numerous and estimated imprecisely. This, in general, causes an incidental-parameter problem in the estimator of the parameters of interest. The problem can be alleviated by working with an estimating equation whose expectation is insensitive to the value of the nuisance parameters. We discuss and contrast three notions of insensitivity, also called orthogonality, in the context of likelihood models: Neyman orthogonality, Neyman orthogonality to order q, and full orthogonality. Orthogonal moments are obtained by projecting the estimating equation on nested subspaces, which are spanned by, respectively, the scores of the nuisance parameters, the first q derivatives of the likelihood ratio with respect to the nuisance parameters, and all likelihood ratios of the model. We give explicit constructions in binary-choice, count-data, and nonlinear regression models.
We consider a robust delegation problem in which the principal does not know the distribution from which the underlying state is drawn. The principal can choose a general randomized mechanism and maximizes her worst-case expected payoff over all state distributions. Our main result characterizes the robustly optimal mechanism. The mechanism has up to three regions: (i) accommodation, where the agent's ideal action is taken, (ii) calibrated randomization, where each type receives a distinct lottery, and (iii) pooling, where all types receive the same lottery. We then extend our model to a multidimensional setting in which preferences are additively separable and satisfy a cross-dimensional symmetry condition, and show that delegating separately across dimensions is robustly optimal.
Numerical inflation targets anchor beliefs. Across euro-area and US professional forecasts, inflation swaps and options, and realized inflation, uncertainty about inflation is compressed at the announced number and kinks exactly there. This paper identifies a cost of the same design that, to our knowledge, has not been shown before, and that appears in second moments only. Within the workhorse New Keynesian model, tolerating part of the inflation a supply shock produces is optimal, yet the optimal tolerated share is not identified: optimal look-through and an unwarranted drift of the effective target are observationally equivalent in the inflation history, so a central bank cannot demonstrate that a warranted deviation is warranted, ex post as much as in real time. Agents holding finite, heterogeneous patience then generate predictive variance that is flat below the target and rises linearly with the expected overshoot above it -- an observational-equivalence bill, zero at the announced number and accumulating with the point-years inflation spends above it. The first-order benefit of the number stands; what the cost changes is how inflation uncertainty must be measured. The distance from target explains 70% of the variation in professional forecast variance, and the bill lies within that component; Normalized Uncertainty -- to our knowledge the first such correction -- removes it. The purge changes inference substantially: on French loan-level data, raw dispersion is unrelated to corporate loan rates, while one standard deviation of the purged measure is associated with rates 89 basis points higher, and cross-country growth and time-series results shift similarly. Conventional measures of inflation uncertainty partly record inflation's distance from its anchor rather than uncertainty about the outlook.
LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \$247 to \$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.
We study recursive-design wild bootstrap inference for dynamic panel data models with unobserved common factors estimated by Common Correlated Effects. In the large N,T setting, the bootstrap reproduces the biased limiting distribution in pure autoregressive models, but fails to capture all bias and factor-estimation variance components in models with additional regressors, particularly under weak exogeneity. We trace this failure to holding regressors fixed across bootstrap replications. We propose to combine bootstrap procedure with available bias-correction methods to conduct adjusted inference. Monte Carlo evidence shows substantial improvements over conventional strategies of using bias-correction paired with cross-sectional bootstrap methods.
The return migration of skilled workers provides origin countries many important benefits. The likelihood of return migration is thought to depend on relative economic opportunities, but quantitative evidence of its relative importance remains limited. Here we measure return migration rates using the public profiles of software developers on GitHub, geocoded across ten waves between 2012 and 2026, in a panel containing more than 270,000 international movers. Within five years of departure, 8.8% of movers have returned. Origin-country rates range from under 1\% to 16\%. Hazard models with corridor fixed effects show that migrants whose home countries experience greater catch-up growth return at significantly higher rates. Immutable features of country pairs such as cultural or geographic distance do not explain the estimate. We forecast future return rates using IMF growth projections, predicting, for example, that India's return rate will rise 12\% by 2035 and China's 7\%. The difference is driven mostly by India's faster projected growth at home. As developing economies close income gaps with the destinations their emigrants have settled in, they recover more of the workers they lost.
This paper studies risk-averse treatment allocation when individuals self-select into treatment based on unobserved characteristics. We develop a framework that combines the marginal treatment effect approach to endogenous selection with a general class of coherent risk measures that capture distributional preferences over welfare outcomes. We show that the planner's problem admits equivalent interpretations in terms of uncertainty aversion, distributional robustness, and worst-case welfare. For law-invariant coherent risk measures, we derive a Kusuoka representation that expresses the planner's objective as a weighted evaluation of different regions of the welfare distribution and characterize the resulting optimal allocation rule. We further establish finite-sample regret guarantees for empirical risk-averse policy learning, showing how the statistical difficulty of learning a policy depends on the planner's sensitivity to adverse welfare outcomes. The framework nests the risk-neutral policy learning model of \cite{Kiatagawa_Tetenov_2018} as a special case. An application to the \cite{Card1995} college proximity data demonstrates that incorporating risk aversion can lead to economically meaningful changes in optimal college admission policies.
We derive the limiting distributions of the $M$-test family of unit root statistics in the nearly integrated nearly white noise (NINW) framework introduced by Nabeya and Perron (1994) in the case of an unknown linear time trend. In the case of known long run variance (LRV), the limiting distributions of the $M^{GLS}$ tests are contaminated by additional noise terms as a result of quasi differencing whereas these terms are less present in the $M^{OLS}$ limiting distributions, both of which display conservative properties under conventional critical values. Furthermore, we prove the Gaussian power envelope in the NINW model is asymptotically equivalent to the standard envelope of Elliott, Rothenberg, and Stock (1996), and that the oracle $M$-tests have inefficient power relative to this benchmark. We then derive the limiting distributions of the feasible statistics and show that the autoregressive estimate of the LRV commonly used overestimates the LRV, creating altered limiting distributions. Finally, finite sample simulations illustrate that, of the procedures considered, no uniformly satisfactory solution exists for handling a series with a large negative moving average coefficient.
Researchers increasingly rely on third-party platforms such as Similarweb and Semrush to measure web traffic when first-party analytics are unavailable. Yet these platforms report model-generated estimates rather than raw data, raising questions about whether their measures preserve the temporal and cross-source variation required for causal inference. As a motivating diagnostic, we examine reported referral traffic around two documented search-engine outages; the absence of visible discontinuities illustrates why preservation of identifying variation cannot be taken for granted. We then characterize three mechanisms, within-source smoothing, cross-source leakage, and treatment-induced calibration error, through which platform processing can generate nonclassical outcome measurement error. Analytical results and a stylized difference-in-differences simulation show that this error can attenuate, amplify, or reverse estimated treatment effects. Our findings caution against using third-party traffic measures based on black-box proprietary models for causal inference.
Bayesian monotonicity is a necessary condition for full implementation in Bayes-Nash equilibrium. Under the sequential equilibrium refinement, however, multi-stage mechanisms can expand the scope for implementation. We show that this additional implementation power is fragile. Whenever a multi-stage mechanism sequentially implements a social choice correspondence that is not Bayesian monotone, we construct arbitrarily small "information perturbations'' and show that there are sequential equilibria under these perturbations that yield inadmissible decisions with probability bounded away from zero. Under these perturbations, players' types are not independent conditional on the state and some player is sometimes uncertain of their payoff type. We show by example that the general result fails if perturbations are required to satisfy conditional independence or "known payoff types.''
Why does a superior technology sometimes spread and sometimes stall, even when its expected returns are much higher? We develop an agent-based model in which firms adopt a digital technology by imitating successful neighbours through a fast-and-frugal heuristic. We compare frequency-based imitation, where firms follow the local majority, with performance-conditioned imitation, where they follow better-performing neighbours. We define a cascade as high final adoption reached through rapid, accelerating takeoff rather than slow accumulation. In the calibrated regime, four results emerge. First, dynamic cascades arise under performance-conditioned imitation, but not under an exogenous adoption hazard or frequency-based imitation. Second, on degree-controlled networks, adoption tips once the network provides sufficient reach, at a small-world rewiring threshold near 0.41, with a 95 percent bootstrap confidence interval of 0.39 to 0.43. This reflects short paths, low clustering, and high-reach nodes rather than one graph statistic. Third, at a fixed mean threshold, concentrating receptiveness in the lower tail can generate cascades that homogeneous thresholds cannot. Fourth, cascade probability depends strongly on the positions of initial adopters and low-threshold pioneers. At human-scale observation degrees, the mechanism is genuine complex contagion: adoption requires reinforcement from multiple successful exemplars, survives a requirement of at least two successful digital neighbours when seeded by a clustered critical mass, and weakens beyond Dunbar-scale neighbourhoods. The model identifies where interventions may have leverage rather than forecasting industries. Broad cultural change is not required for system-wide adoption: a small minority of receptive firms, positioned where the network can transmit success, can be enough. In short, a few misfits can change the world.
Online experiments must often be evaluated before long-term outcomes mature. Under rolling enrollment, these outcomes are observed only for early enrollees, while short-term surrogates are available for everyone. We compare seven estimators across eleven data-generating processes, spanning partial mediation, drift, outcome sparsity, and enrollment-time labeling, with up to $R = 2,000$ replications over more than 500 method-by-scenario cells. We find a sharp robustness-efficiency tradeoff: the surrogate index delivers large efficiency gains when surrogacy holds but its coverage collapses under violations, while PPI-family methods stay asymptotically valid under random labeling at smaller gains. We give the finite-sample variance of PPI++ in the all-units parameterization for a fixed predictor, a joint asymptotic distribution for the two estimators under cross-fitting, and a Hausman-type estimator-disagreement diagnostic, then quantify the detection-damage gap: in the partial-mediation design, where we locate both edges, a band of violations destroys surrogate-index coverage yet is too small to detect on most datasets. On the 64,000-customer Hillstrom experiment the diagnostic rarely flags a violation that biases the surrogate index, and a Cauchy-kernel hybrid of the two estimators inherits 10.7% relative bias; on the 14-million-user Criteo experiment the violation is detected, and subsampling traces detection turning on with scale as damage persists. PPI++ has limits: at a rare-conversion $n = 30,000$ Criteo subsample its empirical coverage is 85.0%. We recommend prespecifying PPI++ with the exact variance as the primary analysis under random labeling and adequate labeled outcome counts, reading the diagnostic as a warning, not a certificate, and treating the surrogate index as a sensitivity analysis.