2026-09-16 | | Total: 14
Telework has expanded rapidly, and understanding the behaviors employees enact under it has become correspondingly important. This study develops an integrative conceptual framework specifying the multidimensional nature of telework behavior. A systematic literature review following PRISMA identified 114 review articles, which were analyzed using constructivist grounded theory. Six behavioral dimensions were identified, covering performance, communication, environmental, task, policy, and well-being conduct. Antecedents were grouped into individual factors, job characteristics, organizational norms, technological factors, and work environment factors. Outcomes were grouped into job satisfaction, productivity, turnover, health and well-being, work-life balance, and social isolation. Three contextual moderators were identified, namely telework modality, telework preference, and cultural and national context. Fourteen propositions link these categories and are advanced as testable claims rather than established findings, since they are derived from published review evidence and have not been empirically tested. The framework proposes telework behavior as the construct through which the conditions of remote work are translated into employee and organizational outcomes, and it specifies a content domain from which measures of the construct can be developed. It also indicates where organizations can direct policy, training, and intervention.
I study dynamic treatment effects in panel data under staggered adoption when treatment timing depends jointly on unobserved time-invariant heterogeneity and time-varying pretreatment covariates, including lagged outcomes. Untreated potential outcomes follow a nonparametric dynamic panel model that allows flexible interactions between time-varying covariates and latent heterogeneity. I use pretreatment outcome histories to find individuals with similar time-invariant latent factors, and the key requirement is that these histories are sufficiently informative about those latent factors. I develop an identification strategy for the dynamic average treatment effect on the treated (ATT) and propose kernel-based doubly robust estimators for the dynamic ATT. I further combine double cross-fitting with undersmoothing and show that, under suitable regularity conditions, the proposed estimators are $\sqrt{n}$-consistent, asymptotically normal, and asymptotically unbiased. The simulation study demonstrates that the proposed method provides accurate inference across a wide range of data-generating processes. I illustrate the method with an application to the U.S. family planning program studied by Bailey (2012) and reestimate its effect on fertility rates.
We investigate Empirical Bayes (EB) methods in the context of compound adaptive experiments, where the arm distribution in each experiment follows a normal distribution with an unknown mean that we seek to estimate. There are two main EB strategies: $g$-modeling, which estimates the prior by maximizing the marginal likelihood, and $f$-modeling, which derives posterior means directly from the empirical distribution of the observations. We show that $g$-modeling continues to be a valid EB procedure even when it incorrectly assumes that data are collected exogenously; its validity does not depend on the particular sampling algorithm or on whether sample sizes are endogenous. In practice, one can apply standard $g$-modeling techniques by acting as though the data were exogenously sampled. We extend regret guarantees from exogenous sampling to adaptively generated data. By contrast, naively applying the Tweedie formula based on the marginal density of the observed data, as in standard $f$-modeling, can produce biased rules under adaptive sampling. We corroborate the robustness of $g$-modeling through simulations with widely used adaptive algorithms and demonstrate its applicability using a real-world dataset consisting of multiple sequential experiments.
This paper develops Wald inference for least-squares estimation of linear regression models on dyadic data, accommodating configurations where multiple observations share the same pair of units (e.g., directed flows, multilayer networks, and dyadic panels). We establish that the dyadic-robust Wald statistic is asymptotically $χ^2_q$ for an arbitrary nonrandom sequence of full-rank restrictions, under a single condition on the accumulation of dependence. Throughout, no convergence rate is assumed or estimated, permitting the condition number of the score's variance matrix to diverge. We further propose a delete-one-unit jackknife alternative that is positive semidefinite by construction. This jackknife statistic attains the same asymptotic limit under one additional condition on dyad multiplicity and remains asymptotically conservative when that condition fails. A supplement contains all proofs, Monte Carlo experiments featuring estimated coefficients that converge at heterogeneous rates, and an empirical gravity application to bilateral trade.
Consider a situation wherein a decision maker sequentially searches for the best alternative among heterogeneous options with an arbitrary search order. The agent partially learns the value of an option when inspecting it. The information structure jointly determines the ex-ante and ex-post value of investigating each option, thereby shaping the entire learning path. We characterize the set of all search behaviors compatible with some information structure, which forms a polytope. A single information structure rationalizes all these search behaviors, which minimizes the agent's welfare among all information structures. Under certain symmetry assumption over primitives, we also examine the set of all feasible pairs of search and choice behaviors, with and without option sellers' price competition, and prove the same results. Our study not only provides a simple framework for analyzing how information shapes ordered learning paths, but also offers sharp insights into information design and identification problems in such environments.
Numbers this large invite a defensive reflex: reach for the reassuring figure and move on. By the most widely cited measure, the World Bank's extreme-poverty line of \$3.00 per day (2021 PPP), approximately 847 million people, or 10.4% of the world's population, lived in poverty in 2024. That figure is accurate as measured. It is also, by design, a floor: a threshold built to mark bare survival in the poorest economies, not to describe what it takes for a family anywhere to live with basic security. The central argument of this report is that the reassurance offered by that single line is largely an artifact of how we chose to measure. Held to the \$3.00 floor, roughly 847 million people are poor; held to the standards of their own societies, the number is several times larger, and even the most cautious, honest count exceeds a billion. Measured against income standards appropriate to each country's level of development, some 3.7 billion people fall below the World Bank's \$8.30/day upper-middle-income reference line, roughly 1.1 billion are multidimensionally poor, and around 2.3 billion are food insecure. These are not competing errors but answers to different questions. This report does not dispute the official data. It takes those figures as accurate and asks what they do and do not capture. It disaggregates by geography, age, and gender and supplements the monetary count with multidimensional measures of food security, working poverty, social protection, and relative poverty. It also documents a data-quality problem that runs in one direction: survey coverage is weaker in regions with the worst poverty, so official figures are far likelier to understate deprivation than to overstate it.
The Money Pump Index (MPI) of Echenique et al. (2011) measures the severity of consumer irrationality, but computing the exact mean and median MPI over all revealed preference cycles is NP-hard (Smeulders et al., 2013). Existing solutions rely on heuristic proxies, such as evaluating only shorter cycles or bounding the MPI. By framing revealed preferences as a directed graph, this paper projects choice violations onto fundamental cycle bases, which are minimal sets of linearly independent cycles that span the graph's entire cycle space. This yields computationally tractable estimators for the mean and median MPI that are asymptotically equivalent to the original MPI. Applying this methodology to the scanner dataset analyzed by Echenique et al. (2011) and Smeulders et al., (2013), the proposed estimators compute quickly and with negligible small-sample bias.
This paper is the first work at the intersection of game theory and property testing, giving algorithms and lower bounds for efficiently testing whether an allocation mechanism is incentive compatible (IC). We propose distinguishing whether a mechanism is $ε$-far from being IC, i.e., when it observes many monotonicity "violations." Conceptually, inspired by the literature on Boolean function monotonicity testing, we construct a tester for discrete single-parameter allocation rules. Technically, our work is the first to consider monotonicity testing of vector-valued functions on the hypergrid. We give a $\tilde{O}(n/ε)$-query algorithm to test whether a function (representing n-player allocation mechanisms) is coordinate-wise monotone versus $ε$-far from it. We also show a matching lower bound: the class of coordinate-wise monotone vector-valued functions on a Boolean hypercube or hypergrid requires $\tildeΩ(n/ε)$ queries to test whether it is $ε$-far from monotonicity, and this holds even if the tester is two-sided and allowed to make adaptive queries. Finally, we extend our upper bound to and give a tester of the same query complexity for pricing functions of allocation mechanisms. This requires overcoming the technical challenge that the path in function space to the closest IC mechanism may involve interdependent changes to both the price and the allocation rule.
Policy learning aims to determine who should be treated based on individual characteristics. In high-stakes settings such as medicine and public policy where safety is a central concern, improving the average outcomes alone may not be sufficient: decision makers may also seek to protect individuals from harm, in line with the Hippocratic principle of ``do no harm.'' In this paper, we propose \textit{conformal policy learning} (CPL), a policy learning procedure with a new distribution-free safety guarantee that controls the probability of assigning treatment to an individual who would be harmed relative to control. CPL views each treatment decision as testing a hypothesis of counterfactual harm and assigns treatment by thresholding conformal p-values. These p-values use observable proxies and selective calibration to address the challenge that the potential outcomes under comparison are never simultaneously observed. For randomized experiments, under standard exchangeability conditions, CPL provides finite-sample safety guarantee at a user-specified level, without imposing any outcome modeling assumptions. Moreover, when the outcome model is consistently estimated, CPL achieves asymptotically optimal welfare subject to the safety constraint. In observational studies, CPL with learn-then-balance weights achieves doubly robust safety guarantees. We evaluate CPL through extensive simulations and apply it to an empirical study of AI-powered interventions designed to reduce conspiracy beliefs.
Green economic complexity provides a generalizable framework for examining countries' productive capabilities in a defined product set. We apply this framework to AI-enabling goods within the full product space, linking current specialization with adjacent diversification opportunities. Using BACI exports for 2007-2023 and 103 AI-enabling goods, we measure complexity-weighted specialization (AECI), product-level adjacent opportunities (AIAP), and average complexity-weighted relatedness of remaining candidates (AECP). In 2023, Japan leads AECI, while China leads AECP; portfolio breadth accounts for much of the variation in raw AECI. Initial raw potential is positively associated with subsequent changes in the AI-enabling export share, but its associations with changes in AECI and specialization counts are not statistically significant at the 5% level. Our contribution is a trade-based assessment of AI-enabling productive capabilities and related opportunities. The results and public dashboard provide a preliminary complement to publication and patent indicators, not a comprehensive measure of national AI performance or a validated forecast of diversification.
Model robustness analysis estimates an effect across a multiverse of specifications that pools control sets identifying the declared estimand with sets that condition on mediators or colliders. We propose stating rival assumptions about contested controls as a small set of candidate causal graphs, enumerating the adjustment sets each graph licenses, and reporting robustness metrics conditional on each graph. A finite-mixture identity splits the licensed multiverse's dispersion into within-graph and between-graph components; the between-graph share is a conditional descriptive summary whose reading depends on the candidate set, the weights, and a common estimand. Simulations examine misleading pooled robustness assessments and the limits of the decomposition. Applications to hurricane fatalities, job training, and union wages show fragility that survives every graph, instability produced by unlicensed specifications, and a fragility verdict concealing a significant premium in each adjustment-identified candidate world. An R package implements the workflow.
The return-on-investment (ROI) constraint is central to many auctions, particularly in online advertising, where a bidder is unwilling to pay more than a fixed fraction of the value obtained. We study truthful and revenue-maximizing auctions for ROI-constrained bidders. We first characterize truthful auctions when both valuations and ROI constraints are private, showing that the allocation rule uniquely determines the payment rule. Building on this characterization, for multiple bidders we introduce $σ$-increment mechanisms that resemble Myerson's optimal mechanism~\cite{journals/mor/Myerson81}; as $σ$ vanishes, these mechanisms become asymptotically optimal among deterministic truthful mechanisms, and their revenue approaches at least a $1/\bar r$ fraction of the optimal expected revenue over all truthful mechanisms, where $\bar r$ is the largest possible ROI constraint. In the single-bidder setting, we prove that every truthful auction can be replaced by a convex pricing function with weakly higher payments for every type, and we derive the optimal pricing functions when either the valuation or the ROI constraint is public.
Anthropogenic forcing components follow different long-run paths, while persistent temperature change can involve distributional changes beyond the mean. Scalar regressions aggregate these components and retain only mean temperature, obscuring how distinct forcing paths relate to persistent distributional change. We develop new testing, estimation, and inference methods for long-run relations between an integrated predictor vector and a density-valued response. These comprise a residual-based test of between-cointegration (whether predictor trends account for all stochastic trends in the response density), a fully modified least-squares estimator of predictor-specific functional responses, and simulation-based inference for interpretable projections. We apply the methods to densities of observed local temperature anomalies and anthropogenic effective radiative forcing divided into CO$_2$ and non-CO$_2$ portfolios. The test results are consistent with persistent movements in these portfolios statistically accounting for the persistent evolution of the anomaly distribution, with no additional stochastic trend detected in the residual. A joint test rejects the common-response restriction imposed by aggregating the two portfolios. The fitted CO$_2$ response mainly shifts mass toward warmer anomalies and increases central concentration, whereas the non-CO$_2$ response produces a smaller shift but greater dispersion and off-center reshaping. Positive fitted mean responses for both portfolios conceal these contrasts, demonstrating the information lost through scalar aggregation.
As artificial intelligence's capabilities improve, it is increasingly viewed as a general scientific method. But how true are these claims? Does AI outperform all techniques, or only some, and how is this changing? To assess the claims, we assemble a corpus of 2,507 head-to-head comparisons between AI and other scientific analysis techniques across 27 scientific disciplines from papers published between 2000 and early 2025. We find a profound dichotomy. Relative to traditional statistics, AI often outperforms, but at a significantly higher computational cost. But there are also nearly a quarter of cases where AI is both more expensive and performs worse than traditional statistical techniques and this fraction has been stable for a decade. Relative to scientific computing, AI often underperforms, but at lower computational cost. This has begun to change: since 2020, AI's performance against scientific computing has notably strengthened and it now outperforms on more than half of comparisons. These patterns suggest that AI is therefore not a universal replacement for existing methods, but rather a valuable -- and improving -- part of a new AI-enabled scientific frontier.