Computers and Society

2026-07-16 | | Total: 20

#1 AI-Augmented Human Resource Management? Insights from German companies [PDF] [Copy] [Kimi] [REL]

Authors: Yannick Kalff, Katharina Simbeck

This study examines the integration of AI into Human Resource Management in German companies. We ask if and how AI-based technologies are \enquote{augmenting} human resource management. Organisations employ generative AI or predictive analytics to transform traditional human resource functions, to streamline routine tasks and to reallocate resources toward strategic, people-centred activities. Our findings from interviews and group discussions and a survey (N=410) reveal that while AI tools enhance HR analytics capabilities, their adoption mainly serves efficiency and rationalising goals. The introduction of AI tools is shaped by organisational transformation factors such as digital infrastructure, co-determination frameworks, and ethical implications. The research highlights both the strategic potential for improved talent development and the challenges posed by data governance and algorithmic transparency. Overall, this work contributes to understanding the ambiguous role of technological change in HR, which promises to augment predictive capabilities yet serves the ends of efficiency and rationalisation.

Subjects: Computers and Society , Artificial Intelligence

Publish: 2026-07-15 13:46:51 UTC


#2 Persona Migration and Expectation Recalibration in Generative AI Adoption: A Longitudinal Study at a State Department of Transportation [PDF] [Copy] [Kimi] [REL]

Authors: Omidreza Shoghli, Fatemeh Banani Ardecani, Amin Mohamadi Hezaveh

Generative AI tools are increasingly being piloted in public agencies, but limited evidence explains how employee acceptance changes after hands-on use. This study examines Microsoft 365 Copilot adoption during an eight-week pilot at a state Department of Transportation. A matched two-wave survey measured perceived usefulness, perceived ease of use, behavioral intention, and trust before and after participation. After matching and response-quality screening, the sample included 124 employees. Nonparametric tests assessed aggregate changes, k-means clustering identified baseline acceptance personas, and fixed-centroid assignment tracked migration. Open-ended responses were examined using keyword-based content mapping. Perceived usefulness declined significantly after use, suggesting recalibration of expectations, while perceived ease of use, behavioral intention, and trust showed only small, nonsignificant changes. Three baseline personas emerged: Skeptics, Cautiously Positive users, and Champions. Although persona counts changed modestly, individual movement was substantial: 40 percent of Skeptics moved to Cautiously Positive, while 68 percent of Champions moved to less enthusiastic personas. Upward movement was associated with gains in usefulness, behavioral intention, and trust; downward movement was associated with declines in usefulness and trust. Communication and summarization remained stable use cases, while data, chart, and presentation tasks declined. Accuracy and privacy concerns decreased, but job and skills concerns increased. Public-sector AI adoption should be monitored dynamically and supported through persona-specific training, workflow examples, verification routines, and trust-calibration safeguards. The study offers a framework for tracking workforce heterogeneity during enterprise generative AI implementation.

Subjects: Computers and Society , Human-Computer Interaction

Publish: 2026-07-15 13:02:55 UTC


#3 The Environmental Cost of Digital Sovereignty: Water, Energy, and Emissions Impacts of Sovereign AI Infrastructure in the Global South [PDF] [Copy] [Kimi] [REL]

Authors: Muntaser Syed, Marius C. Silaghi, Sheikh Abujar, Sharun Akter Khushbu, Amal El Ahmad

Sovereign AI has become a strategic priority across the Global South, with over \$200 billion in state-led commitments announced between 2024 and 2026. Yet the physical infrastructure that compute sovereignty demands, above all data centers, imposes water, energy, and carbon costs that fall hardest on countries least equipped to absorb them. This paper presents a comparative environmental stress analysis across four cases: the United Arab Emirates, Bangladesh, India, and Africa (with a focus on Kenya). Using publicly available water stress data, grid carbon intensity factors, and GPU power specifications, we model the water consumption, energy demand, and carbon emissions of hypothetical sovereign AI deployments under multiple cooling technology scenarios. We find that a 1,024-GPU cluster using evaporative cooling in the UAE would consume over 30 million liters of water annually in a country classified as ``extremely high'' water stress. In Bangladesh, sovereign AI policy documents call for centralized GPU procurement but do not address where to site data centers in a country where more than a fifth of the land floods in an average year and the power grid struggles to deliver reliable supply. We identify a sovereignty-sustainability trilemma in which no country can simultaneously maximize AI sovereignty, minimize environmental impact, and maintain affordable resource access for citizens. We propose design principles for environmentally responsible sovereign AI, including mandatory water usage effectiveness reporting, climate-vulnerability siting assessments, and a preference for frugal small language models over frontier pre-training in resource-constrained settings.

Subjects: Computers and Society , Distributed, Parallel, and Cluster Computing

Publish: 2026-07-15 05:00:42 UTC


#4 Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System [PDF] [Copy] [Kimi] [REL]

Authors: Teri Rumble, Javad Zarrin, P. George Lovell, Ruth Falconer

This paper is an extension of a paper presented at the ICAART 2026 conference, which introduced LEA (Learning Engagement Assistant), an adaptive AI tutoring agent combining course-specific Retrieval-Augmented Generation (RAG) with structured Knowledge Component (KC) models across integrated Chat, Tutor, and Quiz modes. That prior work validated LEA on a single STEM course (CMP511) exclusively through simulation, using synthetic learner agents. This paper extends that work by reporting the first classroom deployment of LEA with real students (n = 8, CMP511) and the first empirical test of its cross-course scalability, deploying the system across three courses spanning two academic levels and two disciplinary domains. The study reveals a divergence from simulation predictions across modes, showing that synthetic evaluation alone cannot anticipate all aspects of real deployment. A RAGAS-based cross-course scalability evaluation (660 questions) finds Answer Relevancy and Context Precision broadly stable across courses (0.88-0.94 and 0.88-0.90 respectively), while Faithfulness declines with curriculum distance from the system's original course (0.69 to 0.50), a preliminary finding that may reflect generation logic tuned to the system's original subject rather than a scalability limitation. These findings suggest that while the orchestration layer requires no modification, full course-agnosticism of all downstream components requires further investigation.

Subjects: Computers and Society , Artificial Intelligence , Human-Computer Interaction

Publish: 2026-07-15 01:25:22 UTC


#5 The tragedy of the cognitive commons: collective intelligence beyond AI-induced knowledge collapse [PDF] [Copy] [Kimi] [REL]

Authors: Maher Kallel, Mohamed El Louadi

In a recent dynamic model by Acemoglu, Kong and Ozdaglar (2026a) agentic AI can cause a self-reinforcing deterioration of humanity's common knowledge base in what they call knowledge collapse. The model is based on a natural complementarity between the cumulative general knowledge of humans and locally generated context-specific knowledge, and on a learning externality that means that we all contribute to the private signal and the thin public signal that feeds the collective stock. If agentic AI can substitute for the private signal but not rebuild the public signal, and when human effort is sufficiently elastic, we can reach a low-knowledge equilibrium. This paper offers a measured appraisal of the model. It assesses the model in terms of what the popular summaries say is common knowledge, the world literature against knowledge commons and model collapse, and the partial empirical evidence, such as a 25% decline in public knowledge sharing on Stack Overflow. It then presents five structural criticisms of the model: its fixed knowledge taxonomy, the assumption that AI cannot provide general knowledge, the unmeasured effort elasticity parameter, and the political futility of the model to be implemented. We argue that the model's value lies not in any guaranteed catastrophe but in identifying a credible negative externality, the tragedy of the cognitive commons, whose remedy is a better aggregation of human knowledge. The paper outlines a research agenda for the study of effort elasticity and knowledge commons.

Subject: Computers and Society

Publish: 2026-07-14 21:13:58 UTC


#6 Analyzing Curricular Pattern Complexity Using AI to Improve On-Time Graduation Rates [PDF] [Copy] [Kimi] [REL]

Authors: Lynn Vonderhaar, Juan Couder, Siri Siqveland, Omar Ochoa, James Pembridge

The rise of Artificial Intelligence (AI) enables automatic analysis of large amounts of data. Previously time-consuming and labor-intensive tasks can be completed much more efficiently with the use of AI. This work uses AI techniques to analyze and revise curricular patterns in an undergraduate degree for Software Engineering. Curricula often have long sequences where failure to pass a class within the sequence may jeopardize completion of the degree within four years. Manual analysis and revision of curricula by university faculty is a lengthy and labor-intensive process, causing changes to occur rarely and making it impossible to keep up with the changing needs of students. This work reduces the time-to-change for curricula and reduces bottlenecks and graduation delays by using Large Language Models (LLMs) to analyze curricular patterns and suggest revisions.

Subjects: Computers and Society , Artificial Intelligence

Publish: 2026-07-14 02:02:01 UTC


#7 The Hitchhiker's Guide to Monoculture [PDF1] [Copy] [Kimi] [REL]

Author: Gordon Burtch

Large language models (LLMs) often produce homogeneous outputs, raising concerns that AI coding assistants may lead to convergence in the software artifacts that developers create. Whether this occurs in practice is unclear because developers interactively prompt, evaluate, modify, and reject model outputs, and because outputs vary with prompt and repository context. I examine code homogenization using Kaggle contest submissions from 2019 to mid-2026. I first document widespread convergence toward the random seed value 42, consistent with LLMs reinforcing a longstanding convention in programming culture. I then study homogenization more broadly, at two levels of aggregation and abstraction. At the submission level, I measure the average pairwise similarity of submissions within contests. At the contest level, I measure the conceptual span of submitted code, motivating distinct measures for each: TF-IDF representations, which capture surface syntax, and Voyage 3 code embeddings, which capture code intent and semantics. The results demonstrate substantial syntactic homogenization at both the individual and collective levels: individual submissions have become more alike in literal syntax and code structure, while the latent dimensionality of syntactic variation has narrowed. In contrast, I find little evidence of semantic homogenization, individually and collectively. Average semantic distance remains essentially flat, and the contest-level latent dimensional span of semantic approaches remains stable, with evidence suggesting it has even expanded modestly. These findings suggest that AI coding assistants are certainly standardizing implementation details, yet they have not yet produced evidence of homogenization in the approaches and problem-solving strategies coders employ.

Subjects: Computers and Society , Artificial Intelligence , Software Engineering

Publish: 2026-07-12 21:01:06 UTC


#8 LessonBench-V1: A Benchmark Dataset for Evaluating AI Lesson Generation Agents [PDF] [Copy] [Kimi] [REL]

Authors: Ravidu Suien Rammuni Silva, Ahmad Lotfi, Isibor Kennedy Ihianle, Golnaz Shahtahmassebi, Jordan J. Bird

Large Language Model (LLM) based AI educational content generation systems are increasingly being developed, yet no standardised benchmark exists to systematically evaluate them. This study introduces LessonBench-V1, a benchmark dataset comprising 647 human-written lessons paired with LLM-based reverse-engineered lesson plans across 240 STEM topics spanning mathematics, physics, chemistry, and computer science. The lessons are drawn from 97 trusted open sources, including LibreTexts, Brilliant.org and GeeksForGeeks. Each lesson plan is human-reviewed and produced through a pedagogically grounded methodology that synthesises Bloom's Taxonomy, Gagné's Events, Merrill's First Principles, and the 5E Instructional Model. The lesson plans capture 3,620 learning objectives with pedagogical metadata, enabling systematic, reproducible evaluation of lesson-generation AI agents and supporting further research. The study further proposes a three-dimensional evaluation pipeline for use with the dataset.

Subjects: Computers and Society , Artificial Intelligence , Machine Learning , Multimedia

Publish: 2026-06-12 07:00:36 UTC


#9 Final Authority in AI Governance: Frontier-Provider Sovereignty and Action-Centered Deployer Governance [PDF] [Copy] [Kimi] [REL]

Author: Zexun Wang

This paper examines where final authority should sit once capable AI systems are embedded in organizational workflows. It compares two governance models. The first, frontier-provider sovereignty, assigns privileged authority to the provider of the most capable models and is reflected in contemporary arguments for frontier-model testing, release gating, transparency duties, and compute-related controls. The second, action-centered deployer sovereignty, places final authority over high-impact actions with the organization that authorizes the action, embeds it in a business process, and bears the downstream legal, operational, and commercial consequences. The paper combines comparative reading of public governance frameworks with implementation-informed analysis of runtime heterogeneity and enterprise control requirements. It compares EU AI Act guidance, the NIST AI Risk Management Framework, Singapore's Model AI Governance Framework for Agentic AI, recent Japanese AI policy instruments, and Canada's voluntary code and managerial guidance. Across these materials, the paper finds stronger support for distributed operational accountability than for unilateral frontier-provider control. It further argues that rapid enterprise adoption, declining provider transparency, and widening control gaps increase the value of a portable governance layer centered on governed action rather than on provider-native session objects. The conclusion is layered rather than absolutist: strong upstream authority remains justified for frontier capability gating, but final authority over concrete enterprise action is better located with the deployer and consequence-bearer.

Subjects: Computers and Society , Artificial Intelligence

Publish: 2026-06-12 03:28:06 UTC


#10 Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants [PDF] [Copy] [Kimi] [REL]

Author: Dipesh Tharu Mahato

Safety evaluations for dual-use biology assistants often measure base-model capability, refusal behavior, or jailbreak success. These metrics miss a deployment question: for a fixed base model, how does the access condition users actually see change benign utility and harmful actionable assistance? I introduce safeguard-conditioned uplift, a protocol for comparing deployed access conditions through a human-judged utility-risk frontier. I evaluate Claude Sonnet 4.6 and Gemini 3.5 Flash under helpful prompting, safety prompting, and an external safeguarded assistant on a 108-task surrogate benchmark, with the headline claim restricted to a locked 18-task held-out split. In a 600-row blinded human audit, the safeguarded assistant reduces harmful actionability relative to helpful prompting by -0.063 over 49 matched response pairs, with bootstrap 95% interval [-0.117, -0.011], while correctness changes by +0.009 with interval [-0.057, +0.077]. Adaptive, Test-B, cue-ablation, and controller-baseline checks support the measurement story but also show non-dominance: safety prompting is often strongest for Claude, while external control helps more for Gemini and can reduce benign utility. The contribution is not a universal defense. It is a deployment-level evaluation target, plus a learned risk-budgeted calibration procedure, for measuring how user-facing access conditions move the utility-risk frontier.

Subjects: Computers and Society , Artificial Intelligence

Publish: 2026-06-12 00:57:14 UTC


#11 Designing Safety-Constrained LLM Systems for Public Health Information Access [PDF] [Copy] [Kimi] [REL]

Authors: Ben Torkian, Jun Zhou

We present the design and implementation of a safety constrained large language model (LLM) system for public health information access, focusing on maternal and child health (MCH) resource navigation. While LLM based systems offer flexible and natural interfaces for information retrieval, their deployment in healthcare contexts introduces risks related to safety, trust, and uncontrolled generation. This work explores practical design patterns for constraining LLM behavior in safety critical environments. We introduce a multi-layered architecture that integrates domain-restricted retrieval augmented generation (RAG), strict boundary enforcement to prevent medical advice, anonymous multiuser session management, and comprehensive audit logging for monitoring and compliance. A key aspect of the design is a controlled data pipeline that grounds all responses in curated public health resources, avoiding reliance on the model pretrained medical knowledge. We implement the system in a real world public health setting and conduct scenario-based validation across in scope, out of scope, and emergency queries. Results show consistent enforcement of safety constraints, reliable resource grounding, and stable system performance, with an average response time of 5.3 seconds. Beyond the specific application, we discuss design trade offs and lessons learned in balancing safety, usability, and system flexibility. Our findings provide practical guidance for deploying LLM based systems in healthcare and other domains where strict information boundaries and accountability are required.

Subjects: Computers and Society , Artificial Intelligence

Publish: 2026-06-11 21:06:55 UTC


#12 Early Adoption of Agentic Coding Tools by GitHub Projects [PDF] [Copy] [Kimi] [REL]

Authors: Maliha Noushin Raida, Daqing Hou

Agentic coding tools are increasingly capable of generating and submitting pull requests (PRs) to software projects, introducing new forms of human-agent collaboration in software development. While prior studies have examined PR-level outcomes of agent-generated contributions, less is known about how agentic coding tools are adopted and managed at the project level. In this paper, we analyze 25,264 agentic PRs from 2,361 popular GitHub repositories to investigate (1) the adoption of agentic coding tools, (2) project-level agentic PR productivity, and (3) human-agent collaboration patterns. Our results show that the median repository generates only one to two agentic PRs during a three-month period, indicating that intensive adoption remains concentrated in a small subset of projects. At the same time, small projects (1-5 contributors) exhibit higher participation ratios and average levels of agentic PR activity than medium-sized and large projects. We also observe substantial variation in project-level agentic PR productivity. While a small number of projects exceed an industry-reported estimate of 36 PRs per participant during the three-month observation period, most projects remain below this threshold. Finally, human-agent collaboration is dominated by a single-human oversight model, in which one developer reviews and/or modifies the agent's contributions, while multi-human collaboration patterns remain uncommon. These findings provide early empirical evidence on how open-source projects organize human oversight around agentic coding tools and suggest that successful integration of agent-generated contributions depends not only on advances in agent capabilities but also on the human and organizational processes that govern their use. Because this study captures an early snapshot of agent adoption, future work should continue to track how adoption patterns evolve over time.

Subjects: Software Engineering , Artificial Intelligence , Computers and Society , Machine Learning

Publish: 2026-07-15 17:05:06 UTC


#13 Epidemic Informatics and Control: A Holistic Approach from System Informatics to Epidemic Response and Risk Management in Public Health [PDF] [Copy] [Kimi] [REL]

Authors: Hui Yang, Siqi Zhang, Runsang Liu, Alexander Krall, Yidan Wang, Marta Ventura, Chris Deflitch

This paper presents a holistic systems informatics approach, i.e., Define, Measure, Analyze, Improve, and Control (DMAIC), for epidemic response and management through the intensive use of data, statistics and optimization. Despite the sustained successes of system informatics in a variety of established industries such as manufacturing, logistics, services and beyond, there is a dearth of concentrated review and application of the data-driven DMAIC approach in the context of epidemic outbreaks. First, we define specific challenges posed by epidemic outbreaks to populational health, health systems, as well as economic challenges to different industries such as retailing, education and manufacturing. Second, we present a review of medical testing and statistical sampling methods for data collection, as well as existing efforts in data management and data visualization. Third, we discuss the importance to realizing the full potential of data for epidemic insights, and emphasize the need to leverage analytical methods and tools for decision support. Fourth, an epidemic brings imperative changes to health systems. We discuss the new trend of healthcare solutions to improve system resilience, including telehealth, artificial intelligence, resource allocation, and system re-design. In closing, prescriptive approaches are discussed to optimize the health policies and action strategies for controlling the spread of virus. We posit that this work will catalyze more in-depth investigations and multi-disciplinary research efforts to accelerate the application of system informatics methods and tools in epidemic response and risk management.

Subjects: Systems and Control , Computers and Society , Applications

Publish: 2026-07-15 14:54:09 UTC


#14 Design of policy digital twins incorporating multi-level agent based modelling [PDF] [Copy] [Kimi] [REL]

Authors: Matt Tipuric, Alejandro Beltran, Nick Malleson, Dani Arribas-Bel, David Wagg

Digital twins are used across many industries to enable better decision making. However, while policy makers at all levels (including city, national and supranational scales) have expressed a desire to integrate digital twins into their workflows, this adoption has been slow to materialise. In this paper, we discuss the key issues associated with policy digital twins, and the ways in which they differ from, and are similar to, their counterparts in other areas. We describe how multi-level agent based modelling can be used within policy digital twins to include the effects of human behaviours on outcomes; an aspect that is often largely overlooked. We also describe how digital twins can be designed for policy use cases, and present as a case study the design of a policy digital twin incorporating multi-level agent based modelling to aid a UK city council (local authority) in delivering energy transition policy. After describing both the design method used and the resultant digital twin, we discuss the effectiveness of both, as well as how the ways in which different contexts might shape the future architecture of the digital twin.

Subjects: Systems and Control , Computers and Society

Publish: 2026-07-15 12:28:24 UTC


#15 Extending Liquid Rank Toward Multi-Source Reputation Aggregation [PDF] [Copy] [Kimi] [REL]

Authors: Nejc Znidar, Anton Kolonin

In this paper, we present an extension of liquid rank reputation systems that enables the aggregation and blending of multiple heterogeneous reputation sources into a unified reputation score. The proposed framework supports the incorporation of external reputational signals alongside internally generated reputation, allowing influence to reflect participation and contribution across multiple contexts and subsystems. By introducing explicit weighting and blending mechanisms, the model provides fine-grained control over the relative impact of individual reputation sources, making it adaptable to diverse governance and coordination scenarios involving both human and machine agents. The resulting approach extends existing liquid rank systems and offers a flexible foundation for designing reputation-based governance mechanisms in complex socio-technical environments.

Subjects: Multiagent Systems , Computers and Society

Publish: 2026-07-15 09:04:33 UTC


#16 AI advice suppresses people's willingness to say "I don't know", even when the advice is wrong and accuracy is incentivized [PDF] [Copy] [Kimi] [REL]

Authors: Chiara Marcoccia, Walter Quattrociocchi, Valerio Capraro

Knowing when to say "I don't know" is fundamental to human judgment, yet AI assistants offer a fluent answer to almost any question. In five experiments (N = 3,132; four preregistered, one direct replication), participants answered difficult questions and could always decline to respond. We engineered the questions so that AI advice was wrong, separating AI use from its accuracy. Merely having access to AI nearly eliminated participants' willingness to suspend judgment, and this held whether the advice was actively requested or simply displayed. Consequently, participants answered more questions but were correct about a third as often as when AI was unavailable-yet their confidence nearly doubled. Incentivizing accuracy and penalizing inaccuracy led participants to seek and follow AI advice less, answer more accurately, and suspend judgment more often, though still far less than when AI was unavailable. As AI suggestions grow ubiquitous and unsolicited, they may not simply affect answer accuracy; they may even alter the metacognitive threshold at which people decide whether they know enough to answer.

Subjects: Artificial Intelligence , Computers and Society , Human-Computer Interaction

Publish: 2026-07-15 08:04:14 UTC


#17 When Rubrics Change: Cross-Rubric Generalization for Critical Thinking Essay Scoring [PDF] [Copy] [Kimi] [REL]

Authors: Nischal Ashok Kumar, Payu Wittawatolarn, Sana Kang, Marisa C. Peczuh, Blair Lehman, Ryan Baker, Caitlin Mills, Sherry Lachman, Ruochen Sun, Andrew Lan

Automated essay scoring (AES) research has largely focused on cross-prompt generalization, where essays from unseen prompts are scored while the scoring criteria are typically held constant. In practice, however, educators may revise or even introduce new rubrics in their scoring task, to evaluate different aspects of essays. We study cross-rubric generalization: training on essays labeled under one set of rubrics and evaluating on previously unseen rubrics, which target different aspects of the essay. We use a Large Language Model (LLM) fine-tuning framework with two components: rubric-agnostic intermediate representations, called traits, and target-essay supervision under seen rubrics during training. On an AES dataset augmented with multiple rubric-defined labels of student critical thinking skills, we find that traits improve macro F1 by 5.0% over a baseline without traits in the hardest setting, where both target rubrics and target essays are unseen during training. We further find that increasing target-essay supervision improves performance, with our best fine-tuned open-source Llama-based model outperforming GPT-5-mini prompting by 2.1% macro F1 and trailing GPT-5 by 1.9%. These results show that trait-based intermediate structure and controlled supervision improve generalization to unseen rubrics.

Subjects: Computation and Language , Computers and Society

Publish: 2026-07-15 04:34:14 UTC


#18 xChk: Bring Your Own Identity -- Heterogeneous Assurance with Verifier-Determined Sufficiency [PDF] [Copy] [Kimi] [REL]

Author: Sean MacGuire

We present xChk, a reference identity provider for Bring Your Own Identity (BYOI): users enroll via heterogeneous proofs (government KYC, corporate SSO, WebAuthn/FIDO2, professional networks, live verification, longitudinal activity, behavioral signals) and disclose them as portfolio claims in standard OAuth 2.0 / OpenID Connect (OIDC) tokens, while each relying party applies its own sufficiency policy - the IdP transports claims and may evaluate an RP-supplied evidence policy for consent, but does not adjudicate access. Enrollment depth varies by modality (some paths are user-initiated; org KYB and officer binding are operator-assisted). xChk also supports human-in-the-loop attestation for high-risk actions: humans can initiate attestations directly (browser UI / POST /api/attestations), and AI agents acting under those principals can trigger the same gateway via scope-gated authorize/attest - hash-chained human approvals on a shared verification graph (humans via OIDC; agents via API keys). A production deployment at https://in.xchk.io ships both initiation paths with bilateral RP evaluation at consent; one documented relying party (https://crabbyed.com, Appendix B) exercises Login with xChk.

Subjects: Cryptography and Security , Computers and Society

Publish: 2026-07-15 01:19:35 UTC


#19 AI-Native Insurance for Agentic AI: Pricing, Underwriting, and End-to-End Automation [PDF] [Copy] [Kimi] [REL]

Author: Quanyan Zhu

Agentic AI introduces new insurance challenges because autonomous AI systems can make decisions, invoke tools, modify external environments, and interact with third-party services. This paper develops an AI-native mathematical framework for underwriting, pricing, and contract design for agentic AI deployments. A deployment is represented by a risk state that captures autonomy level, operational authority, permission exposure, governance maturity, and dependency concentration. The framework maps the risk state to event probabilities, loss severities, governance costs, premiums, deductibles, coverage allocation, and policy covenants, and formulates an optimization problem for insurance contract design under participation, profitability, and incentive compatibility constraints. The paper establishes structural properties of insurability, including characterization of an insurability region, monotone deterioration of feasibility with increasing exposure, and governance certification thresholds. Insurance is further interpreted as both an operational cost and a regulatory mechanism for AI deployment. A healthcare case study illustrates contract optimization, sensitivity analysis, and automated claims processing for agentic AI systems.

Subjects: Artificial Intelligence , Cryptography and Security , Computers and Society

Publish: 2026-07-14 19:46:21 UTC


#20 Beyond AI-Generated Labels: Watermarking, Co-Creation, and Conflation of AI-Generation with Disinformation [PDF] [Copy] [Kimi] [REL]

Authors: Federico Germani, Giovanni Spitale

Watermarking is often presented as a straightforward solution for distinguishing AI-generated from human-generated content, enabling platforms and regulators to trace synthetic content and detect AI-generated outputs at scale. This paper examines whether such mechanisms meaningfully address the epistemic and ethical challenges that arise in domains where the central concern is not the automation of content production, but the accuracy, intent, and deceptive potential of messages. We argue that extending watermark-based approaches to these settings is conceptually and practically misguided. Invisible watermarking encodes only model origin; when operationalized into visible AI-generated labels, it reduces complex creative processes to a misleading binary and provides no information about truthfulness. Such labels may stigmatize legitimate uses of generative tools while encouraging misplaced trust in unmarked content. Here we propose an alternative approach centered on process transparency and information literacy. We argue that these measures address the epistemic and ethical dimensions of AI-generated disinformation more effectively than visible watermark labels, reframing authorship as a transparent human practice rather than a binary indicator of machine involvement.

Subjects: Cryptography and Security , Computers and Society

Publish: 2026-07-13 08:44:12 UTC