2026-09-11 | | Total: 23
WiFi gesture recognition is accurate in fixed deployments but often degrades when user orientation, available links, or transceiver placement changes. Unlike ordinary domain shifts, these changes alter the wireless observation operator, so the same motion is expected to produce different measurements. Existing methods nevertheless pursue domain-invariant features and largely overlook changing layouts and observation configurations. Yet changing the observation operator also changes which task-relevant motion cues are physically observable, rather than merely altering the appearance of a fixed set of cues. Under a local linearization of the WiFi forward process, we derive a common task-observability condition under which a strict common linear representation is recoverable from every geometry-induced operator while preserving the gesture task. When the condition fails, enforcing stronger alignment across additional heterogeneous source operators may discard task-relevant cues still observable under individual operators. We therefore present MotionQ, which generates an operator-conditioned two-support motion measure for each candidate geometry. A motion quotient removes only the arbitrary ordering of its unlabeled supports and is represented by permutation-invariant central moments. Rather than matching quotients across operators, single-link-retention interventions encourage each view to retain information sufficient for gesture recognition. Extensive evaluations show that MotionQ is robust to extrapolative observation operators.
Maritime Autonomous Surface Ships (MASS) and AI- supported decision assistants are expected to transform maritime operations, but their safe integration depends on how maritime professionals perceive and trust such systems. This paper presents a survey study on maritime stakeholders' attitudes toward an AI-supported assistant in collision-avoidance scenarios. Participants evaluated technology anxiety, trust in automation, and explanation quality using established and adapted questionnaires, complemented by sentiment and thematic analysis of open-ended responses Results indicate a generally positive disposition toward maritime technology, no clear age-related differences in openness, stable trust across scenarios, and more scenario-sensitive, multidimensional explanation ratings. Open responses showed that participants valued support for decision-making, situation awareness, and confidence-building, while raising concerns about AI reliability, over- reliance and loss of expertise. The findings suggest that maritime AI systems should not focus solely on increasing automation or trust, but on supporting calibrated reliance through transparent, reliable, and operationally meaningful design with domain experts in the loop.
Extended Reality (XR) technologies have received growing attention in accessibility research involving Deaf and Hard of Hearing (DHH) communities. Yet less attention has been given to the assumptions shaping this work. We present a theory-grounded review of XR research involving DHH users. Drawing on Disability Studies, Deaf Studies, and DeafSpace, we develop a theoretical framework with four analytical dimensions: orientation toward access, distribution of responsibility, conceptualization of DHH communities, and spatial and perceptual assumptions. We apply this framework to 53 XR studies involving DHH users, identified through a search and screening of ACM publications from 2015 to 2025. Our analysis shows that XR accessibility is frequently framed as supporting communication within hearing-default environments, while overlooking the diversity of DHH communities. We identify directions for redistributing accessibility labor and reconfiguring space, and offer our framework as a tool for future XR research involving DHH communities.
Group-based ethics training for Science, Technology, Engineering and Mathematics (STEM) students is a complex challenge, requiring substantial resources and expertise. While activity-based teaching methods, such as role-playing and discussions, are commonly employed to simulate real-world scenarios, current practices are often manual and lack integration with effective online platforms for supporting group-based ethical discussions. In this work, we propose Ethics Training Agents, a group discussion system that leverages multiple LLM participants embodying distinct ethical orientations, along with a moderator agent, to enable structured human-AI group ethical discussions for collaborative reflection. We conduct a user study with 45 undergraduate STEM students to evaluate the learning outcomes and user experience. The results show that our system supports engagement, coordination, and perspective-taking in group discussions and has a positive influence on ethical sensitivity. We also discuss practical design strategies for integrating multiple LLM agents into multi-human group settings to facilitate ethics training for STEM students.
This study examines data visualization design evolution over 12.5 years, reflecting on the impact of Large Language Models over the last 3.75 years. Using a longitudinal corpus of 55 visualizations from a single-subject design record, the study identifies how LLMs have aided design-space exploration: reducing coding effort, enabling new design opportunities, shock, excitement, accomplishments, and shifts to the design process.
Live captions on TV often contain errors and timing issues, making it hard for deaf and hard-of-hearing (DHH) viewers to follow dialog. It is essential that caption quality metrics reflect the lived DHH TV viewing experience. To this end, we describe a U.S.-based large-scale online survey with 216 validated participants, who provided 302 responses containing a cumulative 4,832 data points. Participants viewed videos drawn from a pool of 70 clips recorded from live TV, and were asked to rate the caption quality and subjective understanding of the content across four conditions: TV captions as originally recorded with up to 7-12 seconds delay, TV captions synchronized with audio, Automatic Speech Recognition (ASR)-generated captions synchronized with audio, and ASR captions with an average two-second delay. All captions were evaluated against the Word Error Rate (WER), Automated Caption Evaluation (ACE2) and Number, Edition and Recognition (NER) metrics. Results show that TV and ASR captions were rated similarly. For TV captions, all three metrics were moderately-to-highly correlated with viewer ratings, but far less so for ASR captions, making them far from technology-neutral. Additionally, caption latencies significantly impact the viewer experience, especially typical 7-12-second TV delays. We discuss the implications for the adoption of caption quality metrics.
Sports data analysts translate domain questions into insights by combining computation with sport-specific domain expertise. Large language models ease programming, but prompt-to-report workflows may obscure decisions and evidence. We present AI Soccer Analyst, a mixed-initiative system with revisable stages: Data Understanding, Problem Definition, Structured Planning, Execution, Evidence-Grounded Reporting, and Interaction and Refinement. A formative study with five analysts first informed design goals for automation, verifiability, human control, and accessibility. Subsequently, a task-based evaluation with 16 participants combined system logs, retained artifacts, ratings, and open responses; 33 of 48 tasks met the operational completion criteria. Exploratory tests supported favorable participant perceptions of completed-task output quality, task achievement, reliability, and verifiability after Holm correction. Interaction records showed domain knowledge emerging through clarification, planning, and refinement. These findings position stage-aware human-AI collaboration as a practical approach for producing inspectable, revisable, and verifiable analyses while retaining domain-expert involvement in consequential decisions.
Providing guidance is frequently referenced as a key capability of digital health interventions targeting physical activity, yet the term remains poorly defined and inconsistently applied. Existing work often conflates guidance with related constructs such as personalisation, feedback, or persuasion, limiting both theoretical clarity and design progress. This paper conceptualises an initial design space of guidance in the context of digital physical activity support. We define guidance for physical activity as situated, action-oriented support that scaffolds users' embodied engagement in physical activity. Drawing on literature from behaviour change, human-computer interaction, embodied cognition, and digital health, we outline a design space that characterises guidance along multiple dimensions: scope, purpose, timing, context, modality, embodiment, adaptivity, autonomy, and affective quality. By offering a structured vocabulary and conceptual foundation, this work aims to support more coherent research, comparisons, and responsible design of digital health interventions featuring guidance for physical activity support.
The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.
Multi-display signage (MDS), now ubiquitous in urban environments, has the potential to influence human behavior and experience in public spaces. However, despite its unique capability to present spatially distributed dynamic visual stimuli, its current use is mainly limited to advertising. In this study, we propose a perception-based approach for laterally modulating pedestrian trajectories as a nonverbal means of guiding pedestrians in public spaces. The approach is motivated by vection, the illusion of self-motion, and uses laterally moving monochrome stripes, a standard stimulus in vection research, presented across spatially distributed displays to elicit postural responses that may bias pedestrian trajectories. We evaluated the approach through a controlled laboratory experiment and a real-world field deployment involving actual pedestrian flows in a national museum. The laboratory experiment examined whether the MDS setup induced trajectory shifts in the direction predicted by prior research on the behavioral effects of vection. The field deployment investigated whether comparable effects would emerge in aggregate pedestrian behavior during unconstrained movement under conditions closer to those of urban public spaces. In the laboratory, full-screen motion significantly biased walking trajectories in the direction of visual motion, whereas partial-stripe motion produced no significant directional effect. In the field deployment, opposing full-screen motion conditions produced direction-consistent differences in aggregate pedestrian positions. The field results, observed despite the substantial variability in real-world pedestrian flows, extend the controlled laboratory findings and provide ecologically valid evidence supporting practical MDS-based pedestrian modulation in public settings. The results further suggest that sufficient visual-motion coverage may be important.
We present X-Hinges, a design and fabrication method for self-sensing compliant mechanisms based on multi-material FDM 3D printing. By co-printing two conductive filaments of different conductivities within a compliant body, we embed resistive sensing elements directly during fabrication without post-assembly, enabling continuous motion sensing across multiple degrees of freedom in a single print. The structure supports three degrees of freedom, each equipped with a dedicated sensing element configuration for multi-DOF motion estimation. We develop a precision data acquisition system and data-driven regression models that enable continuous, real-time motion sensing. We also introduce an interactive design tool for customizing the geometry, mechanical properties, degrees of freedom, and sensing configurations of X-Hinges. The tool also supports augmenting existing 3D models with self-sensing structures, endowing ordinary objects with continuous multi-DOF sensing capabilities. Finally, we present a set of application examples demonstrating the capability of X-Hinges for fabricating personalized interactive interfaces.
Live programming provides visibility to programmers by running and tracing programs as they are edited. However, for programs with potentially harmful side effects, liveness can turn mistakes into disasters. We propose enabling live programming in environments with side effects via sandboxing: confining effects to a simulation of the true environment. We apply sandboxed live programming in the challenging context of shell scripting: a ubiquitous and powerful---yet notoriously opaque and error-prone---tool. ShellVis provides line-by-line feedback on a shell script's run-time behavior, with file operations sandboxed via a safe overlay of the file system. A qualitative user evaluation finds ShellVis to be helpful to participants, replacing tedious existing practices and instilling confidence. Participant responses also reveal areas for future research, particularly bridging the gulf of execution alongside the gulf of evaluation. ShellVis serves as a case study of how sandboxing can bring live-programming techniques into the many real-world programming contexts where side effects are important.
Blind or low vision (BLV) students are less likely to graduate from college, particularly in computing. Prior work documents accessibility challenges in high school and college, but we lack understanding of the transition process that produces this "leaky pipeline." To address this, we interviewed ten BLV college students about going to college to study computing. We analyzed our data through the lens of life transition, specifically Intersecting Liminality. Our findings reveal that some BLV students face such immense digital accessibility and college acclimation barriers that the only way forward as coders is to take on a "second job" as a blind advocate or drop out of the computing major. We argue that the college transition is a critical point for analysis and technological intervention, and further, that Intersecting Liminality provides a useful lens for HCI scholars to unpack the compounding challenges that prevent some BLV students from completing computing degrees.
Generative artificial intelligence (GenAI) significantly expands creators' productive capacity, but this does not necessarily entail a corresponding increase in creative agency or authorship. This paper distinguishes creativity at the level of the work from creative agency at the level of the creator, and argues that human authorship cannot be determined solely by manual intervention, degree of automation, the origin of an initial idea, or final selection authority. Rather, authorship depends on whether human judgment and reasons genuinely shape the development of the work. To articulate this requirement, the paper introduces Meaningful Human Control (MHC) into generative creation and identifies a limitation of its classical tracking condition. Creative reasons are not always fully specified prior to interaction with AI; they may emerge, change, or be abandoned as the creative process unfolds. The paper therefore proposes dynamic-reflexive tracking (DRT), which requires that a creator's evolving reasons undergo reflective uptake, exert genuine influence on the subsequent trajectory of creation, and remain capable of rejecting and redirecting the system's default direction. DRT consists of four conditions: diachronic reason formation, reflective uptake, trajectory efficacy, and contestability and redirection, together with a minimal tracing requirement. The paper argues that human authorship under generative AI depends not on how many steps a person personally performs, but on whether that person's reasons continuously, reflectively, and effectively shape what the work becomes.
Long-video understanding on edge devices must reason over hours of content under tight compute and bandwidth budgets. Subsampling visual tokens loses temporal structure, while text-only video memories lose fine-grained visual attributes. We observe a visual-textual duality: language memories carry long-range temporal structure better than dense frames, while pixels remain decisive for attribute-level perception. Building on this insight, we propose Caption-once, Frames-onDemand (CFD), a budget-aware edge-cloud agentic framework. The edge runs a single offline captioning pass that builds a dual-track narrative index, an event-level story skeleton plus a clip-level micro-log, cached and reused across queries without re-captioning. At query time, a cloud-side MLLM reasons over the index in a story-first loop centered on a lightweight Visual-Need Router: a per-query gating module that triggers bounded keyframe retrieval only for perceptual questions (appearance, on-screen text, attribute disambiguation) and keeps temporal-structural questions in language space. The router turns visual access into a first-class, query-conditioned cost, capping per-query frame consumption regardless of video length. Experiments on long-video benchmarks demonstrate strong accuracy-efficiency trade-offs while substantially reducing online visual processing.
Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (hanab.live), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, $-$0.7~pp in AI pairs, and +16.4~pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46~pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38--41\%), but human failure rates ranged from 14.4\% to 34.4\% and the gap from +24.1 to +6.2~pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6~pp at the convention-free level, rising monotonically to +21.7~pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.
The development of autonomous driving demands comprehensive testing in mixed-traffic scenarios involving vulnerable road users (VRUs), where purely artificial agents often fail to capture authentic human social negotiations. While human-in-the-loop (HITL) simulators enable safe investigation of these interactions, existing multi-agent platforms struggle with the network latency and synchronization constraints required for high-fidelity haptic feedback. To resolve this, we present CARLAverse, an open-source, multimodal simulation ecosystem. Extending modular hardware abstraction, CARLAverse integrates driving (DrivoCARLA), cycling (CycloCARLA), and pedestrian (WalkoCARLA) simulators into a shared virtual environment. Its core methodological contribution is a distributed physics architecture: latency-critical ego dynamics and high-frequency force feedback are computed locally on client nodes, while a central CARLA server orchestrates non-player character (NPC) physics and global traffic. By decoupling haptic control loops from network bottlenecks, CARLAverse enables scalable, cross-institutional HITL experiments without compromising physical immersion. Code and documentation: https://git.ieem-ka.de/simulator-environments/carlaverse
This paper presents SurgicalRoomAgent, a voice-interactive multi-agent system for smart operating rooms based on large language models (LLMs). The system achieves natural language understanding, device control, intraoperative recording, and surgical report generation through a layered architecture comprising a voice interaction pipeline (wake, ASR, turn detection, agent reasoning, TTS) and an agent core (skill registry, task planner, device manager). Three key technologies are investigated: (1) KV Cache prefix warming for low-latency inference, reducing recomputation overhead from approximately 500 ms to tens of milliseconds via byte-level Longest Common Prefix reuse; (2) streaming partial JSON parsing with early parallel task execution, reducing end-to-end latency by approximately 30%; and (3) progressive skill prompt disclosure, which dynamically filters system prompts based on user role, connected devices, and surgical phase to maximize information density within limited context windows. The system is implemented using the Qwen3-27B model with llama.cpp/sglang inference engines. Experimental analysis demonstrates effective operation within a 16,384-token context limit and multi-device parallel control response times meeting OR real-time requirements.
Generative AI and the practice of "vibe coding" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?
Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.
Digital technologies are deployed in urban built environments with the aim of supporting social dimensions. However, the research offers no unified guidance for such digital technology design. Additionally, evidence shows that week social wellbeing contributes to mental and physical health outcomes which suggests that the technology designed to strengthen social wellbeing could also function as a form of health promoting and preventive intervention. This chapter addresses this research gap by developing a conceptual framework that supports the design of digital technologies for social wellbeing in built environments. We propose a conceptual framework composed of three core components which are drawn on the synthesis of selected empirical studies on technologies embedded in built environments for social wellbeing. First, a social wellbeing dimensions model that identifies what digital technology could address. Second, a digital technology contribution matrix that distinguishes the types of contributions a digital technology could make. Third, levels that maps the scope at which technology could support social wellbeing. This conceptual framework could help researchers, practitioners and policymakers to design and guide digital technology interventions that target social wellbeing in the built environment.
Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.
The pulling illusion induced by asymmetric vibration stimuli has attracted attention for its potential applications in rehabilitation and sensory assessment. However, the underlying mechanism of the pulling illusion remains unclear. This study addressed the central question of whether peripheral vibrotactile sensitivity alone is sufficient for the illusion to emerge or whether processing beyond basic vibration detection is also required. Neurological disorders can involve impairments at different levels of the nervous system, providing an opportunity to examine this question. Accordingly, we evaluated directional discrimination performance for the pulling illusion and fingertip vibration detection thresholds in 25 participants with diverse neurological disorders affecting different levels of the nervous system, from peripheral to central. Clustering analysis identified contrasting profiles, with high directional discrimination performance despite elevated vibration detection thresholds and chance-level performance despite relatively low-to-intermediate thresholds. In the generalized linear mixed model, motor-related signs, including hemiplegia and tremor, showed a robust negative association with directional discrimination performance, whereas vibration detection threshold was not robustly associated with performance. Furthermore, in participants with hemiplegia, directional discrimination performance was around chance level on the affected side and close to 100% on the unaffected side, despite stimulus amplitudes well above the measured vibration detection thresholds on both sides. Collectively, these findings suggest that the pulling illusion depends on perceptual processing beyond basic vibration detection, through which asymmetric vibration is experienced as directional pulling.