A longitudinal study of Microsoft 365 Copilot adoption at a state Department of Transportation found that perceived usefulness of the tool declined significantly after an eight-week pilot, while perceived ease of use, behavioral intention, and trust showed only minor, nonsignificant changes. The research identified three baseline user personas—Skeptics, Cautiously Positive users, and Champions—with substantial individual movement between groups, including 68% of Champions shifting to less enthusiastic personas. The study also observed that while some concerns (accuracy, privacy) decreased, worries about job impact and required skills increased.
Why it matters: This study provides empirical evidence that employee enthusiasm for generative AI tools can wane after real-world use, highlighting the need for ongoing expectation management and tailored support in enterprise AI deployments.
A new preprint audits nine leading large language models (LLMs) for asset-specific preferences and finds that Bitcoin's ranking among money-like instruments is highly dependent on the scenario presented, with models ranking it higher in crisis and autonomous-agent contexts. The study identifies a dominant internal feature in Gemma 3 that selectively represents Bitcoin; manipulating this feature can causally shift the model's portfolio allocation toward or away from Bitcoin by up to 5.2 percentage points. This demonstrates that internal model representations can be both audited and causally linked to real financial decisions.
Why it matters: This work provides a concrete method for auditing and influencing asset-specific preferences in financial LLMs, laying groundwork for transparency and accountability as such models are deployed in real-world financial decision-making.
A new preprint finds that large language models (LLMs) generate responses that are more concentrated and mainstream than the diverse, long-tail outputs produced by humans. The researchers tested interventions such as increasing temperature sampling, prompting for diverse perspectives, and aggregating outputs from multiple models, finding that these methods can improve diversity, but single-model outputs still fall short of human-level diversity.
Why it matters: The study raises concerns that LLMs could reduce cultural diversity in generated content, which has implications for AI policy and the preservation of democratic values.
Policy & Safety→Official→arXiv Computers and Society
A new benchmark, NOHARM, evaluates 20 large language models (LLMs) and 4 clinical AI tools on 1,100 medical consultation cases, finding that direct use of AI-generated recommendations could result in severe harm in up to 24.6% of cases, with omission errors accounting for over 80% of severe errors. In a randomized study of 101 physicians, AI assistance improved performance, but physicians often omitted valuable AI recommendations, indicating complementary strengths in human-AI teaming.
Why it matters: This study provides the first systematic measurement of clinical safety in LLM-generated medical advice, revealing that widely used AI tools can produce potentially harmful recommendations and highlighting the need for explicit safety evaluation.
Researchers introduce EG-VAR, a Lean 4-based tool-calling architecture that ensures every verified output is grounded in attested tool calls and kernel-checked inference, thereby eliminating unsupported claims. On a subset of TableBench numerical reasoning tasks (n=120), EG-VAR achieves perfect accuracy (120/120) compared to a 95% baseline, and maintains 100% source-faithfulness on stress tests where baselines drop to 80-90%. The system also provides explicit audit trails for abstentions and formalization errors.
Why it matters: EG-VAR offers a practical and auditable approach to eliminating LLM hallucination in empirical inference, potentially transforming trustworthiness in high-stakes AI applications.
A new preprint compares large language models (LLMs) such as GPT, Twitter-roBERTa, and LLaMA to traditional machine learning methods for analyzing open-ended survey responses. The study finds that LLMs achieve higher classification accuracy, especially in sentiment and thematic analysis, but exhibit significant variation in consistency and the explicitness of their reasoning. These results highlight important trade-offs between predictive performance and interpretability in large-scale qualitative research.
Why it matters: The study offers practical insights for researchers seeking to balance automation with interpretive rigor when applying LLMs to qualitative data analysis.
AgentSociety 2 is a new research environment that integrates large language model (LLM) agents as both AI social scientists and simulated participants, automating the end-to-end workflow of social science research. The system enables hypothesis generation, experiment design, simulation execution, result interpretation, and manuscript drafting across micro, meso, and macro social scenarios. It demonstrates the ability to reproduce qualitative patterns from prior studies and supports large-scale, auditable simulations.
Why it matters: This work advances computational social science by providing a scalable, reproducible, and auditable platform for automating complex social experiments with AI agents.
A new preprint examines the use of large language models (LLMs) for content moderation through 'policy-as-prompt' methods, where moderation policies are given to LLMs as natural-language prompts. The authors argue that this approach introduces specific risks and harms, and that simply writing prompts is not sufficient for effective or meaningful community governance. They propose several considerations for improving prompt governance but conclude that prompt-writing alone cannot ensure robust moderation outcomes.
Why it matters: This research highlights important limitations and governance challenges for AI-driven, prompt-based content moderation systems as their use expands in online communities.
A new preprint compares institutional and course-level generative AI policies at U.S. research-intensive universities. The study finds that while institutions tend to be more supportive of GenAI use, course-level guidance in computing education remains cautious. The authors propose an instructor-centered framework to guide future GenAI adoption in courses.
Why it matters: This research highlights a disconnect between university-wide AI policies and classroom practices, offering a framework to help computing educators navigate GenAI adoption.
Policy & Safety→Official→arXiv Computers and Society
AAAI-26 organizers report a significant increase in dual submissions—papers submitted to multiple venues without disclosure—during the conference's review process. By combining similarity assessment, LLM-based overlap tools, and manual review, they desk-rejected 141 main-track submissions. The organizers warn that generative AI may be enabling more sophisticated forms of dual submission and propose several policy and technical recommendations to address the issue.
Why it matters: This development exposes a growing integrity challenge in AI research, with generative AI potentially exacerbating threats to the peer-review process and the reliability of the scientific record.
A new preprint examines how transfer learning can be applied to adaptive multi-agent systems facing policy regime changes. The authors compare blank-slate learners with transfer learners that reuse structural knowledge from previous regimes, using an emissions-regulation simulation. Results show that transfer learning improves performance when the policy-outcome relationship remains stable, but can lead to negative transfer when a regime change introduces a threshold break. The paper provides a methodological framework for determining when regulatory experience should be reused or discarded.
Why it matters: This work offers a formal approach to understanding the risks and benefits of transfer learning in policy modeling for adaptive socio-technical systems, informing the design of AI-driven regulatory tools.
Policy & Safety→Official→arXiv Computers and Society
A new preprint argues that as AI models approach top performance on existing benchmarks, the remaining discriminating items are those requiring elite expert judgment, which is structurally scarce. The authors present a formal model showing how benchmark signal depreciates as models improve, and document a scarcity premium for high-judgment evaluation labor. The paper also discusses the governance implications of these findings for AI capability measurement.
Why it matters: This work highlights a fundamental bottleneck in evaluating advanced AI systems, raising concerns about the reliability of current benchmarks and the challenges for AI governance.
A new preprint explores 'model collapse,' the phenomenon where recursively training AI models on AI-generated data leads to degraded performance, such as repetition and noise. The author argues that while this is typically seen as a technical failure, it also has creative and aesthetic dimensions, drawing parallels to early video feedback art. The paper examines how model collapse challenges certain technological ideals and underscores AI's ongoing reliance on human-generated data.
Why it matters: This research reframes model collapse as not only a technical issue but also a source of artistic and philosophical insight, broadening our understanding of AI's creative potential and limitations.
A new empirical study analyzes 6 million images from the open-source image generation ecosystem, examining how creators use 22,400 base models and 154,000 LoRA models. The research identifies the ecosystem's unique strengths and challenges, offering insights that could inform its sustainability and future innovation. The dataset compiled for the study is publicly available for further research and practical use.
Why it matters: This is the first large-scale empirical analysis of creative workflows in the open-source image generation ecosystem, providing foundational insights for researchers and practitioners.
A large-scale study analyzing over 300 million scientific works across 26 fields finds that the decades-long decline in solo-authored papers halted and partially reversed following the public release of ChatGPT in late 2022. This reversal is most pronounced in fields where coauthors' contributions are more easily replaced by AI, and is observed among both established researchers and newcomers who previously only coauthored papers. The study suggests that generative AI is enabling more researchers to publish solo work, particularly in computationally oriented topics.
Why it matters: This research provides empirical evidence that generative AI is reshaping scientific collaboration by substituting for human labor, altering the traditional division of cognitive work in research.
A new bilingual benchmark study shows that freely accessible large language models (LLMs) fabricate legal citations for Saudi data protection law (PDPL) in 60-77% of cases, while achieving near-perfect accuracy (94-100%) on the EU's GDPR. The research tested 120 questions in both Arabic and English across three models, revealing that fabrication rates are driven by the jurisdiction of the law, not the language of the query. The study also found that high model confidence does not prevent fabricated citations, highlighting a significant reliability gap.
Why it matters: This research highlights a critical jurisdiction-based reliability gap in LLM-generated legal citations, raising concerns for regulatory compliance and legal decision-making.
Researchers present Gauntlet, an open-source pipeline that uses five independent expert-persona LLM reviewers and an adversarial synthesis stage to analyze computer architecture papers. In evaluations on 20 ISCA and HPCA papers, human judges preferred Gauntlet's analyses over those by human experts in 15 out of 20 cases, with statistically significant advantages in critical rigor. Ablation studies show that the multi-agent structure, especially the synthesis stage, is key to Gauntlet's performance gains over single-agent LLM baselines.
Why it matters: This work suggests that structured multi-agent LLM pipelines can exceed human experts in deep technical critique, indicating potential new roles for AI in peer review and research evaluation.
A four-year analysis of undergraduate programming submissions examines how students use natural-language comments to guide AI code generation. The study introduces a taxonomy covering comment type, code expression level, and code construct, and finds that students primarily write 'What' comments but shift to 'How' comments for procedural tasks. Students tend to focus more on verifying generated code than on revising their comments.
Why it matters: This research offers new empirical insights into how students interact with AI code assistants, informing the evolving role of natural language specifications in programming education.
Researchers have introduced DeepBias, an adaptive framework designed to probe social biases in Large Vision-Language Models (LVLMs) more deeply than traditional static datasets allow. DeepBias uses a dynamic loop involving a ProposerAgent that generates test data and a DiggerAgent that iteratively rewrites these tests based on model responses, enabling the exposure of progressively deeper biases. The team also developed DeepBiasBench, a benchmark constructed using an ensemble of five state-of-the-art LVLMs to identify vulnerabilities shared across different architectures.
Why it matters: This work advances LVLM safety assessment by introducing an adaptive, evolutionary approach that reveals deeper and more nuanced model biases than static datasets can uncover.
A new preprint systematically compares four AI agent architectures—monolithic, chain-based, multi-agent, and iterative—across 50 journalism tasks using the same language model and tools. The study finds that architecture explains 82% of the variance in processing behavior. Multi-agent collaboration achieved the highest accuracy (84.7%) but required about twice the time of other designs, while the monolithic architecture exhibited a 71.7% source rejection rate, paralleling classic human gatekeeping.
Why it matters: The findings provide evidence-based guidance for newsrooms on selecting AI architectures based on priorities such as speed, accuracy, or auditability.