What changed in AI — Page 92

ResearchOfficialarXiv Computer Vision

DM-KG: A Structured Knowledge Graph Framework Improves Spatial Reasoning in Vision-Language Models for Street View Imagery

A new framework called DM-KG introduces a structured knowledge graph to enhance spatial reasoning in vision-language models (VLMs) applied to street view imagery. By extracting and encoding directional and metric relationships between entities from 2D images, DM-KG significantly reduces distance estimation and direction judgment errors on spatial question-answering benchmarks.

Why it matters: This work offers a notable advance in improving the spatial cognition of VLMs, addressing a key limitation for their deployment in geospatial and geographic visual question answering tasks.

ResearchOfficialarXiv Computer Vision

Self-Consistent Flow: Unifying Velocity and Endpoint Prediction for Rectified Flow Models

Researchers introduce Self-Consistent Flow (SC-Flow), a method that jointly trains a single network to predict both local velocity and data endpoint in rectified-flow generative models. By adding a lightweight consistency loss, SC-Flow unifies the strengths of both parameterizations, stabilizing training and improving the straightness of generation paths. Experiments on image generation tasks show that SC-Flow achieves notable quality improvements over standard rectified-flow baselines with minimal computational overhead.

Why it matters: This work offers a principled approach to combining two key training targets in rectified flow models, potentially enhancing the stability and quality of generative image models.

ResearchOfficialarXiv Computer Vision

VLM-Assisted Framework Improves EEG-to-Image Reconstruction Evaluation

A new evaluation framework uses four vision-language models (VLMs) to assess EEG-to-image reconstructions, introducing Tolerant Perceptual Alignment Scores (T-PAS) and Tolerant Semantic Alignment Scores (T-SAS). The distilled BCI-Coherence Score (BCS) demonstrates lower mean absolute error and higher correlation with human judgments than traditional pixel or representation-based metrics, addressing the challenge of distinguishing visual fidelity from semantic recoverability.

Why it matters: This framework offers a more meaningful way to evaluate EEG-to-image reconstructions, potentially advancing brain-computer interface research by better capturing both perceptual and semantic accuracy.

ResearchOfficialarXiv Computer Vision

RINO: Unified Vision Model Treats All Visual Tasks as RGB-to-RGB Editing

Researchers introduce RINO (RGB In and RGB Out), a unified vision model that reformulates diverse visual tasks—including segmentation, depth estimation, and pose-to-image generation—as RGB-to-RGB image editing problems. Using a single architecture and shared parameters, RINO achieves robust zero-shot performance across tasks without task-specific fine-tuning. This approach aims to establish a shared visual interface, analogous to how language models operate over text.

Why it matters: RINO demonstrates a significant step toward unified vision models capable of handling multiple tasks with a single architecture, potentially simplifying deployment and advancing general visual understanding.

ModelsOfficialarXiv Computer Vision

SpikeDS: Dual Sparsity Spiking Transformer Improves 3D MRI Analysis for Cancer Prognosis

Researchers have developed SpikeDS, a spiking neural network architecture that predicts perineural invasion (PNI) in cholangiocarcinoma from 3D MRI scans. SpikeDS employs dual sparsity—combining activation sparsity from binary spike communication and spatial sparsity from window pruning—to reduce computational costs while maintaining diagnostic accuracy. In a clinical cohort study, SpikeDS achieved higher accuracy (AUC 0.753) and greater energy efficiency than existing methods.

Why it matters: This work demonstrates a promising approach for efficient and accurate AI-based medical imaging analysis, which could improve cancer prognosis in clinical practice.

ResearchOfficialarXiv Computer Vision

Representation and Reference Selection in Training-Free Synthetic Image Attribution

A new preprint investigates how the choice of representation space and reference construction methods affect the performance of training-free synthetic image attribution. The study finds that attribution accuracy is highest when using intermediate layers of CLIP and DINOv2 models, and that semantically constrained references further improve results, particularly when only a small number of references are available. The analysis highlights the importance of both representation selection and reference construction in building effective attribution systems.

Why it matters: The findings offer practical insights for designing scalable, training-free attribution systems that can adapt to new image generators without retraining.

ResearchOfficialarXiv Computers and Society

LLMs Outperform Traditional ML in Open-Ended Survey Analysis but Face Consistency and Explainability Issues

A new preprint compares large language models (LLMs) such as GPT, Twitter-roBERTa, and LLaMA to traditional machine learning methods for analyzing open-ended survey responses. The study finds that LLMs achieve higher classification accuracy, especially in sentiment and thematic analysis, but exhibit significant variation in consistency and the explicitness of their reasoning. These results highlight important trade-offs between predictive performance and interpretability in large-scale qualitative research.

Why it matters: The study offers practical insights for researchers seeking to balance automation with interpretive rigor when applying LLMs to qualitative data analysis.

ResearchOfficialarXiv Computers and Society

AgentSociety 2: An Integrated Research Environment for Executable Social Science

AgentSociety 2 is a new research environment that integrates large language model (LLM) agents as both AI social scientists and simulated participants, automating the end-to-end workflow of social science research. The system enables hypothesis generation, experiment design, simulation execution, result interpretation, and manuscript drafting across micro, meso, and macro social scenarios. It demonstrates the ability to reproduce qualitative patterns from prior studies and supports large-scale, auditable simulations.

Why it matters: This work advances computational social science by providing a scalable, reproducible, and auditable platform for automating complex social experiments with AI agents.

ResearchOfficialarXiv Computer Vision

GEST-Engine: From Text to Fully-Annotated Synthetic Video via Explicit World Models

The GEST-Engine is a system that generates fully-annotated synthetic multi-actor video from natural language input by maintaining an explicit, inspectable world model represented as a Graph of Events in Space and Time (GEST). It produces frame-aligned RGB video, depth, segmentation, pose, and other annotations at zero marginal annotation cost. The system guarantees object permanence and temporal consistency, making it suitable for generating training data and evaluation benchmarks for video understanding.

Why it matters: GEST-Engine enables scalable production of richly annotated synthetic video with guaranteed consistency, potentially reducing reliance on manual annotation in video research.

ResearchOfficialarXiv Computer Vision

VLM-Based Method Extracts Expert Actions and Decision-Making Scenes from Maintenance Videos

A new method uses vision-language models (VLMs) to detect anomalous frames between maintenance task videos, enabling automatic extraction of expert-specific actions and contextual decision-making scenes. In simulated maintenance experiments, the approach achieved extraction rates of 65% for actions and 61% for decision-making scenes, outperforming conventional methods. The technique leverages frame-wise visual descriptions and intra-video self-similarity to identify key moments of expert know-how.

Why it matters: This method could facilitate the transfer of expert knowledge to less experienced workers by automatically identifying and extracting critical scenes from maintenance videos.

ResearchOfficialarXiv Computers and Society

Policy-as-Prompt Moderation with LLMs: Risks and Governance Considerations

A new preprint examines the use of large language models (LLMs) for content moderation through 'policy-as-prompt' methods, where moderation policies are given to LLMs as natural-language prompts. The authors argue that this approach introduces specific risks and harms, and that simply writing prompts is not sufficient for effective or meaningful community governance. They propose several considerations for improving prompt governance but conclude that prompt-writing alone cannot ensure robust moderation outcomes.

Why it matters: This research highlights important limitations and governance challenges for AI-driven, prompt-based content moderation systems as their use expands in online communities.

ResearchOfficialarXiv Computers and Society

Study Finds Gap Between Institutional and Course-Level GenAI Policies in Computing Education

A new preprint compares institutional and course-level generative AI policies at U.S. research-intensive universities. The study finds that while institutions tend to be more supportive of GenAI use, course-level guidance in computing education remains cautious. The authors propose an instructor-centered framework to guide future GenAI adoption in courses.

Why it matters: This research highlights a disconnect between university-wide AI policies and classroom practices, offering a framework to help computing educators navigate GenAI adoption.

ResearchOfficialarXiv Computer Vision

DeGuNet: Depth-Guided Ultra-Compact Backbones for Efficient LiDAR-Camera 3D Detection

A new preprint introduces DeGuNet, an ultra-compact image backbone designed for LiDAR-camera 3D detection in autonomous driving. DeGuNet uses depth-guided representation learning to address parameter redundancy in current multi-modal frameworks, reducing GPU memory usage by up to 66.5% and achieving a 1.16x speedup. The method also improves mean average precision (mAP) by up to 6.20 points on the nuScenes dataset, demonstrating both efficiency and accuracy gains.

Why it matters: DeGuNet could enable more efficient and accurate 3D perception for autonomous vehicles by significantly reducing computational demands.

Policy & SafetyOfficialarXiv Computers and Society

AAAI-26 Desk-Rejects 141 Papers Amid Surge in Undisclosed Dual Submissions

AAAI-26 organizers report a significant increase in dual submissions—papers submitted to multiple venues without disclosure—during the conference's review process. By combining similarity assessment, LLM-based overlap tools, and manual review, they desk-rejected 141 main-track submissions. The organizers warn that generative AI may be enabling more sophisticated forms of dual submission and propose several policy and technical recommendations to address the issue.

Why it matters: This development exposes a growing integrity challenge in AI research, with generative AI potentially exacerbating threats to the peer-review process and the reliability of the scientific record.

ResearchOfficialarXiv Computer Vision

Continual Learning for Heterogeneous Medical VQA: An Empirical Analysis

A new preprint systematically evaluates continual learning (CL) methods for medical visual question answering (MedVQA) across a range of clinical tasks, including classification, detection, cell counting, and report generation. The study finds that current CL methods have difficulty maintaining a balance between retaining old knowledge and learning new tasks when faced with diverse objectives and supervision formats. The authors also analyze the impact of task ordering and the evolution of model parameters during continual learning. Code and experimental setup will be made publicly available.

Why it matters: This work exposes key limitations in current continual learning approaches for medical VQA, highlighting challenges that must be addressed for robust real-world deployment.

ResearchOfficialarXiv Computer Vision

SymbOmni: Agentic Omni Model with Symbolic Concept Learning for Cumulative Evolution

A new preprint introduces SymbOmni, an agentic omni-model designed to address the 'perpetual novice' problem in visual generation by leveraging Symbolic Concept Learning. The model features a Symbolic Concept Box that abstracts experiences into reusable instructions, enabling cumulative learning and compositional generalization. Experimental results show that SymbOmni outperforms existing agent-based and closed-source systems in image quality and task success, reduces token consumption by over 40%, and achieves state-of-the-art continual learning performance.

Why it matters: This work presents a novel approach for enabling AI models to learn cumulatively and evolve autonomously, potentially overcoming a key limitation of current monolithic models.

ResearchOfficialarXiv Cryptography and Security

AutoTrace: Agentic Pipeline Localizes Vulnerability Triggers Beyond Patched Functions

AutoTrace is an agentic pipeline that uses LLM agents to explore code property graphs and localize vulnerability triggers, even when they are located outside patched functions. On the InterPVD benchmark, AutoTrace achieves 75.0% VulnHit and 80.8% FuncHit, surpassing previous methods. The authors also introduce SinkTrace-Bench, a dataset of 1,542 source-to-sink causal chains, which reveals that current frontier LLMs struggle with causal reasoning in vulnerability analysis.

Why it matters: This work advances automated vulnerability analysis by enabling interprocedural trigger localization and provides a new benchmark that highlights the causal reasoning limitations of current LLMs.

ResearchOfficialarXiv Cryptography and Security

StableAML: Machine Learning for Behavioral Wallet Detection in Stablecoin Anti-Money Laundering on Ethereum

A new preprint introduces StableAML, a machine learning framework for detecting money laundering in stablecoin transactions on Ethereum. The study finds that domain-informed tree ensemble models outperform graph neural networks in identifying suspicious wallets, and can distinguish between behavioral patterns of cybercrime syndicates and sanctioned entities. The approach is designed to support compliance with emerging regulations such as the EU's MiCA and the U.S. GENIUS Act.

Why it matters: This work presents a novel, interpretable method for high-precision behavioral classification in stablecoin anti-money laundering, potentially improving compliance and reducing unjustified asset freezes under new regulatory frameworks.

ResearchOfficialarXiv Cryptography and Security

Agent Identity URI Scheme Enables Decentralized, Topology-Independent Naming and Discovery for Multi-Agent Systems

A new preprint introduces the agent:// URI scheme, designed to decouple agent identity from network location in multi-agent systems. The scheme incorporates trust roots, hierarchical capability paths, and sortable unique identifiers, with cryptographic attestation using PASETO tokens. Evaluation demonstrates full capability expressiveness on 369 tools, perfect discovery precision across 10,000 agents, and sub-5-microsecond performance, suggesting a robust and scalable approach to agent identity and discovery.

Why it matters: This work proposes a novel, decentralized solution to agent identity and discovery, potentially enabling more robust and scalable multi-agent ecosystems.

Policy & SafetyOfficialarXiv Cryptography and Security

Large-Scale Study Reveals LLM Agents Hallucinate Skill Names, Enabling Supply-Chain Attacks

A large-scale study analyzing 15,000 prompts across 12 LLM and agent configurations found that hallucination of skill names is widespread, with rates averaging 36.0% for standalone LLMs and 36.9% for agents, and rising to 43.1% on real-world developer questions. These hallucinated names can be exploited by adversaries who pre-register malicious skills, enabling supply-chain attacks. The study evaluated four defenses and found that the strongest, retrieval grounding, reduced hallucination to 3.2% but significantly reduced the system's usefulness, with correct skill recommendations dropping to about one in six.

Why it matters: This vulnerability exposes LLM agent ecosystems to easy supply-chain attacks, and current defenses severely compromise usability, highlighting the need for structural changes to registries and recommendation pipelines.