A new preprint introduces a white-box instrument based on hidden deterministic finite automata to separately measure reward attainment and latent-state learning in reinforcement learning (RL) agents. The study demonstrates that high reward does not necessarily indicate that an agent has learned the underlying task state, distinguishing between 'perception gaps' (where latent state is not recoverable from observations) and 'planning gaps' (where state is recoverable but not used). The authors show that optimizer strength, task structure, and observation informativeness each influence the relationship between reward and state learning.
Why it matters: This work challenges the common assumption that high reward in RL implies genuine task understanding, providing a method to diagnose when agents exploit shortcuts rather than learning true latent states.
CARE-LoRA is a method designed to reduce the memory bottleneck of activations during LoRA fine-tuning by replacing full input activations with low-rank compressed activations from the LoRA branch. It introduces a lightweight reconstruction matrix computed during the forward pass, enabling gradient reconstruction in backpropagation while keeping LoRA matrices fully trainable. Experiments across various models and tasks show that CARE-LoRA achieves competitive or superior performance compared to standard LoRA and its variants, with a substantially reduced memory footprint.
Why it matters: This approach addresses a major memory limitation in fine-tuning large models, potentially making such processes more feasible on hardware with limited memory.
Researchers present Support Vector Attention (SV-Attention), a novel max-margin memory mechanism for attention models that enables certified token eviction and exact unlearning. SV-Attention achieves higher rare-item recall (0.86 vs. 0.32) and better retention (0.80 vs. 0.05 deterioration hours) on MIMIC-IV data streams compared to baselines, and demonstrates improved compression (2.178 BPC vs. 2.383) on enwik8 relative to a sliding-window Transformer. The method also supports surgical forgetting, exact editing, and patient-record deletion in real-world scenarios.
Why it matters: This work introduces a principled and practical approach to certified unlearning in attention mechanisms, addressing a key challenge in machine learning privacy and data management.
A new training-free attribution method is introduced to analyze the dependencies of feedforward network (FFN) neurons in Transformer models. The study finds that small, sparse subsets of upstream activations and attention outputs can preserve neuron activations with high fidelity, even when other inputs are masked. This reveals that FFNs, despite their dense parameterization, have sparse and structured inter-layer dependencies at the neuron level. The method is scalable and can be used for circuit-level interpretability and identifying sparse pathways for potential efficiency gains.
Why it matters: This work offers a practical and scalable approach to understanding and potentially optimizing the internal structure of Transformer models, which could lead to more efficient inference and improved interpretability.
Researchers introduce the first pseudo-polynomial-time exact algorithm for computing Data Shapley values in weighted k-nearest-neighbor (KNN) regression, addressing a longstanding computational barrier. The work also presents a certified fully polynomial-time approximation scheme (FPTAS) with machine-checkable error bounds and extends the approach to soft-label multi-class prediction. An open-source implementation and the first exact ground-truth dataset for weighted-regression Data Shapley are provided.
Why it matters: This advance enables deterministic, certified data valuation in weighted KNN regression, providing a reliable reference for auditing and improving approximate methods.
A new preprint introduces an agentic AI scientific community composed of virtual laboratories that autonomously discover neural operator architectures. Each lab uses LLM agents for planning, training, and peer review, operating under a citation-based economy. The system demonstrates the ability to find high-accuracy, low-parameter-count neural operator designs across five PDE benchmark problems, outperforming rule-based alternatives in maintaining architectural diversity.
Why it matters: This work represents a significant advance in automating scientific discovery using LLM-driven agentic systems, highlighting the potential for AI to autonomously innovate in complex domains.
Airbnb has developed and deployed Proximity Features, a privacy-compliant system that addresses the cold-start personalization problem by grouping users based on geographic proximity using geo-IP data and adaptive clustering. This approach aggregates signals for groups of around 1,000 users, avoiding the need for persistent individual identifiers and operating within consent-gated privacy controls. Online A/B tests in production show statistically significant increases in bookings, particularly for users lacking recent or any history.
Why it matters: This work provides a scalable, privacy-preserving solution to cold-start personalization with demonstrated real-world impact at a major online platform.
Researchers introduce LiteTopK, a fused GPU kernel designed to accelerate Indexer-TopK operations in long-context sparse attention by leveraging the curse of dimensionality. LiteTopK samples data to estimate score ranges and partitions candidates into bins, reducing memory traffic and overhead while maintaining exact top-k correctness. Experiments show a 1.2x speedup in the prefill stage of GLM 5.2 during real-world deployment, along with lower memory usage.
Why it matters: This work offers a practical advance in the efficiency of sparse attention for large language models, enabling faster and more memory-efficient processing of long contexts.
A new preprint demonstrates that up to 77% of features recovered by sparse autoencoders (SAEs) with high cosine similarity to ground-truth directions are causally inert, meaning the matched atom does not activate when the feature is present. The authors introduce sae-causal-audit, a model-agnostic tool for causal validation, and identify two types of inertness: structural (due to antipodal-pair geometry) and competitive (arising from TopK pathologies in degraded SAEs).
Why it matters: This work challenges the use of correlational metrics alone for evaluating SAE interpretability, showing that high recovery scores can mask a lack of causal relevance, which is important for mechanistic interpretability research.
A new audit of six KV-cache compression methods finds that their performance rankings reverse when compression is performed before seeing the query (query-agnostic) compared to after (query-aware). In the more realistic query-agnostic setting, only KeyDiff consistently outperforms trivial baselines, while the widely used SnapKV method underperforms a simple 'keep start and recent window' baseline.
Why it matters: This study challenges the validity of current evaluation protocols for KV-cache compression, suggesting that some popular methods may be less effective in real-world LLM deployments than previously thought.
A new method equips speculative execution in LLM agents with three online memory systems—a contrastive transition table, episodic memory, and confusion tracker—to improve prediction quality by learning from past agent trajectories. The approach achieves 19–39% relative accuracy improvement on action prediction and up to 2.5× increase on observation prediction tasks, with gains increasing as memory accumulates. All speculation occurs during idle time, resulting in no added wall-clock cost and preserving identical agent trajectories compared to non-speculative execution.
Why it matters: This work demonstrates a way for LLM agents to learn from experience and improve efficiency without sacrificing correctness, potentially advancing the deployment of faster and more capable agentic systems.
A new framework called PFAdapter introduces hierarchical LoRA decomposition to separate global-shared and local-private parameters for federated fine-tuning of multimodal large language models (MLLMs). By synchronizing only the global-shared components and keeping local adaptations private, PFAdapter reduces communication costs by nearly 50% and achieves accuracy improvements of 2.4% to 4.8% on several medical and multimodal datasets. The approach also uses orthogonality regularization to enforce strict separation between parameter types, preventing redundant feature learning.
Why it matters: This work offers a practical advance for deploying personalized, communication-efficient AI at network edges, addressing key challenges in federated learning for multimodal models.
Researchers introduce FastAlign, a sparsity-aware framework for optimal transport-based network alignment. FastAlign maintains alignment quality comparable to state-of-the-art methods while significantly improving computational efficiency, achieving up to 9.45x speedup on CPU and 32.54x on GPU. The method leverages mixed sparse-dense operations and custom kernel fusion to address scalability challenges in large-scale network alignment.
Why it matters: This work enables efficient analysis of large-scale networks, addressing a major scalability bottleneck in network alignment tasks relevant to fields such as social network analysis and fraud detection.
Researchers have introduced an explicit multimodal routing framework for clinical prediction using electronic health record (EHR) data, allowing for interpretable reasoning across structured variables, clinical notes, and chest X-rays. The model constructs discrete unimodal, bimodal, and trimodal routes, and uses inference-time route masking to audit the contribution of each modality and assess robustness without retraining. Evaluated on phenotype and mortality prediction tasks with MIMIC-IV data, the framework reveals systematic differences in modality reliance across clinical conditions.
Why it matters: This work advances interpretability and robustness in AI-driven clinical decision support by providing a transparent method to understand how different data modalities influence predictions.
A new preprint reports that, contrary to conventional wisdom, data imbalance can actually promote robust generalization in sufficiently capable models when spurious correlations are present. In a synthetic task, a 2-layer transformer achieved 100% adversarial accuracy in 77% of training runs at a high spurious ratio (0.90), compared to 0% at a balanced ratio (0.50). This effect was not observed in 1-layer models, where data imbalance led models to rely on the shortcut feature instead.
Why it matters: This finding challenges standard assumptions about data balance and suggests new strategies for training models to resist spurious correlations.
A new preprint introduces SALT-GNN, a lightweight graph neural network architecture that combines degree-aware statistical aggregation with attention mechanisms to improve anti-money laundering (AML) detection, particularly in dense transaction neighborhoods. The authors show that SALT-GNN uses up to 77% fewer parameters than task-specific graph-transformer baselines and improves dense-context F1 scores by 3-6 points on two datasets, and by 16-20 points on a third dataset for highest-degree nodes. The improvements are consistent across both Transformer- and GAT-style attention mechanisms.
Why it matters: This work addresses a key operational challenge in AML detection—reduced model performance on high-activity accounts—by proposing an efficient architectural modification that could enhance detection accuracy and reduce investigation costs.
A new framework called GAE (Graph-Augmented Evolution) integrates graph neural networks, reinforcement learning, and online fine-tuning to enhance large language model (LLM)-guided evolutionary program search. In experiments on symbolic regression for nonlinear oscillators, GAE efficiently discovers closed-form equations and achieves state-of-the-art out-of-distribution performance compared to static LLM-driven baselines.
Why it matters: GAE offers a significant advance in automated scientific discovery by enabling more directed and adaptive search, potentially accelerating the discovery of physical equations and other scientific insights.
TabLoRA is a parameter-efficient neural ensemble method designed for large-scale tabular learning. By sharing a common backbone across predictors and introducing predictor-specific low-rank adaptations, TabLoRA enables ensemble-style prediction without duplicating all parameters. Benchmark results show that TabLoRA achieves a favorable balance between predictive performance and efficiency compared to gradient-boosted decision trees (GBDTs) and recent deep learning baselines.
Why it matters: TabLoRA offers a practical approach to neural ensemble learning for large-scale tabular data, potentially challenging the dominance of GBDTs in this area.
Researchers have developed Vilya-1, a deep learning model that predicts the structures and properties of macrocycles using an all-atom representation. Vilya-1 demonstrates improved geometric accuracy over existing computational methods and supports the generative design of novel macrocycles across diverse chemical classes.
Why it matters: Vilya-1 could significantly accelerate the development of macrocycle-based therapeutics by enabling more accurate and generalizable structure prediction and design.
FlashTrie is a system that optimizes constrained beam search for generative retrieval tasks on GPUs, featuring a succinct trie layout and a cooperative CUDA kernel to perform decoding entirely on-device. On a library of 800 million keywords with beam widths up to 1000, it reduces trie-search latency to under 3 ms and achieves up to 24x speedup over a highly optimized multi-threaded CPU baseline. In a large-scale online A/B experiment on a commercial search engine, FlashTrie delivered a statistically significant +0.71% revenue lift.
Why it matters: FlashTrie enables real-time, large-scale constrained decoding for generative retrieval, directly improving commercial search engine performance and revenue.