What changed in AI — Page 51

Companies & FundingReportedThe New York Times / AI

American A.I. Giants Like Alphabet Face Fresh Tests

Rapid advancements in Chinese artificial intelligence models are raising questions about costly technology spending as Google’s parent, Alphabet, prepares to report earnings. These developments point to intensifying competition between U.S. and Chinese AI firms.

Why it matters: This highlights the growing competitive pressure on American AI leaders from Chinese advancements, which could influence investment decisions and market dynamics.

ModelsReportedThe Verge / AI

China's Moonshot and Alibaba unveil AI models challenging US leaders

Chinese companies Moonshot and Alibaba have introduced new AI models, asserting that their performance rivals leading systems from OpenAI and Anthropic while operating at a lower cost. These swift developments indicate that the gap between US and Chinese AI capabilities may be narrowing.

Why it matters: This development highlights intensifying competition in advanced AI between China and the US, with potential implications for global AI leadership and market dynamics.

Policy & SafetyReportedThe Register / AI & ML

Auditors tell UK government to do the math before banking on £45B AI savings

UK government auditors have warned that departments have not assessed how AI will reshape staffing, roles, and skills across the public sector, casting doubt on projected £45 billion savings. The report urges more rigorous analysis before relying on such figures.

Why it matters: This highlights a critical gap between AI adoption promises and the practical workforce planning needed to realize them in the public sector.

Companies & FundingReportedThe Guardian / AI

Jeff Bezos and UK government invest in £2bn British startup CuspAI

Jeff Bezos and the UK government have invested in CuspAI, a Cambridge-based AI startup valued at $2.6bn. The company aims to develop AI software that accelerates research and reduces the use of rare metals in chipmakers’ supply chains.

Why it matters: This investment underscores increasing interest from both government and private sector leaders in leveraging AI to address critical material supply chain challenges.

Companies & FundingReportedThe Decoder

Moonshot pauses new Kimi K3 subscriptions after GPU demand maxes out in 48 hours

Moonshot has temporarily halted new subscriptions for its Kimi K3 model after demand nearly maxed out its GPU capacity within 48 hours. The company intends to split its subscription model to distribute computing power more evenly.

Why it matters: This underscores the high demand for advanced AI models and the infrastructure challenges in scaling GPU resources.

ResearchOfficialarXiv Statistical ML

Improving Backward Conformal Prediction via Non-Conformity Score Transformation

A new method, ST-BCP, introduces a data-dependent transformation of non-conformity scores to address the coverage gap in Backward Conformal Prediction (BCP). The approach is theoretically justified and, in experiments on common benchmarks, reduces the average coverage gap from 4.20% to 1.12%.

Why it matters: This work advances uncertainty quantification in machine learning by making prediction sets more reliable under size constraints.

ModelsOfficialarXiv Statistical ML

Density-Informed Pseudo-counts Enhance Uncertainty Calibration in Evidential Deep Learning

A recent arXiv preprint presents Density-Informed Pseudo-count EDL (DIP-EDL), a new method designed to improve uncertainty calibration in Evidential Deep Learning (EDL) models. DIP-EDL addresses the issue of overconfidence, particularly on out-of-distribution data, by decoupling class prediction from uncertainty estimation through separate modeling of label distribution and input density. The paper provides both theoretical justification and empirical evidence that DIP-EDL leads to better interpretability, robustness, and uncertainty calibration under distributional shift.

Why it matters: Accurate uncertainty calibration is essential for deploying deep learning models in real-world and safety-critical scenarios, where overconfidence can have serious consequences.

ResearchOfficialarXiv Statistical ML

Latency-Response Theory Model: Evaluating LLMs via Accuracy and Chain-of-Thought Length

Researchers introduce the Latency-Response Theory (LaRT) model, which jointly models large language model (LLM) response accuracy and chain-of-thought (CoT) length for evaluation purposes. The model incorporates a correlation parameter between latent ability and latent speed, and is shown through theoretical analysis, simulations, and real LLM benchmark data to outperform traditional Item Response Theory (IRT) in estimation accuracy and evaluation efficiency. LaRT also produces different LLM rankings and demonstrates improved predictive power and ranking validity compared to IRT.

Why it matters: This approach could lead to more nuanced and statistically robust assessments of LLM reasoning by leveraging both accuracy and reasoning process length.

ResearchOfficialarXiv Statistical ML

Stable Signal Principle Explains Retraining Convergence in Performative Prediction

A new theoretical framework, the stable signal principle, demonstrates that retraining predictive models converges to a stable direction when a nonzero model-independent signal exists, even if the model's influence on the data is strong. The analysis generalizes to affine retraining operators and applies to language model training with synthetic data, offering a unified explanation for stability in performative prediction loops.

Why it matters: This work provides a theoretical explanation for the convergence and stability of retraining in real-world learning systems, including language models, even under strong feedback effects.

ResearchOfficialarXiv Software Engineering

CodeCoR: An LLM-Based Self-Reflective Multi-Agent Framework for Code Generation

Researchers have introduced CodeCoR, a multi-agent framework for code generation that employs four large language model (LLM) agents to generate prompts, code, test cases, and repair advice. By pruning low-quality outputs at each stage, CodeCoR aims to reduce error propagation in sequential code generation workflows. Experiments on HumanEval and MBPP datasets show that CodeCoR achieves an average Pass@1 score of 77.13%, outperforming existing baselines.

Why it matters: This work demonstrates a significant advance in code generation reliability by integrating self-reflection and output pruning in a multi-agent LLM framework, addressing a key limitation of sequential LLM-based approaches.

ResearchOfficialarXiv Software Engineering

AutoSpec: Automated Generation of Neural Network Specifications

AutoSpec is a framework that automatically generates and evaluates neural network specifications for learning-augmented systems, particularly in safety-critical domains. It uses a tree-based algorithm to partition the input space and a statistical certification framework to provide accuracy guarantees for each specification. Experiments across four applications show that AutoSpec improves F1 score by up to 53% over human-defined specifications and 73% over the strongest baseline.

Why it matters: This work automates the specification process in neural network verification, addressing a key bottleneck and making formal safety guarantees more practical for learning-augmented systems.

ResearchOfficialarXiv Software Engineering

Attention to Detail: Evaluating Energy, Performance, and Accuracy Trade-offs Across vLLM Configurations

A large-scale empirical study examines how different vLLM inference engine configurations—specifically attention kernel type, prefix caching, and chunked prefill—affect energy consumption, latency, and accuracy across multiple open-weight LLMs and tasks. The findings show that attention kernel type and prefix caching have significant impacts on energy and performance, while chunked prefill has limited effect under default settings. The optimal configuration is highly dependent on the specific model and workload, with no single configuration performing best in all scenarios.

Why it matters: This research provides practical guidance for optimizing LLM inference deployments, demonstrating that while configuration tuning can yield local improvements, the choice of model has a greater overall impact on trade-offs between energy, performance, and accuracy.

ResearchOfficialarXiv Software Engineering

Empirical Study Identifies Agent-Reactive Bugs at the Model-Harness Boundary in LLM Agents

A new empirical study analyzes 255 bug reports from Codex, Gemini-CLI, LangChain, and CrewAI, introducing a new class of bugs termed agent-reactive (AR) bugs that arise at the interface between LLM model outputs and harness code. These AR bugs often manifest as silent errors without clear test oracles, and their reproduction is complicated by the stochastic nature of LLM responses. The study also finds a mismatch between user-proposed fixes, which often suggest harness-side guardrails, and developer responses.

Why it matters: This research exposes a critical challenge in LLM agent reliability, emphasizing the need for new testing and debugging tools to address bugs that emerge from the interaction between model and harness.

ResearchOfficialarXiv Robotics

BayesContact: Simulation-Based Inference for Visuo-Tactile Pose Estimation in Peg-in-Hole Insertion

BayesContact is a Simulation-Based Inference framework that combines depth and force/torque observations to estimate object pose during peg-in-hole insertion tasks. The method maintains a particle belief over pose, using a renderer and physics simulator to score hypotheses against real observations. In both simulated and real-robot experiments, BayesContact improves pose observability and increases insertion success by 30% compared to vision-only inference.

Why it matters: This work offers a practical approach to fusing vision and tactile sensing for precise pose estimation in contact-rich robotic manipulation, potentially reducing the need for retraining across different environments and geometries.

ResearchOfficialarXiv Robotics

Handroid: A Desktop-Scale Robot That Reconfigures from Dexterous Hand to Humanoid

Researchers have introduced Handroid, a 27-degree-of-freedom (DoF) reconfigurable robot capable of switching between a dexterous hand and a desktop-scale humanoid form (0.33 m tall, 2.05 kg). The platform supports teleoperation, dexterous grasping, in-hand manipulation, humanoid locomotion, and long-horizon tasks that involve changing its embodiment. Handroid is validated through real-world manipulation, reinforcement-learning-based locomotion, and tasks requiring reconfiguration and coordinated actions.

Why it matters: Handroid enables unified research on dexterous manipulation and humanoid mobility within a single, compact, and reconfigurable platform, advancing the field of morphology-reconfigurable robotics and cross-embodiment learning.

ResearchOfficialarXiv Robotics

Orbis 2: A Hierarchical World Model for Driving

Researchers introduce Orbis 2, a hierarchical world model for autonomous driving that separates future prediction into two levels: a high-level predictor for coarse scene structure over long time horizons and a low-level generator for detailed predictions. The model is trained in two stages, first with diffusion forcing pretraining to enhance internal representations, followed by teacher forcing fine-tuning for stable rollouts. Orbis 2 achieves state-of-the-art results on standard driving world model benchmarks, including long-horizon generation fidelity and steering responsiveness.

Why it matters: This work demonstrates a significant advance in autonomous driving world models by combining long-horizon spatial reasoning with high perceptual fidelity, leading to improved performance and internal representations.

ResearchOfficialarXiv Software Engineering

Not All Tokens Matter: Data-Centric Optimization for Efficient Code Summarization

A new preprint investigates token-level reduction techniques to lower computational costs in large language model (LLM)-based code summarization. The study finds that optimal token reduction strategies are highly language-dependent: AST-based reduction improves Java summarization performance by 37% but degrades Python by up to 49%, while CrystalBLEU-guided pruning achieves robust cross-language token reduction of 60-72%. Function signature-based reduction is optimal for Python, achieving 83% token reduction while maintaining summary quality.

Why it matters: The work demonstrates that language-aware token curation is crucial for efficient code summarization, challenging assumptions about the transferability of token reduction strategies across programming languages.

ResearchOfficialarXiv Software Engineering

NexForge: Requirement-Driven Task Synthesis Scales LLM Agent Capabilities

NexForge is a framework that automatically synthesizes diverse, executable agent tasks and expert trajectories from high-level capability requirements, removing the need for domain-specific infrastructure. The system generated thousands of terminal and office tasks, significantly improving the performance of Qwen3.5-35B-A3B on benchmarks such as Terminal-Bench 2.0 and GDPval. Scaling up to 43,200 tasks enabled the training of Nex-N2 models, which achieve state-of-the-art open-source results and surpass several proprietary systems.

Why it matters: NexForge demonstrates a scalable approach to expanding LLM agent capabilities by automating task synthesis from requirements, reducing manual engineering and enabling broader domain coverage.

ResearchOfficialarXiv Robotics

DPNeXt: Lightweight Multi-Task Decoder for Efficient ViT-Based Dense Prediction

A new preprint introduces DPNeXt, a lightweight multi-scale feature fusion decoder designed to replace the standard Dense Prediction Transformer (DPT) for multi-task dense prediction in robotics. DPNeXt employs dual depthwise separable inverted bottlenecks and a Multi-Task Boundary Guidance strategy to improve efficiency and mitigate negative transfer between tasks. The method reduces trainable parameters by 78.6% compared to DPT and achieves state-of-the-art results on Cityscapes and NYUv2 benchmarks, with faster inference on resource-constrained hardware.

Why it matters: This work advances efficient multi-task perception for robotics, enabling real-time scene understanding on devices with limited computational resources.

ResearchOfficialarXiv Robotics

SLAC: Safe and Efficient Real-Robot RL via Unsupervised Simulation Pre-Training

SLAC is a reinforcement learning method that uses a low-fidelity simulator to pretrain a task-agnostic latent action space through unsupervised skill discovery, enabling efficient and safe real-world learning for high-degree-of-freedom robots. The approach achieves state-of-the-art performance on bimanual mobile manipulation tasks, learning contact-rich, whole-body behaviors in under an hour of real-world interaction without demonstrations or hand-crafted priors.

Why it matters: SLAC advances the feasibility of real-world reinforcement learning for complex robots by combining simulation-based pretraining with safe, sample-efficient real-world learning.