What changed in AI — Page 118

Open SourceReportedMarkTechPost / AI

How to Build a T4-Friendly Autonomous Data Science Agent with DeepAnalyze-8B

A tutorial demonstrates building an autonomous data science agent using DeepAnalyze-8B, sandboxed code execution, and iterative analysis. The agent runs on a T4 GPU in Colab, loads the model in 4-bit mode, and performs data cleaning, analysis, and visualization. It generates an analyst-grade report from a multi-file e-commerce dataset.

Why it matters: This tutorial shows how to run a capable data science agent on limited hardware, making autonomous analysis accessible to more developers.

ResearchReportedMarkTechPost / AI

NeuroVFM: A Neuroimaging Foundation Model Trained on 5.24M Clinical MRI and CT Volumes

Researchers at the University of Michigan have developed NeuroVFM, a generalist neuroimaging foundation model trained on 5.24 million clinical MRI and CT volumes. The model uses Vol-JEPA, an extension of I-JEPA and V-JEPA, to learn brain anatomy and pathology without requiring radiology-report labels.

Why it matters: This model could enable more scalable analysis of brain imaging data without the need for labeled datasets.

Policy & SafetyReportedMETR

METR Analysis: Anthropic's Researcher Uplift from Coding Agents Plausibly >2x

METR researcher Thomas Kwa analyzes Anthropic's reported 8x increase in code merged per day in Q2 2026 versus 2021-2024. Using economic production models and assuming code quality equivalence, he estimates that coding agents alone yield a researcher uplift of at least 2x, with most models predicting uplift between 2.33x and 2.91x. The analysis notes that these estimates do not account for potential uplift from non-coding tasks.

Why it matters: This analysis provides a quantitative framework for understanding how AI coding agents may amplify researcher productivity, with implications for AI development speed and economic impact.

ResearchReportedMarkTechPost / AI

A Coding Guide to NVIDIA’s Tile-Based GPU Programming: From cuTile and Triton Kernels to Flash Attention

A tutorial explores NVIDIA tile-based GPU programming using TileGym, building a Colab workflow that runs across different hardware. It covers core tile concepts and implements vector addition, fused GELU, row-wise softmax, tiled matrix multiplication, and flash attention, checking each against PyTorch.

Why it matters: This guide provides practical, hands-on instruction for developers to leverage NVIDIA's tile-based GPU programming techniques, which can significantly improve performance for AI workloads.

Products & AgentsOfficialAWS Machine Learning Blog

AWS Launches UI for Generative AI Inference Recommendations in SageMaker AI

AWS has introduced a user interface for generative AI inference recommendations in Amazon SageMaker AI Studio, offering a low-code/no-code experience. The new UI guides users through preset use-case profiles, visual comparisons, and one-click deployment, enabling teams without deep infrastructure expertise to obtain validated configurations.

Why it matters: This update makes it easier for more teams to deploy optimized generative AI models without requiring deep infrastructure knowledge.

Companies & FundingReportedThe Verge / AI

Apple Lawsuit Alleges OpenAI Stole Trade Secrets and Spied on Hardware Prototypes

Apple has filed a lawsuit accusing OpenAI of stealing confidential documents, spying on hardware prototypes, and tricking an employee. The suit claims OpenAI's hardware head asked Apple job applicants to bring unreleased product samples to interviews. These allegations are based on a report from The Verge.

Why it matters: This lawsuit could escalate tensions between two major tech companies and raise questions about corporate espionage in the AI industry.

ResearchOfficialarXiv AI/ML

LLM-Driven Evolution Generates Multi-Objective Bayesian Optimization Algorithms Outperforming Human Designs

Researchers extended the LLaMEA framework to evolve multi-objective Bayesian optimization (MOBO) algorithms, using large language models as mutation and crossover operators. The best generated algorithm achieved a mean normalized hypervolume of 0.971 on synthetic problems (vs. 0.869 for baseline qParEGO) with about 60x less wall-clock time, and 0.985 on real-world engineering problems (vs. 0.971) at roughly 3.4x lower cost. The study demonstrates that LLM-driven evolutionary search can discover algorithm designs with superior Pareto efficiency compared to manual approaches.

Why it matters: This work shows that LLMs can autonomously design optimization algorithms that match or surpass human-crafted ones, potentially accelerating progress in multi-objective optimization.

ResearchOfficialarXiv AI/ML

Neuro-Agentic Control Framework Uses LLM and Time-Series Foundation Model for Industrial Cybersecurity

Researchers propose a neuro-agentic control framework that combines an LLM-based planner (such as Gemini 2.5 Flash-Lite) with a pre-trained Time-Series Foundation Model (TimesFM) for autonomous defense in industrial IoT. The framework introduces a 'Counterfactual Physics Injection' mechanism, which simulates LLM-proposed interventions in the foundation model's latent space before actuation, allowing the system to reject hallucinated or unsafe actions. Evaluated on the SWaT dataset, the framework prevented five breaches (33.3%) below threshold, outperforming LSTM (26.7%) and TCN (13.3%) baselines, with zero physically invalid actions executed.

Why it matters: This work demonstrates a practical method to ground LLM-based agents with physics-aware foundation models, addressing safety concerns for closed-loop control in critical infrastructure.

ResearchReportedArs Technica / AI

Simulating everything, sort of: The promise and limits of world models

World models aim to simulate environments for AI training, but experts highlight both their potential and current limitations. The article explores how these models work and what remains unsettled.

Why it matters: World models could revolutionize AI training by enabling safe, scalable simulation, but their limits must be understood to avoid over-reliance.

Policy & SafetyReportedArs Technica / AI

Defenders Embrace Prompt Injection to Thwart Hacking AI Agents

Security researchers are employing a technique known as 'context bombing' to defend against malicious AI agents. By injecting overwhelming or confusing prompts, they can cause hacking agents to shut down before executing harmful actions.

Why it matters: This represents a shift in the use of prompt injection from an attack method to a defensive strategy in AI security.

Products & AgentsReportedThe Decoder

Meta Removes Muse Image Feature Allowing AI Generation of Instagram Users Without Consent

Meta has removed a controversial feature from its Muse Image model that allowed users to generate AI images of other people by @-mentioning their public Instagram accounts without consent. The company acknowledged the feature "missed the mark" and shut it down just days after its announcement, following widespread criticism.

Why it matters: This incident underscores the challenges tech companies face in balancing AI innovation with user privacy concerns.

Companies & FundingReportedThe Decoder

S&P Global downgrades Oracle credit rating, cites OpenAI as key risk

S&P Global has downgraded Oracle's credit rating to 'BBB-', one notch above junk status, citing OpenAI as a key credit risk. OpenAI accounts for roughly half of Oracle's $638 billion in contractual obligations, and if OpenAI walked away, Oracle would be left with significant unused data center capacity.

Why it matters: This downgrade highlights the financial risk for cloud providers when a single customer represents a large portion of their contractual obligations.

ResearchReportedWIRED / AI

Scientists Use AI and Quantum Computing to Generate New Peptides for Rare Diseases

Researchers have combined AI and quantum computing to generate new peptides, with the goal of aiding drug development for underserved populations and rare diseases. The project was supported by limited funding and conducted in researchers' spare time.

Why it matters: This approach could help advance drug discovery for neglected diseases by leveraging emerging quantum and AI technologies.

Products & AgentsReportedThe Decoder

Anthropic: Claude Cowork's Main Use Case Is Routine Office Work

Anthropic analyzed 1.2 million Claude Cowork sessions from over 600,000 organizations and found that about half of all usage is dedicated to business processes and text creation, which it refers to as 'the work around the work.' Tasks include compiling status reports, building onboarding checklists, and creating slide decks. Software development is rarely done in Cowork, as developers prefer Claude Code for those tasks.

Why it matters: This highlights that enterprise AI adoption is currently focused on automating routine office tasks rather than technical work.

ModelsOfficialPartnership on AI

Partnership on AI Highlights Progress in Responsible AI Development

Partnership on AI has published an update detailing advancements in responsible AI practices. The organization outlines ongoing efforts to shape ethical standards and best practices for AI development and deployment.

Why it matters: Establishing responsible AI frameworks is crucial for ensuring ethical and safe AI technologies.

Policy & SafetyOfficialPartnership on AI

Partnership on AI Announces New Global Initiatives to Measure Progress in Responsible AI

Partnership on AI has announced new global initiatives aimed at measuring progress in responsible AI. The initiatives focus on developing metrics and frameworks to assess responsible AI practices worldwide.

Why it matters: This effort provides a standardized way to evaluate and compare responsible AI progress across organizations and countries.

Policy & SafetyOfficialPartnership on AI

When Companies Listen to Employees About AI, Everyone Benefits

Partnership on AI reports that companies involving employees in AI implementation processes can achieve better outcomes for both businesses and workers. The article emphasizes the value of inclusive AI governance and employee engagement.

Why it matters: This highlights that employee involvement is important for ethical and effective AI adoption in the workplace.

Products & AgentsOfficialAdobe Research

Adobe Research unveils AI Assistant for Photoshop that edits images via natural language

Adobe Research has developed an AI Assistant for Photoshop that allows users to edit images by describing their desired changes in natural language. The assistant can either perform the edit automatically or guide users through the editing process.

Why it matters: This AI Assistant could make advanced image editing more accessible to users without professional expertise.

ResearchOfficialMIT News / Artificial Intelligence

New Spatial Memory System Helps Robots Remember Object Locations

MIT researchers have developed a spatial memory system that enables robots to efficiently capture and recall details about objects in their environment. The system works by storing object locations and features during exploration, potentially allowing robots to help find misplaced items like keys.

Why it matters: This advance could improve human-robot interaction in homes and workplaces by enabling robots to assist with everyday tasks such as locating lost objects.

Products & AgentsOfficialAdobe Research

Adobe Firefly Adds Alt Text for Blind and Low-Vision Creators

Adobe Research has introduced a new feature in Adobe Firefly that generates alt text for AI-generated images, enhancing accessibility for blind and low-vision creators. This feature provides descriptive text to help users understand and select images.

Why it matters: This feature increases inclusivity in AI image generation by enabling blind and low-vision users to independently create and choose images.