TODAY — WED 02 SEP 2026 31 in view
◆ FRONTIER · INDUSTRY IMPACT
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are commonly scalarized with fixed aggregation weights. We identify a failure mode in which aggregation itself induces reward hacking: static projection al… read on
reward-hacking multi-reward-rl rlhf reward-aggregation
◆ FRONTIER · INDUSTRY IMPACT
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-Verified Benchmark Rewriting (RePro), the first framework to integrate Lean-oriented neural automated theorem provers (ATPs) into benchmark rewriting,… read on
llm-evaluation benchmark-rewriting proof-verification lean-atp
THE INDEX — 29 MORE CLEARED
● AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs cs.CR
● Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning cs.CL · cs.AI
● What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness cs.CR
● Instella-MoE Technical Report cs.CL · cs.AI
● AI Morbidity and Mortality: A Framework for Clinical AI Failure Review cs.AI · cs.HC
● Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry cs.CL · cs.CR
● Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption cs.DS · cs.LG
● Validity-Aware Jailbreak Evaluation for Large Language Models cs.AI
● The Privacy-Hallucination Tradeoff in Differentially Private Language Models cs.AI · cs.CL
● ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training cs.CV
● CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs cs.LG
● When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning cs.CR · cs.AI
● Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching cs.AI
● LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark cs.AI · cs.CL
● Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades cs.AI · cs.CR
● EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities cs.CL · cs.AI
● When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation cs.AI
● The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems cs.AI
● Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees cs.AI · cs.LG
● The Constitutional Coverage Trilemma in AI Governance cs.LG · cs.AI
● TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning cs.CL · cs.CR
● Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models cs.AI · cs.CL
● Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation cs.CV · cs.AI
● Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain cs.CV
● LatentPress: Context Compression Beyond Text and Vision cs.LG · cs.AI
● From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling cs.CL · cs.AI
● Effective Interventions Against AI-Enhanced Scams cs.CR · cs.CY
● Causal Evidentiary Governance for High-Risk Machine Learning Systems cs.CY · cs.AI
● Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate cs.AI · cs.MM
TUE 01 SEP 2026 61 in view
◆ FRONTIER
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length. We assume inputs are fully parenthesized, well-fo… read on
algorithmic-generalization transformers circuit-computation length-generalization By IBM
◆ FRONTIER · INDUSTRY IMPACT
A Causal Model for Locating and Unlocking Sandbagging in Model Organisms Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later lay… read on
sandbagging causal-mechanistic residual-stream model-governance ↗ llama-3 ↗ llama-3-8b ↗ qwen2.5-7b ↗ mistral-7b
◆ FRONTIER · INDUSTRY IMPACT
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events i… read on
long-horizon llm-agents benchmark autonomous-ecommerce By Alibaba ↗ glm 5 ↗ Qwen3 ↗ GPT-5.6 ↗ qwen3.8-max
◆ FRONTIER · INDUSTRY IMPACT
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three… read on
reward-hacking agent-evaluation benchmarks autonomous-ml
◆ FRONTIER · INDUSTRY IMPACT
A.X K2 Technical Report We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percent… read on
mixture-of-experts foundation-model agentic-ai token-efficiency
THE INDEX — 56 MORE CLEARED
● Authority-Inference Separation in Agentic Finance: First-Line Control, Blockchain Enforcement, and Replayable Assurance q-fin.GN · cs.CR
● Efficient GPU Retrieval for Semantic Search cs.AI · cs.LG
● SingProbe Technical Report cs.CR · cs.AI
● Verification-Time Dependency on a Disappearing Evaluator cs.CY · cs.SE
● CogEvol: Towards Efficient and Reliable Learning Environment Generation cs.CL · cs.AI
● CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Target cs.AI · cs.IR
● CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding cs.RO
● Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agents cs.CR · cs.AI
● Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steering cs.CL · cs.AI
● EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agents cs.AI · cs.CL
● Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verification cs.SE · cs.AI
● TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimization cs.CL · cs.LG
● Membership is Ownership: A Robust Ownership Verification Framework for Diffusion Models cs.CR · cs.CV
● TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Training cs.LG
● ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systems cs.CR
● STEP: A Modular Silent Trial Engine for Operational Evaluation of Digital Pathology AI in Routine Workflow cs.SE · cs.CV
● Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Models cs.CL · cs.CV
● GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Silicon cs.CL · cs.AI
● Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D Modeling cs.CV
● Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities cs.LG · cs.AI
● The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and Refusal Meta cs.LG
● Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation Alibaba cs.AI · cs.RO
● IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamil cs.CL · cs.AI
● SemTrace: Source-Grounded Semantic Signatures for Tracing LLM Exposure to Protected Documents cs.CL
● Extracting Knowledge from Tools in LLM Agents cs.CR
● Do VLMs Share Safety Neurons Across Modalities? cs.LG
● AI Can Be Easily Persuaded in Clinical Decision Making cs.CL · cs.AI
● Can escalation channels redirect reward hacking toward defect disclosure? cs.AI · cs.CR
● WebWorld: The Browser as a World Model for Self-Improving Web Code cs.CL · cs.SE
● Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injection cs.CR · cs.AI
● Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollback cs.CR · cs.AI
● RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciences cs.AI · cs.CL
● Zero-Knowledge Predicate Proofs Between AI Agents: A Measured, Cross-Protocol Gateway and the Source-Integrity Gap cs.CR · cs.MA
● The Fragility of Jailbreak Robustness Across Operational States cs.CR · cs.CL
● ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternatives cs.CL · cs.SE
● One Capability or Many? Testing the Economic Validity of Frontier AI Evaluation cs.LG · cs.CY
● Defending Wearable VLMs Against Private Attribute Inference cs.CV · cs.AI
● Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents cs.AI · cs.CL
● Adversarial Calibration Attack on Autonomous Vehicles cs.RO · cs.CV
● Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration cs.HC · cs.CV
● Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language Models NVIDIA cs.CV
● Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data cs.AI · cs.CL
● Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Text cs.CL · cs.AI
● Manacá-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluation cs.CL
● DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution Alibaba cs.CV · cs.SD
● One note in three: a verified census of three deployed AI scribes, and the instrument that counted it cs.CL · cs.AI
● SIR: Self-improving Red-teaming for Compute Use Agents IBM cs.CR · cs.AI
● Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimization cs.LG
● Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering cs.LG
● Identity by Design, Demographics by Accident: Demographic Leakage and Suppression in Behavioral Biometric Embeddings cs.CR
● The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoning cs.LG · cs.AI
● Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture cs.CL · cs.AI
● The Price of Intelligence: A Quality-Adjusted Price Index for AI Services econ.GN · cs.LG
● WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenance cs.CR
● JITterFlip: Uncovering Fault Attack Surfaces in JIT-Compiled LLM Serving cs.CR
● Drishti: AI-Led Human-Directed Vulnerability Auditing for 5G Cores cs.CR
MON 31 AUG 2026 8 in view
8 CLEARED
● OpenStamp: A Watermark for Open-Source Language Models cs.CL · cs.AI
● FISGuard: Defending Against Membership Inference via Fixed Input Subspaces cs.CR · cs.AI
● Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Control cs.CR
● When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AI cs.AI · cs.CL
● EvoHarmBench: Breaking Content Moderation with Iterative Human-Like Evasion Alibaba cs.CL
● Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification cs.CR · cs.AI
● CURA: Certified Runtime Alarms for Computer-Use Agents cs.AI · cs.CV
● LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails cs.AI
loading earlier editions…