TODAY — FRI 04 SEP 2026 20 in view
◆ FRONTIER · INDUSTRY IMPACT
Scaling Bimanual Household Manipulation from 1,500 hours of Demonstrations to On-Policy Corrections Learning generalist policies for robust bimanual manipulation is bottlenecked by the scarcity of high quality large scale human demonstration data. In this work, we release 1,500 hours of diverse bimanual manipulation demonstrations covering everyday household tasks, and use this comprehensive corpus to train XR-2, a powerful vision-language-action (VLA) model. Enabled by a purpose built high throughput data pipeline… read on
bimanual-manipulation vision-language-action imitation-learning robotics-datasets
THE INDEX — 19 MORE CLEARED
● Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints cs.AI · cs.LG
● Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor Triggers cs.CV · cs.CR
● Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression cs.CY · cs.AI
● A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors cs.CR · cs.AI
● Seeing Less Is Not Seeing Safely: Privacy Leakage from Task-Scoped Robot Perception Exports cs.RO · cs.CR
● IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks cs.CL · cs.AI
● Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning cs.CR · cs.AI
● A Black Box for Agentic Processes: Blockchain-Anchored Evidence for AI Agent Communication, Human Oversight, and GRC Audits cs.CR
● FiMI Banking: A Sovereign Model for Indian Retail Banking cs.AI · cs.CL
● The Natural Language Interaction Protocol and Standard for AI Agents IBM cs.AI
● Mind the Gap: Robustness Risks in PII Detection Systems cs.LG
● Flip, Don't Shuffle: Watermarking LLMs at the Speed of Inference cs.CR · cs.CL
● BRIDGE: An Open-Source Humanoid Platform via Morphology-Control Co-Design for Physical AI cs.RO · cs.AI
● EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders cs.CV · cs.AI
● When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization cs.CR · cs.AI
● AlcaTRAz - Anchored Tree-Rule Defense Against Jailbreaks cs.CR
● A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms Google cs.AI
● LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes cs.CV · cs.AI
● The Illusion of Independent Quorums: Epistemic Fault Domains and Correlated Cognitive Failures in Agentic Quorums cs.DC · cs.MA
THU 03 SEP 2026 37 in view
◆ FRONTIER · INDUSTRY IMPACT
EvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language Models Frontier large language models can often recognize when they are being evaluated, a capability known as evaluation awareness. If models behave differently in evaluations than in deployment, this undermines the validity of evaluation results, which are a crucial component of current AI safety frameworks. We introduce EvalDetectBench, an open pipeline and benchmark for measuring evaluation awareness that works with any… read on
evaluation-awareness llm-evaluation benchmark safety-evals
◆ FRONTIER · INDUSTRY IMPACT
A Finger on the Scale: Covert Policy Steering through Agentic Skills Reusable agent skills extend large language model (LLM) agents with task procedures, tool-use guidance, and output constraints. Yet these skills also act as externalized behavioral policies, which create a supply-chain risk: a third-party skill may preserve the declared task and valid output interface while covertly redirecting agent decisions toward an undisclosed objective. We formalize Skill Policy Integrity, whic… read on
agentic-skills policy-integrity covert-steering supply-chain-risk
THE INDEX — 35 MORE CLEARED
● Stored Is Not Supported: Typed Provenance and Assertion Guardrails for Persistent AI Agents cs.CR
● ASCII Attack: Recontextualising Harmful Requests as Artistic Critique in Large Language Models cs.AI
● Agent Memory Is a Surface for Endogenous Authorization Laundering cs.CR · cs.AI
● SPADE: SPaT Attack Detection from the Connected Vehicle's Perspective cs.CR · cs.LG
● Privacy Washing: Detecting Internal Contradictions in Privacy Policies cs.CY · cs.CL
● Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy Pattern Google cs.AI · cs.SE
● VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages cs.CL · cs.AI
● READY or Not: Reliable Enterprise Agent Deployment cs.AI
● SEAL: Reinforcing Global Safety in Mixture-of-Experts through Shared Expert ALignment cs.LG · cs.AI
● Counter-GEO-Bench: Evaluating Defenses Against Information-Distorting Generative Engine Optimization cs.IR · cs.CL
● Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning cs.AI
● text2ql: Multi-Target Natural Language Querying via a Language-Agnostic Intermediate Representation cs.CL · cs.AI
● Door-in-the-Face Requests and Refusal Behaviour in Large Language Models cs.AI · cs.CL
● Hearing the Whispers: Black-Box Membership Inference Attacks on Finetuned TTS Models cs.CR · cs.LG
● The Implications of Linguistic Illegibility for LLM Security cs.LG · cs.CR
● Learning-Based Reconstruction Attacks on Coordinate-Obfuscated Point Clouds cs.CR · cs.LG
● Humanoid Safe Stop via Learned Stoppability Value cs.RO · cs.LG
● FairLens: Benchmarking Fairness in Vision-Language Models for High-Stakes Decision-Making cs.CV · cs.LG
● Before the Script, Set the Stage: How Worldview Simulation Amplifies Psychologically Grounded Persuasion in Multi-Turn Jailbreaking cs.CL · cs.AI
● Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching cs.CR
● ACLE-MCP: Attested Capability Leases for Execution-Time Trust in Remote LLM Tool Use cs.CR
● Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness cs.CV · cs.HC
● Toward Explainable and Policy-Aware AI for Carbon Credit Price Prediction: A Research Framework for Emerging Carbon Markets cs.LG
● The Shape of Ownership: Verifying LLM Provenance through Semantic Structures cs.CR
● FUSE: An Evaluating Framework for Dangerous Capabilities of LLMs cs.AI
● InfraPatch: Cross-Task Targeted Grayscale Patch Attacks on Infrared-Adapted Vision-Language Models cs.CV · cs.AI
● Skill-as-API: Confidential Multi-Agent Coordination for Agentic Software Engineering cs.CR
● SCX Router: Streaming Zero-Shot Model Selection with a Decoder-KV Classifier and a Real-World Task Ontology cs.AI · cs.CL
● Competitive Market Behavior of LLMs cs.MA · cs.AI
● CodePoisonRAG: Knowledge Poisoning Attacks on Retrieval-Augmented Code Generation cs.CR · cs.LG
● Post-Training Ternarization of Qwen3-4B Capability, Effective Bit Budget, Storage Compression, and Deployment cs.AI · cs.LG
● Evaluating ML-based Intrusion Detection Systems: The Illusion of Model Efficacy cs.CR
● Context Inference Attacks Without Jailbreaks cs.CR · cs.LG
● WeaveMark: Robust and Scalable Multi-bit LLM Watermarking via Coded Payload Spreading cs.CR · cs.LG
● How Fast Do Agents Rot? An Empirical Study of Long-Horizon Degradation in LLM Agents for Production Decision-Making Microsoft physics.soc-ph · cs.AI
WED 02 SEP 2026 41 in view
◆ FRONTIER · INDUSTRY IMPACT
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are commonly scalarized with fixed aggregation weights. We identify a failure mode in which aggregation itself induces reward hacking: static projection al… read on
reward-hacking multi-reward-rl rlhf reward-aggregation
◆ FRONTIER · INDUSTRY IMPACT
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-Verified Benchmark Rewriting (RePro), the first framework to integrate Lean-oriented neural automated theorem provers (ATPs) into benchmark rewriting,… read on
llm-evaluation benchmark-rewriting proof-verification lean-atp
THE INDEX — 39 MORE CLEARED
● AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMs cs.CR
● Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning cs.CL · cs.AI
● What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harness cs.CR
● Don't Let the Model Write the YAML: Deterministic, Minimal-Diff GitOps Remediation from LLM-Proposed Field Changes cs.SE · cs.AI
● Instella-MoE Technical Report cs.CL · cs.AI
● AI Morbidity and Mortality: A Framework for Clinical AI Failure Review cs.AI · cs.HC
● Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetry cs.CL · cs.CR
● ChatDev 2.0: A No-Code Multi-Agent Platform for Developing Everything cs.AI · cs.CL
● Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumption cs.DS · cs.LG
● Validity-Aware Jailbreak Evaluation for Large Language Models cs.AI
● The Privacy-Hallucination Tradeoff in Differentially Private Language Models cs.AI · cs.CL
● ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training cs.CV
● CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMs cs.LG
● When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuning cs.CR · cs.AI
● Workload Identification with Physical Side Channels for AI Governance cs.CR · cs.AI
● Transferable End-to-End Optimization for Indirect Long-Term Memory Poisoning in LLM Agents cs.CR
● Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matching cs.AI
● SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems cs.CR · cs.AI
● LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark cs.AI · cs.CL
● SilentProbe: Measuring Silent Failure in Production APIs Used as Agent Tools cs.IR · cs.SE
● Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades cs.AI · cs.CR
● VerTox: Verifiable Reward-Guided Corpus Poisoning Against Neural Ranking Models cs.CL · cs.IR
● EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilities cs.CL · cs.AI
● When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation cs.AI
● The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systems cs.AI
● Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Trees cs.AI · cs.LG
● The Constitutional Coverage Trilemma in AI Governance cs.LG · cs.AI
● AgentProv: Auditing Agentic LLM API Providers via Tool-use Policy Probes cs.CR · cs.CL
● TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoning cs.CL · cs.CR
● Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models cs.AI · cs.CL
● The Safeguard Worked. Is the LLM System Safer? cs.CR · cs.AI
● Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderation cs.CV · cs.AI
● Federated Trust for Embodied Robot Capability Marketplaces cs.CR · cs.RO
● Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domain cs.CV
● LatentPress: Context Compression Beyond Text and Vision cs.LG · cs.AI
● From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scaling cs.CL · cs.AI
● Effective Interventions Against AI-Enhanced Scams cs.CR · cs.CY
● Causal Evidentiary Governance for High-Risk Machine Learning Systems cs.CY · cs.AI
● Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debate cs.AI · cs.MM
TUE 01 SEP 2026 2 in view
◆ FRONTIER
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length. We assume inputs are fully parenthesized, well-fo… read on
algorithmic-generalization transformers circuit-computation length-generalization By IBM
THE INDEX — 1 MORE CLEARED
● Authority-Inference Separation in Agentic Finance: First-Line Control, Blockchain Enforcement, and Replayable Assurance q-fin.GN · cs.CR
loading earlier editions…