WED 26 AUG 2026 28 in view
28 CLEARED
● Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Information cs.MA · cs.CL
● StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancing cs.AI · cs.CR
● What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructions cs.CR · cs.AI
● WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agents cs.CR · cs.AI
● Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectors cs.AI · cs.CL
● ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergence cs.AI
● Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems cs.AI
● Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matching cs.IR · cs.CL
● PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triage cs.LG
● Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2) cs.CR · cs.AI
● Maia 200: A Software Defined Dataflow System for Large-scale AI Acceleration cs.AR · cs.AI
● Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model cs.MA
● Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attention cs.AI
● Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signals cs.HC · cs.CY
● More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving cs.AI
● Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model cs.SE · cs.AI
● Paritok-4B: Intent-Conditioned Context Compression for Coding Agents cs.AI · cs.CL
● Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMs cs.CL · cs.AI
● Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agents cs.LG · cs.AI
● NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical Robotics NVIDIA cs.RO
● Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI Applications IBM cs.AI
● The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problem math.OC · cs.AI
● Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes cs.CL
● Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems Google cs.AI · cs.CE
● TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers cs.CR · cs.AI
● 'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection cs.CL · cs.AI
● Neurosymbolic Alignment for Physiologically-Safe Clinical Language Models cs.AI
● NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistribution cs.CR · cs.AI
TUE 25 AUG 2026 68 in view
◆ FRONTIER · INDUSTRY IMPACT
OmniCAD: A Large-Scale Benchmark for 3D Spatial Reasoning in Robotics Assemblies Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex mechanical assemblies remains underexplored. We introduce OmniCAD, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery. Om… read on
3d-spatial-reasoning robotics-assemblies benchmark vision-language-models By Microsoft
◆ FRONTIER · INDUSTRY IMPACT
Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing o… read on
rlvr data-leakage memorization pii-extraction ↗ DeepSeek-V3.1 ↗ DeepSeek-V3
◆ FRONTIER · INDUSTRY IMPACT
Apodex 1.1: Scaling Agentic Intelligence for Complex Work General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimens… read on
agentic-ai working-capability environment-scaling verifiable-execution
THE INDEX — 65 MORE CLEARED
● Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models cs.CL · cs.AI
● Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty cs.AI · cs.CL
● GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGeneration Meta cs.CV · cs.AI
● What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine cs.LG · cs.AR
● PURA: Provably Unbiased and Robust Multi-Bit Text Attribution cs.CR
● Redteaming Leading Arabic LLMs with ASAS cs.AI
● SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation cs.CR · cs.AI
● KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inference cs.AI · cs.DC
● Counterfactual, Per-Decision Bias Auditing for Automated Hiring: Localizing and Explaining Disparate Impact in Applicant Tracking Systems cs.CY · cs.CR
● AraDetox: A Multi-Dialect Arabic Detoxification Dataset cs.CL · cs.AI
● ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models cs.CR · cs.AI
● K-Bench: measuring model performance on real scientific agent requests cs.AI · cs.CL
● AdaptPrint: Response-Adaptive Fingerprinting of Black-Box LLM Services cs.CR
● Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testing cs.CL · cs.AI
● The Compaction Cliff in Long-Running AI Agent Memory cs.AI · cs.IR
● DIME: Query-Efficient Framework for Membership Inference on Diffusion Models cs.LG · cs.CR
● Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraints cs.RO · cs.AI
● AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systems cs.CR
● VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generation cs.CV
● Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attacks cs.CR · cs.AI
● PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policies cs.AI
● Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learning cs.LG · cs.AI
● Mitigating Explanation Leakage in Financial Fraud Detection Systems cs.LG · cs.CR
● SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agents cs.CR · cs.CL
● MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feeds cs.AI · cs.CY
● Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifolds cs.CR
● Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Models cs.CL
● Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning cs.AI
● AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systems cs.AI
● Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agents cs.CL · cs.AI
● BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearning cs.LG · cs.AI
● Adversarial Entropy Inflation Against Gumbel-Based Inference Verification cs.CR · cs.AI
● MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance cs.AI · cs.CL
● There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items cs.AI
● LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization cs.AI
● RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authored cs.CL · cs.IR
● AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governance cs.AI · cs.CR
● CIDER: Continual Interactive Distillation for Embodied Reinforcement Learning cs.RO
● Mitigating Database Leakage in RAG Systems with Keyword-Grounded Fact Substitution cs.CL
● Enhancing User Resilience Against AI-Augmented Phishing: A Two-Stage Framework for Detection and Personalized Training cs.CR
● AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations cs.AI · cs.HC
● Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Code cs.SE · cs.AI
● Evaluation Awareness in Language Models: Representation, Verbalization, and Control cs.CL · cs.AI
● What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation cs.CL · cs.AI
● Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation cs.RO · cs.AI
● A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systems cs.AI
● Measuring Activation Control in Large Language Models cs.AI · cs.CL
● Proxy reliance in large language model decisions is uncalibrated to predictive evidence cs.AI · cs.CL
● Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systems cs.CL · cs.AI
● Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation Tencent cs.CL
● TessIndex: Capability Verified Identity System for the Agent Economy cs.AI · cs.CR
● InjecMEM: Memory Injection Attack on LLM Agent Memory Systems cs.CR · cs.AI
● Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequality cs.CL · cs.AI
● No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarios cs.CL
● Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Tasks cs.AI
● GOLEM: Modular Humanoid Autonomy Towards Electric Vehicle Battery Disassembly cs.RO
● Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents cs.CV · cs.AI
● First Demonstration of Multi-Agent LLM System for Million-Scale Optical Link Management in Global Production AIDCs Baidu cs.MA · physics.optics
● RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perception cs.RO · cs.AI
● Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens cs.CY · cs.AI
● Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf cs.AI · econ.GN
● Reviewing Model Collapse and Countermeasures cs.AI · cs.LG
● Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harms cs.CL · cs.AI
● ST²U: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Control cs.LG · cs.CL
● From Inertia to Objectivity: Improving Deep Research Agents with Noise Isolation Alibaba cs.AI
MON 24 AUG 2026 4 in view
4 CLEARED
● AID-Guard: Stateful Authorization for Delegated Agent Effects cs.CR · cs.AI
● AEGIS: Preventing Cross-Domain Resource Abuse in MCP IBM cs.CR · cs.AI
● Who Delegates to AI? Evidence from 53,000 Agent Configurations cs.AI · cs.CY
● Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AI cs.CV · cs.CR
loading earlier editions…