AI Research Radar
Today: 42 papers cleared of 484 scanned, led by Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training. Qwen3 is the literature’s most-referenced model (5 papers in the loaded window).
UPDATED 1M AGO
60 DAYS · 20,387 SCANNED
1,871 CLEARED · 119 ◆
1,871 CLEARED · 119 ◆
TODAY — WED 07 OCT 202642 cleared of 484 · 5 FRONTIER
◆ FRONTIER·INDUSTRY IMPACT
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-TrainingDevelopers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level re…
backdoor-persistencyllm-agentssftreinforcement-learning
◆ FRONTIER·INDUSTRY IMPACT
DecepEval: A Benchmark for Evaluating Deception in LLM AgentsAs large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a bench…
llm-agentsdeceptionbenchmarkevaluation
◆ FRONTIER·INDUSTRY IMPACT
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety AlignmentLarge Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment…
llm-safetyjailbreak-robustnesssafe-role-internalizationalignment-training
◆ FRONTIER·INDUSTRY IMPACT
The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language ModelsSafety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary…
backdoor-attacksllm-safetymulti-turn-dialogueanswer-side-trigger
◆ FRONTIER·INDUSTRY IMPACT
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic CodingAs coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents…
agentic-codingbenchmarksrisk-treatmentevaluation-framework
THE INDEX — 37 MORE CLEARED
●Behavioral Safety Assessment towards Large-scale Deployment of Autonomous Vehicles, Part I: Methodologycs.RO
●Preparing an AI-Augmented SIEM for the EU Cyber Resilience Act: A Practitioner Case Studycs.CR · cs.CY
●Auditable Claims about AI Agentscs.AI · cs.CL
●Toward Trustworthy Physical AI for Human Interactioncs.RO
●Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigationcs.RO · cs.AI
●Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changescs.LG · cs.CL
●APEX: Active Protection at Execution Boundaries for LLM Agentscs.CR · cs.AI
●UNREAL: Unifying Retrieval and Long-Context with a Single ModelNVIDIAcs.CL · cs.IR
●Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steeringcs.LG · cs.AI
●CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footagecs.CV
●AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generationcs.LG · cs.AI
●HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?cs.CR · cs.SE
●Incidental information contaminates patient notes and disrupts clinical reasoning in large language modelscs.CL
●Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Renderingcs.CR · cs.CV
●Surviving the Router: Optimizing Skill Injections for Retrieval and Executioncs.CR · cs.LG
●Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgerycs.CV
●Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluationcs.CV
●Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Drivingcs.CV · cs.AI
●BARE-AI: Bit-Flip Attack Resilience in AI Hardware through Built-in Performance Monitorscs.CR
●RoboCap: A New Platform for Egocentric Robot Learningcs.RO · cs.AI
●Efficient Auditing of Adversarial AI Agent Behavior from Agent Tracescs.CR
●Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AIcs.CL · cs.CR
●Behavioral Safety Assessment towards Large-scale Deployment of Autonomous Vehicles, Part II: Assessment Resultscs.RO
●Does On-Policy Distillation for Safety Pose Backdoor Risks?Amazoncs.LG · cs.AI
●Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulationcs.CL · cs.CY
●The Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessmentcs.CR · cs.CY
●MARCO: The Radioactive Watermark for Protein Generative Modelscs.CR · cs.AI
●Visual Abstention in Unified Multimodal Modelscs.CL · cs.AI
●OpenWAM: An Open Framework for Composable World-Action Modelscs.RO · cs.CV
●PhoneBot: A Low-Cost Open Humanoid Robot Platform Reusing Smartphonescs.RO
●Secure Speculative Decoding for Large Language Modelscs.CR · cs.AI
●SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Modelscs.SE · cs.AI
●Can Power Draw Constrain Covert Compute? Limits of Analogue Verification for AI Governancecs.CY · cs.AI
●Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMscs.CV · cs.AI
●Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updatescs.LG · cs.AI
●Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google ScaleGooglecs.SE · cs.AI
●SkillPoison: Progressive Skill Poisoning via Successful Experiencescs.CR · cs.AI
TUE 06 OCT 202683 cleared of 932 · 8 FRONTIER
◆ FRONTIER·INDUSTRY IMPACT
Backdooring Sparse AutoencodersSparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM a…
sparse-autoencodersbackdoor-attacksmodel-interventionsupply-chain-security
◆ FRONTIER·INDUSTRY IMPACT
Reward Stealing Attack on Large Language ModelsAdversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement…
reward-stealingadversarial-attacksllm-safetyinverse-reinforcement-learning
◆ FRONTIER·INDUSTRY IMPACT
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not EnoughMemory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it cannot yet act on to persistent memory, so that a future aligned agent carries it out when the opportunity arises. We investigate this threat, which we refer to as self-propa…
llm-agentsmemory-poisoningself-propagating-misalignmentagent-safety
◆ FRONTIER·INDUSTRY IMPACT
MLLMs Fail to Refuse when Using Tools AgenticallyAgentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all…
agentic-mlmtool-userefusal-failuresafety-benchmarks
◆ FRONTIER·INDUSTRY IMPACT
Target-free Latent Safety AlignmentLarge language models (LLMs) remain highly vulnerable to jailbreak attacks that induce harmful behaviors and circumvent safety alignment. To defend against such attacks, adversarial training paradigms have been proposed to first simulate failure modes and then train the model to correct them, yielding promising improvements in safety alignment. However, these methods typically construct adversarial samples either by…
jailbreak-defensesafety-alignmentadversarial-trainingtarget-free
◆ FRONTIER·INDUSTRY IMPACT
Better Call Reward: Reward Hacking as Strategic Abstention in Legal Reasoning ModelsWhat happens when a legal AI model learns to look like a lawyer instead of reasoning like one? We fine tune Qwen3-8B with Group Relative Policy Optimisation (GRPO) against a proxy built from three surface features: citation count, legalese density, and response length. The model does not learn to reason more effectively. It learns to withhold commitment. Across 16 yes or no legal reasoning tasks from LegalBench (N=32…
reward-hackingstrategic-abstentionlegal-reasoningrlhf
THE INDEX — 52 MORE CLEARED
●AgentDoxx: Agentic Re-identification of Anonymized Text with Web Searchcs.CR · cs.AI
●What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimescs.LG
●Mechanizing the User's Eye: Pre-Registered Deployment of a Sabotage-Validated Fail-Plausible Observer in a Production LLM Agent Runtimecs.SE
●Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOADcs.CR
●From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deploymentcs.CL · cs.DC
●Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Modelscs.LG · cs.AI
●UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agentscs.SE · cs.AI
●FORGE: Verification-Gated Behavioral Repair for Generative Language Modelscs.CL · cs.LG
●Tracing model-generated DNA with position-independent watermarkingq-bio.GN · cs.CR
●Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attackscs.CR · cs.AI
●Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validationcs.AI
●Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMscs.CV · cs.AI
●StegoMemory: Agentic Memory Acts as Covert Steganographic Channelcs.CR · cs.CL
●BabelFake: A Multilingual Audio-Visual DeepFake Benchmarkcs.CV
●Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Modelscs.CR · cs.LG
●H-CRSPV: Preventing Semantic Omission in Late-Bound Large Language Model Releasescs.CR
●Backdoors in Learning-Based Industrial Robotic Arm Manipulation: An Empirical Security StudyIBMcs.RO
●Lie Rarely, Lie Big: Stealthy Insider Attacks on LLM Robot Teamscs.RO · cs.MA
●Base Models Can Reason By Taking a Cue From Training DataAllenAIcs.LG · cs.AI
●Hidden Risks of Jev: An Empirical Study of Security, Privacy, and Dual Usecs.CR · cs.AI
●CLIFT: Conformal Self-Verification for Web Agent Training and Test-Time Scalingcs.CL · cs.AI
●Readable Before Actionable: Causal Tracing of Indirect Prompt Injectioncs.AI
●Certification of Real Images through Calibrated Content Authenticationcs.CV
●Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potentialcs.LG · cs.AI
●Self-Reflection Fine-Tuning: Enhancing Agent Security against Prompt Injection Attacks from Failure ExperienceNVIDIAcs.LG · cs.CR
●Control OSWorld: An AI Control Environment for GUI Computer Use Agentscs.CR
●Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-ProbabilitiesAmazoncs.AI · cs.CL
●Agent Reliability Profiles in Financial Servicescs.CY · cs.AI
●Red-TTT: Test-Time Training for Automated Jailbreaking Large Language Modelscs.CL · cs.CR
●EnvDreamer: Large-Scale Multimodal-to-Environment Generation for Embodied AIcs.AI
●Who Keeps the Gains from Personal AI Assistants? Seller Adaptation and the Unassisted in a Language-Model Market Simulationcs.MA · cs.AI
●What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Documentcs.CR · cs.CL
●Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinatecs.CR · cs.AI
●Measuring and Reducing Cross-Vendor Mismatch in Language Modelscs.LG · cs.DC
●MemLeak: Cross-User Semantic Leakage in Multi-Tenant AI Agent Memorycs.AI · cs.LG
●TimeNet: An Extensible Unified Data Infrastructure for Next-Generation Temporal Foundation ModelsGooglecs.AI · cs.LG
●IdeaLens: Detecting AI Ideas in Long-form WritingGooglecs.CL · cs.AI
●Safe Image Generation via Reinforcement Learningcs.CV
●Compromise Is Not Consequence: Evaluating Task-Scoped Authorization in LLM Agents with Paired Replaycs.CR
●What Does a Harness Buy? Tokens, Mostlycs.AI · cs.LG
●Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steeringcs.CL
●Cross-Lingual Transferability of Training Data Extraction Attacks to Recover Memorized PIIcs.CL · cs.AI
●Search Engines Never Say No: How Frozen Agents React When the Retrieval Tool Refusescs.IR
●Auditing the Privacy of Synthetic Gene Expression Data: A Unified Weighted-Distance Framework for No-Box Membership Inferencecs.CR
●Adaptive Co-Serving LLM Watermarking on Modern Inference EnginesIBMcs.CR
●TAPDreamer: Transferable Adversarial Patches for World Action Modelscs.CV · cs.AI
●ArtifactArena: Evaluating Models by What They Build in the Physical Worldcs.RO · cs.AI
●Hidden in the Comments: A Context-Injection Attack Surface in Code LLMscs.LG
●Quantifying Collusion Among Autonomous LLM Agents: A Statistical Analysis of the Collusion Wiki Incidentcs.MA · cs.AI
●From Probe Scores to Alarm Policies: Operational Validity of Activation Monitors for Language-Model AgentsAlibabacs.CL · cs.AI
●GrayShield: Bit-Level Sanitization for Transformer Model Supply-Chain Securitycs.CR · cs.AI
●Understanding and Mitigating Hallucination Escape in Tool-Using LLM Agentscs.CL · cs.AI
loading earlier editions…
THE FIELD — 18,135 · 30D
every indexed paper — the whole wire, not just what cleared
IN THE LITERATURE
models referenced in the loaded window · click to pivot
FROM INSIDE THE LABS
papers with a frontier-lab author · click to pivot
Research — every claim one click from the paperarXiv continuous index · FRONTIER ◆ leads each edition