Skip to content
VECTOR WIREAI INTELLIGENCE
UTC

AI Research Radar

Today: 34 papers cleared of 532 scanned, led by UltraText Bench. Qwen3 is the literature’s most-referenced model (2 papers in the loaded window).
UPDATED 1M AGO
MON 10 AUG — 280 scanned · 16 cleared · 0 frontier — click to open this editionTUE 11 AUG — 747 scanned · 53 cleared · 6 frontier — click to open this editionWED 12 AUG — 349 scanned · 31 cleared · 1 frontier — click to open this editionTHU 13 AUG — 303 scanned · 36 cleared · 3 frontier — click to open this editionFRI 14 AUG — 205 scanned · 22 cleared · 1 frontier — click to open this editionSAT 15 AUG — 90 scanned · 12 cleared · 1 frontier — click to open this editionMON 17 AUG — 269 scanned · 20 cleared · 2 frontier — click to open this editionTUE 18 AUG — 702 scanned · 58 cleared · 2 frontier — click to open this editionWED 19 AUG — 311 scanned · 31 cleared · 2 frontier — click to open this editionTHU 20 AUG — 297 scanned · 18 cleared · 1 frontier — click to open this editionFRI 21 AUG — 215 scanned · 18 cleared · 1 frontier — click to open this editionMON 24 AUG — 299 scanned · 34 cleared · 1 frontier — click to open this editionTUE 25 AUG — 636 scanned · 68 cleared · 3 frontier — click to open this editionWED 26 AUG — 345 scanned · 28 cleared · 0 frontier — click to open this editionTHU 27 AUG — 344 scanned · 27 cleared · 2 frontier — click to open this editionFRI 28 AUG — 407 scanned · 51 cleared · 3 frontier — click to open this editionMON 31 AUG — 296 scanned · 24 cleared · 1 frontier — click to open this editionTUE 01 SEP — 773 scanned · 73 cleared · 5 frontier — click to open this editionWED 02 SEP — 463 scanned · 41 cleared · 2 frontier — click to open this editionTHU 03 SEP — 322 scanned · 37 cleared · 2 frontier — click to open this editionFRI 04 SEP — 326 scanned · 20 cleared · 1 frontier — click to open this editionMON 07 SEP — 348 scanned · 37 cleared · 3 frontier — click to open this editionWED 09 SEP — 920 scanned · 80 cleared · 7 frontier — click to open this editionTHU 10 SEP — 296 scanned · 30 cleared · 1 frontier — click to open this editionFRI 11 SEP — 265 scanned · 36 cleared · 2 frontier — click to open this editionSAT 12 SEP — 46 scanned · 5 cleared · 1 frontier — click to open this editionMON 14 SEP — 295 scanned · 18 cleared · 0 frontier — click to open this editionTUE 15 SEP — 709 scanned · 77 cleared · 7 frontier — click to open this editionWED 16 SEP — 361 scanned · 38 cleared · 1 frontier — click to open this editionTHU 17 SEP — 423 scanned · 34 cleared · 3 frontier — click to open this editionFRI 18 SEP — 425 scanned · 32 cleared · 3 frontier — click to open this editionMON 21 SEP — 327 scanned · 31 cleared · 1 frontier — click to open this editionTUE 22 SEP — 797 scanned · 62 cleared · 3 frontier — click to open this editionWED 23 SEP — 451 scanned · 31 cleared · 4 frontier — click to open this editionTHU 24 SEP — 330 scanned · 43 cleared · 5 frontier — click to open this editionFRI 25 SEP — 448 scanned · 50 cleared · 5 frontier — click to open this editionMON 28 SEP — 385 scanned · 37 cleared · 2 frontier — click to open this editionTUE 29 SEP — 1,559 scanned · 109 cleared · 7 frontier — click to open this editionWED 30 SEP — 824 scanned · 93 cleared · 4 frontier — click to open this editionTHU 01 OCT — 669 scanned · 77 cleared · 5 frontier — click to open this editionFRI 02 OCT — 673 scanned · 57 cleared · 1 frontier — click to open this editionMON 05 OCT — 441 scanned · 51 cleared · 1 frontier — click to open this editionTUE 06 OCT — 932 scanned · 83 cleared · 8 frontier — click to open this editionWED 07 OCT — 573 scanned · 49 cleared · 6 frontier — click to open this editionTHU 08 OCT — 532 scanned · 34 cleared · 3 frontier — click to open this editionAUGSEPOCT
60 DAYS · 21,008 SCANNED
1,912 CLEARED · 123 ◆
TODAY — THU 08 OCT 202634 cleared of 532 · 3 FRONTIER
◆ FRONTIER·INDUSTRY IMPACT
UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation
Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation of dense visual text. It contains 432 prompts spanning 24 real-world scene categories and three dif…
image-generationvisual-text-renderingbenchmarkbilingual-evaluation
Deyuan Liu, Yihao Hu, Jingxuan Zhang, et al. (10)cs.CV · cs.AIarXiv ↗PDF ↗Code ↗
◆ FRONTIER·INDUSTRY IMPACT
ASPIRE: Agentic Safety & Prompt Injection Red-teaming Engine
LLM agents retrieve untrusted content and act through tools, creating indirect prompt-injection risks that can cause unauthorized actions or persistent state changes. Existing automated red-teaming largely optimizes payloads for pre-specified scenarios, leaving latent vulnerabilities across the agent's behavior space unexplored. We present ASPIRE, an Agentic Safety & Prompt Injection Red-teaming Engine for open-ended…
prompt-injectionagentic-safetyred-teamingllm-agents
GooglePengfei He, Deep Mitra, Vishesh Sharma, et al. (7)cs.CRarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, loc…
robotics-benchmarkmultimodal-agentsembodiment-diversityphysical-control
Zhiqin Yang, Chenxin Li, Xiaomeng Hu, et al. (10)cs.RO · cs.LGarXiv ↗PDF ↗
THE INDEX — 31 MORE CLEARED
●Robust Decentralized Fairness Auditingcs.LG
●Breaking Adversarial Transferability in Fine-Tuned Speech Recognitioncs.LG
●Purifying Backdoored Large Vision-Language Models by Removing Hijacked Directionscs.CV
●Comprehension Audits to Mitigate Risks from Automated AI Researchcs.CY · cs.AI
●AgentTracer: Tracing Indirect Prompt Injection Attack through Fine-Grained Intention-Execution Alignmentcs.SE
●The Handover Problem: Governing Autonomy Transitions in Human-AI Collaborationcs.AI · cs.HC
●The AI Evaluation Ecosystemcs.AI · cs.CY
●Understanding and Mitigating Token-Pruning-Induced Vulnerabilities in VLMscs.CR · cs.LG
●PatchBench: Measuring Collateral Damage in Activation Patchingcs.LG · cs.CL
●On the Reliability of LLM-Based Vulnerability Patching Benchmarkscs.CR · cs.SE
●Bounded Autonomy and Verifiable Safety for Agentic AI Enabled Automationcs.LG · cs.ET
●SwarmReconGuard: Black-Box Detection of Distributed Collective Reconnaissance by Individually Benign-Looking Agent Populationscs.CR · cs.AI
●Shared and structured inputs undermine collective random choice by reasoning AI agentscs.AI · physics.soc-ph
●Quad-State Safety Evaluation of Open-Weight Large Language Models on Non-Canonical Inputscs.CL · cs.AI
●Multi-Aspect Runtime Verification for Simulation-Based V&V of LLM-Enabled Autonomous Agentscs.CR · cs.AI
●EASE: Entropy-Adaptive Distribution Shaping for Evading AI-generated Text Detectorscs.CL
●Package Hallucination Attacks on Coding Agents through Prompt Injection in Rule Filescs.CR · cs.AI
●Visual Memory Attacks Can Persist Through The KV Cachecs.CR · cs.AI
●Inverting Multi-Vector Visual Document Indicescs.IR · cs.CL
●The Confidence Game: Strategic Miscalibration in Human-AI Delegationcs.GT · cs.AI
●SLDR: Defending Against Malicious Fine-tuning via Selective Layers Recovery and Dynamic Routingcs.CR · cs.AI
●Secure-CUA: Controlling Untrusted Influence in Computer-Use AgentsGooglecs.CR · cs.AI
●Black-Box Adversarial Patch Attacks on VLAs via Ancestor VLM Exploitationcs.CR
●LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Marketscs.AI · cs.CL
●Fault-tolerant foundation modelscs.LG · cs.AI
●AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrailscs.AI
●When Rank Rises as LLMs Degradecs.LG · cs.CL
●Adversarial Images Hijack Web Agents from Visual Grounding to Browser Executioncs.CR · cs.AI
●RoboJEPA: Scaling Robotic Latent World ModelsMetacs.AI · cs.RO
●Sensitive-Topic Leakage Through LLM Routing Metadata: Measurement and Mitigationcs.CR · cs.CL
●Backdooring Acoustic Foundation Models for Physically Realizable Triggerscs.SD · cs.LG
WED 07 OCT 202649 cleared of 573 · 6 FRONTIER
◆ FRONTIER·INDUSTRY IMPACT
Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level re…
backdoor-persistencyllm-agentssftreinforcement-learning
Qiusi Zhan, Nian Lyu, Stephanie Ding, et al. (6)cs.CRarXiv ↗PDF ↗Code ↗
◆ FRONTIER·INDUSTRY IMPACT
ScienceClaw: Benchmarking Continual Self-Evolution of AI-for-Science Agents Across the Natural and Social Sciences
Large language model agents are accelerating scientific automation, yet verified executions rarely become persistent program-level improvements, and existing evaluations do not examine this process across sequential tasks in both the natural and social sciences. We formalize ScienceClaw as fixed-parameter program self-evolution that unifies task solving, scientific verification, and program updates. ScienceClaw-Eval…
ai-for-sciencecontinual-learningagent-evaluationprogram-self-evolution
Mingda Zhang, Wenjin Liu, Tiesunlong Shen, et al. (9)cs.AIarXiv ↗PDF ↗Code ↗
◆ FRONTIER·INDUSTRY IMPACT
DecepEval: A Benchmark for Evaluating Deception in LLM Agents
As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a bench…
llm-agentsdeceptionbenchmarkevaluation
Yiming Xu, Hongyue Yu, Beihua Yang, et al. (10)cs.LGarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment…
llm-safetyjailbreak-robustnesssafe-role-internalizationalignment-training
Jinghao Pang, Jitai Hao, Qiang Huang, et al. (5)cs.AI · cs.CL · cs.IRarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary…
backdoor-attacksllm-safetymulti-turn-dialogueanswer-side-trigger
Yibo Zhang, Tianrong Guan, Liang Lin, et al. (6)cs.CR · cs.LGarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
ParanoiaEval: Benchmarking Unnecessary Defensive Work in Agentic Coding
As coding agents increasingly undertake real-world work autonomously, judging whether their risk treatments are warranted has become important. Existing work evaluates related agent behaviors from separate perspectives, but lacks a systematic framework for unifying these behaviors. To bridge this gap, we introduce ParanoiaEval, the first benchmark for unified evaluation of risk-treatment capabilities in coding agents…
agentic-codingbenchmarksrisk-treatmentevaluation-framework
Hanjun Luo, Xiucheng Zhang, Zhuoning Xu, et al. (7)cs.AI · cs.SEarXiv ↗PDF ↗
THE INDEX — 43 MORE CLEARED
●Behavioral Safety Assessment towards Large-scale Deployment of Autonomous Vehicles, Part I: Methodologycs.RO
●Preparing an AI-Augmented SIEM for the EU Cyber Resilience Act: A Practitioner Case Studycs.CR · cs.CY
●Verifying Coordination in Parallel Coding Agents: NP-Bench and a Scheduling Plannercs.AI
●Auditable Claims about AI Agentscs.AI · cs.CL
●Toward Trustworthy Physical AI for Human Interactioncs.RO
●Seeing the Invisible: Physics-Guided Visual Prompting for Temperature- and Radiation-Aware VLA Navigationcs.RO · cs.AI
●Lost in the bf16 Cast: Exporting Ternary Language Models Can Revert Most Low-Learning-Rate Code Changescs.LG · cs.CL
●APEX: Active Protection at Execution Boundaries for LLM Agentscs.CR · cs.AI
●UNREAL: Unifying Retrieval and Long-Context with a Single ModelNVIDIAcs.CL · cs.IR
●Learning to Decide, Not to Reason: Parameter-Efficient Decision Operators via Low-Rank Activation Steeringcs.LG · cs.AI
●CCDF: A Benchmark Dataset for Deepfake Detection in Real-World Surveillance Footagecs.CV
●AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generationcs.LG · cs.AI
●A Systematic Investigation of Bias in Large Language Models for Advertising RelevanceMicrosoftcs.AI
●HarnessSecurity-Bench: Do Security Mechanisms Really Protect Coding Agent Harnesses?cs.CR · cs.SE
●Incidental information contaminates patient notes and disrupts clinical reasoning in large language modelscs.CL
●Rethinking Visual Provenance: Detection and Watermarking Across Direct Visual Generation and LLM-Driven Code Renderingcs.CR · cs.CV
●Surviving the Router: Optimizing Skill Injections for Retrieval and Executioncs.CR · cs.LG
●Compositional Concept Erasure in Text-to-Image Diffusion Models via Hierarchically Grounded Semantic Surgerycs.CV
●Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluationcs.CV
●Transferable Spatial Temporal Coherence Adversarial Attack on Black-Box Vision Language Models for Autonomous Drivingcs.CV · cs.AI
●BARE-AI: Bit-Flip Attack Resilience in AI Hardware through Built-in Performance Monitorscs.CR
●RoboCap: A New Platform for Egocentric Robot Learningcs.RO · cs.AI
●Efficient Auditing of Adversarial AI Agent Behavior from Agent Tracescs.CR
●Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AIcs.CL · cs.CR
●Behavioral Safety Assessment towards Large-scale Deployment of Autonomous Vehicles, Part II: Assessment Resultscs.RO
●Does On-Policy Distillation for Safety Pose Backdoor Risks?Amazoncs.LG · cs.AI
●Detecting LLM-Assisted Vietnamese Writing via Keystrokes under Behavioral Manipulationcs.CL · cs.CY
●The Amplifier Effect: Human-Factor Risks of AI-Suggested Correlation and Auto-Propagation in Multi-Framework GRC Self-Assessmentcs.CR · cs.CY
●MARCO: The Radioactive Watermark for Protein Generative Modelscs.CR · cs.AI
●Visual Abstention in Unified Multimodal Modelscs.CL · cs.AI
●OpenWAM: An Open Framework for Composable World-Action Modelscs.RO · cs.CV
●PhoneBot: A Low-Cost Open Humanoid Robot Platform Reusing Smartphonescs.RO
●Secure Speculative Decoding for Large Language Modelscs.CR · cs.AI
●zkLLMPoT: Efficient Zero Knowledge Proof of Training for Large Language Modelscs.AI
●SAFESHIELD: A Decision-Organization Framework for Deployment-Time Safety of Small Language Modelscs.SE · cs.AI
●Can Power Draw Constrain Covert Compute? Limits of Analogue Verification for AI Governancecs.CY · cs.AI
●Visual-Invariance-Augmented Feature Optimal Alignment for Transferable Adversarial Attacks against Closed-Source MLLMscs.CV · cs.AI
●SIGMA: Self-Improving Alignment Generalization from a Model Speccs.AI
●Harmful SFT Leaves a Continuous Trace in LLM Checkpoint Updatescs.LG · cs.AI
●Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google ScaleGooglecs.SE · cs.AI
●SkillPoison: Progressive Skill Poisoning via Successful Experiencescs.CR · cs.AI
●OTel: Open Telco AI Datasets, Benchmarks, and Modelscs.AI
●Cascadia: Resident 975B MoE Inference on Eleven AI PCscs.AI · cs.DC
TUE 06 OCT 202683 cleared of 932 · 8 FRONTIER
◆ FRONTIER·INDUSTRY IMPACT
Backdooring Sparse Autoencoders
Sparse autoencoders (SAEs) are increasingly used not only to interpret language models but also to intervene on their internal representations. We show that this creates a supply-chain attack surface: a maliciously modified SAE can induce attacker-chosen behavior when inserted into the forward pass of an otherwise unchanged language model. We introduce a decoder-only SAE backdoor that leaves both the underlying LLM a…
sparse-autoencodersbackdoor-attacksmodel-interventionsupply-chain-security
Enrico Ahlers, Daniel Passon, Tobias Kiecker, et al. (5)cs.CR · cs.AI · cs.CLarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
Reward Stealing Attack on Large Language Models
Adversarial attacks on Large Language Models (LLMs) aim to induce harmful content. However, existing methods suffer from high computational costs or strict model-pairing dependencies, limiting their scalability and transferability. We propose Reward Stealing Attack (ReSA), an adversarial attack framework that targets the latent safety reward underlying LLM alignment. ReSA employs maximum entropy inverse reinforcement…
reward-stealingadversarial-attacksllm-safetyinverse-reinforcement-learning
Jiaming Qian, Pengyang Zhou, Jiahe Xu, et al. (4)cs.CLarXiv ↗PDF ↗Code ↗
THE INDEX — 15 MORE CLEARED
●AgentDoxx: Agentic Re-identification of Anonymized Text with Web Searchcs.CR · cs.AI
●What May an Agent Change About Itself? A Containment Floor for Self-Configuring Agent Runtimescs.LG
●Mechanizing the User's Eye: Pre-Registered Deployment of a Sabotage-Validated Fail-Plausible Observer in a Production LLM Agent Runtimecs.SE
●Autonomous Active Directory Exploitation via Multi-Model Harness Orchestration: A Benchmark Study with NeuroSploit on GOADcs.CR
●From Overloaded to Guaranteed: High-Throughput Multi-SLO Enforcement for LoRA-Assisted On-Premise LLM Deploymentcs.CL · cs.DC
●Your Unlearning Gives You Away: Identifying Erased Concepts in Diffusion Modelscs.LG · cs.AI
●UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agentscs.SE · cs.AI
●FORGE: Verification-Gated Behavioral Repair for Generative Language Modelscs.CL · cs.LG
●Tracing model-generated DNA with position-independent watermarkingq-bio.GN · cs.CR
●Does AI Help Cyber Attackers or Defenders? Evidence from Nonpublic Vulnerabilities and Subsequent Attackscs.CR · cs.AI
●Measurement-First Auditing of Agentic Leaderboards: Contamination Susceptibility, Matched-Control Re-evaluation, and Scorer Validationcs.AI
●Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMscs.CV · cs.AI
●StegoMemory: Agentic Memory Acts as Covert Steganographic Channelcs.CR · cs.CL
●BabelFake: A Multilingual Audio-Visual DeepFake Benchmarkcs.CV
●Localize-and-Detect: Auditing Task-Level Poisoning in Instruction-Tuned Modelscs.CR · cs.LG
loading earlier editions…
THE FIELD — 19,032 · 30D
cs.AI5,273
cs.LG4,307
cs.CL2,651
cs.CV2,313
cs.RO2,052
cs.CR808
every indexed paper — the whole wire, not just what cleared
IN THE LITERATURE
models referenced in the loaded window · click to pivot
FROM INSIDE THE LABS
papers with a frontier-lab author · click to pivot
Research — every claim one click from the paperarXiv continuous index · FRONTIER ◆ leads each edition