VECTOR WIREAI INTELLIGENCE
PKT
Refresh Models Deals Regulatory Sources

Research

WED 26 AUG 2026: 28 in view. GPT-5 is the literature’s most-referenced model (5 papers in the loaded window).
UPDATED 1M AGO
WED 26 AUG 202628 in view
28 CLEARED
Markets, Not Planners: Decentralized Orchestration of LLM Agents with Private Informationcs.MA · cs.CL
StepGuard: Learning Step-Level Guardrails with Scalable Supervision and Safety-Utility Balancingcs.AI · cs.CR
What Guides the Agent? Adjudicating Unauthorized Behavior via Localizing Behavior-Guiding Instructionscs.CR · cs.AI
WebMCP-Phalanx: Enforcing and Characterizing Trust Boundaries for Browser-Integrated LLM Agentscs.CR · cs.AI
Semantic Overlays: Mitigating Prompt Injection with Annotations Beyond Tokens and Steering Vectorscs.AI · cs.CL
ESQ-Bench: A Multi-Tier Enterprise Oracle Benchmark for Evaluating NL2SQL Dialect Generalization and Silent Semantic Divergencecs.AI
Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systemscs.AI
Wontopos Tablet 2: Measuring Multilingual and Multimodal Memory Retrieval Without Lexical Matchingcs.IR · cs.CL
PinSieve: Production Selective VLM Serving and a Governed Memory Flywheel for Enterprise Content-Quality Triagecs.LG
Beyond the Mandate: A Systematic Security Analysis of the Agent Payments Protocol (AP2)cs.CR · cs.AI
Maia 200: A Software Defined Dataflow System for Large-scale AI Accelerationcs.AR · cs.AI
Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Modelcs.MA
Minima-KV: Retention-Preserving KV Cache Compression with Mixed-Format Paged Attentioncs.AI
Adoption Telemetry: Measuring Enterprise AI Adoption from Production Signalscs.HC · cs.CY
More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Servingcs.AI
Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Modelcs.SE · cs.AI
Paritok-4B: Intent-Conditioned Context Compression for Coding Agentscs.AI · cs.CL
Names Can Hurt: Spotting Slopsquatting Risks Caused by Package Name Hallucinations in Local Coding LLMscs.CL · cs.AI
Knowing When to Ask for Help: Bayesian Self-Escalation in Hierarchical LLM Agentscs.LG · cs.AI
NVIDIA Cosmos-H-Dreams: Real-Time Generative Physics Simulation for Surgical RoboticsNVIDIAcs.RO
Granite.Trust Policy Tools: Shareable, Actionable Policies for Generative AI ApplicationsIBMcs.AI
The Shadow Price of Intelligence: Quality Degradation in LLM Inference as a Supply Chain Problemmath.OC · cs.AI
Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probescs.CL
Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading SystemsGooglecs.AI · cs.CE
TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Serverscs.CR · cs.AI
'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detectioncs.CL · cs.AI
Neurosymbolic Alignment for Physiologically-Safe Clinical Language Modelscs.AI
NeuronGuard: Robust LLM Safety Alignment via Ablation-Aware Safety Signal Redistributioncs.CR · cs.AI
TUE 25 AUG 202668 in view
◆ FRONTIER·INDUSTRY IMPACT
OmniCAD: A Large-Scale Benchmark for 3D Spatial Reasoning in Robotics Assemblies
Recent vision-language models (VLMs) show strong capabilities in robotic perception and spatial reasoning, yet their ability to reason about complex mechanical assemblies remains underexplored. We introduce OmniCAD, a large-scale benchmark for assembly-aware 3D spatial reasoning across diverse industrial systems, including robotic mechanisms, automotive components, aerospace structures, and agricultural machinery. Om…
3d-spatial-reasoningrobotics-assembliesbenchmarkvision-language-models
MicrosoftMingjia Wang, Taiting Lu, Ziwei Dong, et al. (10)cs.CVarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
Reinforcement Learning on Benign Facts Amplifies Leakage of Memorized Private Data
Reinforcement learning with verifiable rewards (RLVR) is deployed to make models better at reasoning tasks, but its side effect on what models will divulge is under studied. Here we show that RLVR on facts increases extraction of personally identifiable information (PII) the instruct model had already memorized. We first confirm that instruct models have already memorized PII but leave them latent, rarely surfacing o…
rlvrdata-leakagememorizationpii-extraction
Renfei Zhang, Niloofar Mireshghallahcs.LG · cs.AIarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
Apodex 1.1: Scaling Agentic Intelligence for Complex Work
General-purpose language models can reason and synthesize knowledge, but complex work also requires sustained interaction with files, information sources, and executable code, together with state maintenance, failure recovery, and verifiable delivery. We call this \emph{working capability}: sustained, verifiable progress toward a real-world objective. Apodex 1.1 develops this capability along two complementary dimens…
agentic-aiworking-capabilityenvironment-scalingverifiable-execution
Apodex Team, B. An, B. Li, et al. (10)cs.AI · cs.CL · cs.LGarXiv ↗PDF ↗
THE INDEX — 65 MORE CLEARED
Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Modelscs.CL · cs.AI
Mitigating Reasoning-Induced Misalignment via Safety-Direction Penaltycs.AI · cs.CL
GuardPaint:SpeculativeSafetyDecodingforText-to-ImageGenerationMetacs.CV · cs.AI
What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Enginecs.LG · cs.AR
PURA: Provably Unbiased and Robust Multi-Bit Text Attributioncs.CR
Redteaming Leading Arabic LLMs with ASAScs.AI
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillationcs.CR · cs.AI
KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model Inferencecs.AI · cs.DC
Counterfactual, Per-Decision Bias Auditing for Automated Hiring: Localizing and Explaining Disparate Impact in Applicant Tracking Systemscs.CY · cs.CR
AraDetox: A Multi-Dialect Arabic Detoxification Datasetcs.CL · cs.AI
ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Modelscs.CR · cs.AI
K-Bench: measuring model performance on real scientific agent requestscs.AI · cs.CL
AdaptPrint: Response-Adaptive Fingerprinting of Black-Box LLM Servicescs.CR
Agentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration Testingcs.CL · cs.AI
The Compaction Cliff in Long-Running AI Agent Memorycs.AI · cs.IR
DIME: Query-Efficient Framework for Membership Inference on Diffusion Modelscs.LG · cs.CR
Meta-Ctrl: Guaranteed Plan Generation by Decoupling Syntactic and Semantic Constraintscs.RO · cs.AI
AgentFlow: A Flow-Centric Policy Language and Framework for Securing LLM Agent Systemscs.CR
VISTA: Test-Time Compositional Alignment for Visual Autoregressive Generationcs.CV
Breaking the Assumptions: Auditing Input-Side Jailbreak Defenses Against Semantic Attackscs.CR · cs.AI
PsychJail: Exploring Psychological Jailbreaks via Multi-Turn Persuasion of LLM Policiescs.AI
Anchoring Bias: A Persistent Fairness Backdoor Attack against MLLMs under Continual Learningcs.LG · cs.AI
Mitigating Explanation Leakage in Financial Fraud Detection Systemscs.LG · cs.CR
SkillBloat: Token Amplification Attacks via Skill Injection in LLM Coding Agentscs.CR · cs.CL
MEMORY Wins All: Indirect Bias Injection Attacks via Social Media Feedscs.AI · cs.CY
Beyond Over-Refusal: Defending Indirect Prompt Injection via Latent Instruction Manifoldscs.CR
Text-Anchored Semantic Perturbations for Transferable Jailbreak Attacks on Multimodal Large Language Modelscs.CL
Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learningcs.AI
AUDITA: certified auditing and causal attribution of adverse outcomes in autonomous multi-agent systemscs.AI
Forgotten in Weights, Recovered by Tools: Agentic Tool Unlearning for LLM Agentscs.CL · cs.AI
BLADE: Bilevel Low-rank Augmented-Lagrangian Erasure for LLM Unlearningcs.LG · cs.AI
Adversarial Entropy Inflation Against Gumbel-Based Inference Verificationcs.CR · cs.AI
MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governancecs.AI · cs.CL
There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Itemscs.AI
LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimizationcs.AI
RAG Collapse: LLM Responses Collapse When Retrieved Documents Are Self-Authoredcs.CL · cs.IR
AIREP: A Protocol for Per-Decision Evidence in AI Runtime Governancecs.AI · cs.CR
CIDER: Continual Interactive Distillation for Embodied Reinforcement Learningcs.RO
Mitigating Database Leakage in RAG Systems with Keyword-Grounded Fact Substitutioncs.CL
Enhancing User Resilience Against AI-Augmented Phishing: A Two-Stage Framework for Detection and Personalized Trainingcs.CR
AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversationscs.AI · cs.HC
Evaluating Inference-Time Defenses Against Package Hallucination in LLM-Generated Codecs.SE · cs.AI
Evaluation Awareness in Language Models: Representation, Verbalization, and Controlcs.CL · cs.AI
What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideationcs.CL · cs.AI
Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulationcs.RO · cs.AI
A-CPES: A Reference Framework for Agentic AI in Cyber-Physical Energy Systemscs.AI
Measuring Activation Control in Large Language Modelscs.AI · cs.CL
Proxy reliance in large language model decisions is uncalibrated to predictive evidencecs.AI · cs.CL
Beyond Two Bytes per Letter: Tokenization Overhead in Cyrillic AI Systemscs.CL · cs.AI
Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based ModerationTencentcs.CL
TessIndex: Capability Verified Identity System for the Agent Economycs.AI · cs.CR
InjecMEM: Memory Injection Attack on LLM Agent Memory Systemscs.CR · cs.AI
Unlearning Is Not Just Erasing: Temporal Decoupling via Generation Inequalitycs.CL · cs.AI
No One Model Catches Every Harm: Benchmarking Content Moderation Across Safety Scenarioscs.CL
Hack-Verifiable Terminal Bench: Evaluating Reward Hacking in Terminal Taskscs.AI
GOLEM: Modular Humanoid Autonomy Towards Electric Vehicle Battery Disassemblycs.RO
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agentscs.CV · cs.AI
First Demonstration of Multi-Agent LLM System for Million-Scale Optical Link Management in Global Production AIDCsBaiducs.MA · physics.optics
RoboShape: Information-Theoretic Point Cloud Representations for Privacy-Aware Robot Perceptioncs.RO · cs.AI
Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lenscs.CY · cs.AI
Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalfcs.AI · econ.GN
Reviewing Model Collapse and Countermeasurescs.AI · cs.LG
Register Shifts Break LLM Safety: A Bengali Benchmark with Culturally Grounded Harmscs.CL · cs.AI
ST²U: Stateful Test-Time Unlearning via Restricted Knowledge Boundary Controlcs.LG · cs.CL
From Inertia to Objectivity: Improving Deep Research Agents with Noise IsolationAlibabacs.AI
MON 24 AUG 20264 in view
4 CLEARED
AID-Guard: Stateful Authorization for Delegated Agent Effectscs.CR · cs.AI
AEGIS: Preventing Cross-Domain Resource Abuse in MCPIBMcs.CR · cs.AI
Who Delegates to AI? Evidence from 53,000 Agent Configurationscs.AI · cs.CY
Masking Is Not Enough: Generative Restoration for Multimodal De-Identification in Medical AIcs.CV · cs.CR
loading earlier editions…
THE FIELD — 12,662 · 30D
cs.AI3,902
cs.LG2,452
cs.CL1,800
cs.CV1,762
cs.RO896
cs.SE529
every indexed paper — the whole wire, not just what cleared
IN THE LITERATURE
models referenced in the loaded window · click to pivot
FROM INSIDE THE LABS
papers with a frontier-lab author · click to pivot
Research — every claim one click from the paperarXiv continuous index · FRONTIER ◆ leads each edition