VECTOR WIREAI INTELLIGENCE
UTC
Refresh Models Deals Regulatory Sources

Research

TODAY — WED 02 SEP 2026: 31 in view, led by Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning. Qwen3 is the literature’s most-referenced model (4 papers in the loaded window).
UPDATED 1M AGO
TODAY — WED 02 SEP 202631 in view
◆ FRONTIER·INDUSTRY IMPACT
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. These dimensions are commonly scalarized with fixed aggregation weights. We identify a failure mode in which aggregation itself induces reward hacking: static projection al…
reward-hackingmulti-reward-rlrlhfreward-aggregation
Yu Yuan, Yaoyou Fan, Lili Zhao, et al. (8)cs.CLarXiv ↗PDF ↗Code ↗
◆ FRONTIER·INDUSTRY IMPACT
RePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem Solving
Data contamination undermines the reliable evaluation of large language models (LLMs) on mathematical problem solving. While rewriting-based evaluation mitigates memorization, existing methods lack guarantees of problem validity and answer correctness. We propose Proof-Verified Benchmark Rewriting (RePro), the first framework to integrate Lean-oriented neural automated theorem provers (ATPs) into benchmark rewriting,…
llm-evaluationbenchmark-rewritingproof-verificationlean-atp
Xiyuan Zhou, Zhuoqi Li, Xinlei Wang, et al. (9)cs.CL · cs.AIarXiv ↗PDF ↗Code ↗
THE INDEX — 29 MORE CLEARED
AKRASIA: Stealthy Backdoor Attack on Reasoning-based Code LLMscs.CR
Detecting Hidden Behaviors in LLMs via Activation-matched Finetuningcs.CL · cs.AI
What's in Your Agent's Context? Context Privilege Escalation Attacks against AI Agent Harnesscs.CR
Instella-MoE Technical Reportcs.CL · cs.AI
AI Morbidity and Mortality: A Framework for Clinical AI Failure Reviewcs.AI · cs.HC
Membership Inference in Fine-tuned Diffusion Language Models via Token-level Memorization Asymmetrycs.CL · cs.CR
Prediction-Assisted Pricing and Admission for LLM APIs with Stochastic Token Consumptioncs.DS · cs.LG
Validity-Aware Jailbreak Evaluation for Large Language Modelscs.AI
The Privacy-Hallucination Tradeoff in Differentially Private Language Modelscs.AI · cs.CL
ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-trainingcs.CV
CopyShield: A Cross-Level Benchmark of Copyright Defenses in LLMscs.LG
When Safety Routing Breaks: Understanding Alignment Fragility under Benign Fine-Tuningcs.CR · cs.AI
Autoresearch for Marketplace Catalogs: From Legacy Forms to AI-Native Matchingcs.AI
LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmarkcs.AI · cs.CL
Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascadescs.AI · cs.CR
EvoFlint: An Evolutionary Atlas of Multi-Turn LLM Vulnerabilitiescs.CL · cs.AI
When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluationcs.AI
The Irreversibility Budget: Fleet-Level Risk Accounting and Admission Control for Agent Operating Systemscs.AI
Spawn Freely, Act Sparingly: Progressive Risk Vesting for Recursive LLM-Agent Treescs.AI · cs.LG
The Constitutional Coverage Trilemma in AI Governancecs.LG · cs.AI
TRIS: A Tri-Layer Retrieval Integrity Sieve Against Knowledge Poisoningcs.CL · cs.CR
Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Modelscs.AI · cs.CL
Distributed Implicit Harm: A Compositional Safety Blind Spot in MLLM-Based Video Moderationcs.CV · cs.AI
Forbid Your Attention: Fooling Multimodal Large Language Models by Selectively Removing Intrinsic Focus in Spectral Domaincs.CV
LatentPress: Context Compression Beyond Text and Visioncs.LG · cs.AI
From Detection to Refusal: Safer LLMs via Circuit-Guided Weight Scalingcs.CL · cs.AI
Effective Interventions Against AI-Enhanced Scamscs.CR · cs.CY
Causal Evidentiary Governance for High-Risk Machine Learning Systemscs.CY · cs.AI
Jailbreaking Text-to-Image Models Through Cracks: Navigating Heterogeneous Safety Filters via Multi-Agent Debatecs.AI · cs.MM
TUE 01 SEP 202661 in view
◆ FRONTIER
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers
Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks. We present a provably correct, transformer parameterization (with only 280 learnable parameters for Boolean algebra tasks) capable of learning and evaluating problems of any depth or length. We assume inputs are fully parenthesized, well-fo…
algorithmic-generalizationtransformerscircuit-computationlength-generalization
IBMTakuya Ito, Ruchir Puri, Murray Campbell, et al. (4)cs.LGarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
A Causal Model for Locating and Unlocking Sandbagging in Model Organisms
Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later lay…
sandbaggingcausal-mechanisticresidual-streammodel-governance
Hong Kiat Tan, Linh Le, David Williams-Kingcs.LGarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation
Long-horizon agentic tasks go beyond chaining short tasks over more interaction turns. Their evolving dynamic environments and long-range dependencies require Large Language Models (LLMs) to continually explore, learn from experience, and adapt their policies over thousands of steps. We introduce E-Commerce Bench, the first open-source benchmark that integrates multi-round counterpart negotiation and dynamic events i…
long-horizonllm-agentsbenchmarkautonomous-ecommerce
AlibabaWei Fan, Xinjie Shen, Xudong Guo, et al. (10)cs.LG · cs.CLarXiv ↗PDF ↗Code ↗
◆ FRONTIER·INDUSTRY IMPACT
BAITBENCH: Measuring Agent Reward Hacking with Optional Shortcuts Planted in ML Tasks
LLM agents are increasingly used to run autonomous ML experiments, iterating on target metrics with little human oversight. Prior work has documented reward hacking in these environments, bringing into question the validity of produced research and the broader safety case for AI R&D. Existing benchmarks do not measure exploits that live in the data or the modeling task itself. We introduce BAITBENCH, a suite of three…
reward-hackingagent-evaluationbenchmarksautonomous-ml
Pradyumna Shyama Prasad, Meiri Anto, Leon Eshuijs, et al. (6)cs.LG · cs.AIarXiv ↗PDF ↗
◆ FRONTIER·INDUSTRY IMPACT
A.X K2 Technical Report
We introduce A.X K2, a 688B-parameter Mixture-of-Experts (MoE) language model trained from scratch as a high-performance foundation for \emph{agentic} applications. Trained on approximately 8.5T tokens---fewer than its predecessor, A.X K1---on a smaller but higher-quality mixture with substantially expanded agentic and software-engineering data, it nonetheless improves over A.X K1 across the board, by over 30 percent…
mixture-of-expertsfoundation-modelagentic-aitoken-efficiency
Cheolseung Baek, Dhammiko Arya, Eunki Kim, et al. (10)cs.AI · cs.CLarXiv ↗PDF ↗
THE INDEX — 56 MORE CLEARED
Authority-Inference Separation in Agentic Finance: First-Line Control, Blockchain Enforcement, and Replayable Assuranceq-fin.GN · cs.CR
Efficient GPU Retrieval for Semantic Searchcs.AI · cs.LG
SingProbe Technical Reportcs.CR · cs.AI
Verification-Time Dependency on a Disappearing Evaluatorcs.CY · cs.SE
CogEvol: Towards Efficient and Reliable Learning Environment Generationcs.CL · cs.AI
CHASE: How Content Ecosystems Are Reshaped When Ranking Is the Only Targetcs.AI · cs.IR
CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understandingcs.RO
Influence Is Not Authority: When Causal Guardrail Signals Make Legitimate Tool Use Look Like an Attack in Tool-Using LLM Agentscs.CR · cs.AI
Beyond Token-Level Guidance: Inference-Time Alignment of Specialized LLMs via Cross-Family Representation Steeringcs.CL · cs.AI
EvoSkill Injection: Red-Teaming Autonomous Skill Generation and Evolution in Self-Evolving Agentscs.AI · cs.CL
Auditing Anonymous AI Models: A Four-Stage Protocol for Black-Box Identity Verificationcs.SE · cs.AI
TACS: Trajectory-Aware Candidate Selection for LLM Jailbreak Suffix Optimizationcs.CL · cs.LG
Membership is Ownership: A Robust Ownership Verification Framework for Diffusion Modelscs.CR · cs.CV
TrainSDC: Characterizing and Mitigating Silent Data Corruption in Large Language Model Trainingcs.LG
ECLIPSE: Self-Evolving Stealthy Prompt Injection Attack against Long-Horizon Agentic Systemscs.CR
STEP: A Modular Silent Trial Engine for Operational Evaluation of Digital Pathology AI in Routine Workflowcs.SE · cs.CV
Where Identity Lives: Localized, Retain-Free Identity Unlearning in Multimodal Large Language Modelscs.CL · cs.CV
GreenBench: Benchmarking Energy Efficiency and Carbon Footprint of Open-Source LLM Inference on Apple Siliconcs.CL · cs.AI
Physical Adversarial Examples for Person Detectors in Thermal Images Based on 3D Modelingcs.CV
Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilitiescs.LG · cs.AI
The Safety Relay in Roleplay Jailbreaks: A Component-Resolved Causal Analysis of Harm Recognition and RefusalMetacs.LG
Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon NavigationAlibabacs.AI · cs.RO
IndicDetect: Evaluating Cross-Lingual LLM-Generated Text Detection for Hindi, Telugu, and Tamilcs.CL · cs.AI
SemTrace: Source-Grounded Semantic Signatures for Tracing LLM Exposure to Protected Documentscs.CL
Extracting Knowledge from Tools in LLM Agentscs.CR
Do VLMs Share Safety Neurons Across Modalities?cs.LG
AI Can Be Easily Persuaded in Clinical Decision Makingcs.CL · cs.AI
Can escalation channels redirect reward hacking toward defect disclosure?cs.AI · cs.CR
WebWorld: The Browser as a World Model for Self-Improving Web Codecs.CL · cs.SE
Reachability-Based Capability Confinement for LLM Agents under Indirect Prompt Injectioncs.CR · cs.AI
Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollbackcs.CR · cs.AI
RegDivergence-101: An LLM Benchmark for Cross-Jurisdiction Regulatory Contradiction Detection in Life Sciencescs.AI · cs.CL
Zero-Knowledge Predicate Proofs Between AI Agents: A Measured, Cross-Protocol Gateway and the Source-Integrity Gapcs.CR · cs.MA
The Fragility of Jailbreak Robustness Across Operational Statescs.CR · cs.CL
ALTSTEER: Selective Safety Steering for Moving Beyond Hard Refusals to Constructive Alternativescs.CL · cs.SE
One Capability or Many? Testing the Economic Validity of Frontier AI Evaluationcs.LG · cs.CY
Defending Wearable VLMs Against Private Attribute Inferencecs.CV · cs.AI
Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agentscs.AI · cs.CL
Adversarial Calibration Attack on Autonomous Vehiclescs.RO · cs.CV
Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibrationcs.HC · cs.CV
Guardrail-Agnostic Societal Bias Evaluation in Large Vision-Language ModelsNVIDIAcs.CV
Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Datacs.AI · cs.CL
Hidden Threat in Synthetic Data: Covert Targeted Bias Injection through Benign Textcs.CL · cs.AI
Manacá-1B: An Open, Reproducible Brazilian-Portuguese Language Model and a Tokenizer-Aware, Paired Evaluationcs.CL
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K ResolutionAlibabacs.CV · cs.SD
One note in three: a verified census of three deployed AI scribes, and the instrument that counted itcs.CL · cs.AI
SIR: Self-improving Red-teaming for Compute Use AgentsIBMcs.CR · cs.AI
Sycophantic Agreement Transfers with Neutral Data via Contrastive Preference Optimizationcs.LG
Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steeringcs.LG
Identity by Design, Demographics by Accident: Demographic Leakage and Suppression in Behavioral Biometric Embeddingscs.CR
The Halt Vector: Internalizing a Causal Steering Intervention for Efficient Reasoningcs.LG · cs.AI
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culturecs.CL · cs.AI
The Price of Intelligence: A Quality-Adjusted Price Index for AI Servicesecon.GN · cs.LG
WoE Wrote It? Watermarking Mixture-of-Experts LLMs for Black-Box Text Provenancecs.CR
JITterFlip: Uncovering Fault Attack Surfaces in JIT-Compiled LLM Servingcs.CR
Drishti: AI-Led Human-Directed Vulnerability Auditing for 5G Corescs.CR
MON 31 AUG 20268 in view
8 CLEARED
OpenStamp: A Watermark for Open-Source Language Modelscs.CL · cs.AI
FISGuard: Defending Against Membership Inference via Fixed Input Subspacescs.CR · cs.AI
Recognition Without Enforcement: Configuration-Dependent Failures in LLM Agent Instruction Arbitration and External Controlcs.CR
When Robots Mishear Us: Mapping the Safety Risks of Voice-Controlled Embodied AIcs.AI · cs.CL
EvoHarmBench: Breaking Content Moderation with Iterative Human-Like EvasionAlibabacs.CL
Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verificationcs.CR · cs.AI
CURA: Certified Runtime Alarms for Computer-Use Agentscs.AI · cs.CV
LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrailscs.AI
loading earlier editions…
THE FIELD — 13,720 · 30D
cs.AI4,173
cs.LG2,644
cs.CL2,149
cs.CV1,932
cs.RO925
cs.CR536
every indexed paper — the whole wire, not just what cleared
IN THE LITERATURE
models referenced in the loaded window · click to pivot
FROM INSIDE THE LABS
papers with a frontier-lab author · click to pivot
Research — every claim one click from the paperarXiv continuous index · FRONTIER ◆ leads each edition