Problem Existing Constrained Reinforcement Learning (CRL) methods for Safe Autonomous Driving lack the ability to impose dynamic constraints that adapt to real-time vehicle interactions. This paper addresses this gap by…
Problem Data collection for object detection in critical domains is costly and constrained. This paper addresses the gap in literature regarding the balance of datasets derived from real, simulated, and…
Problem Local language models often lag behind frontier models in capability, particularly when handling complex queries. This limitation can undermine the privacy and cost benefits associated with local deployment. The…
Problem The paper identifies a significant gap in the capability of large language models (LLMs) regarding probabilistic reasoning, particularly under decision costs. The authors highlight that existing models struggle to…
Problem Existing time series anomaly detection (TSAD) methods are constrained to a single temporal granularity, which limits their ability to capture multi-scale interactions in time series data. This paper addresses…
Problem This work addresses the gap in predicting confidence from agentic RAG (Retrieval-Augmented Generation) pipeline signals, specifically within the context of the NTCIR-19 R2C2 task. The authors highlight the need…
Multiverse Computing has developed ProvenanceGuard, a source-aware factuality verification system designed for multi-contextual prompt (MCP) based large language model (LLM) agents. This system aims to address the issue of cross-source…
Problem The paper addresses the significant gap in the availability of animal-fur datasets necessary for realistic and editable animal fur reconstruction from multi-view images. This limitation has hindered advancements in…
Problem This work addresses the challenge of developing a language model capable of operating efficiently across multiple compute budgets without necessitating distinct training or compression processes. The authors highlight the…
Problem Unified multimodal models require the capability to jointly learn reflection text and image generation to facilitate effective self-repair. This paper addresses this gap by proposing a novel approach that…
Problem Token consumption during the execution of large language model (LLM) agents exhibits substantial variability across different runs, complicating the ability to predict resource usage effectively. This paper addresses this…
Problem The paper addresses a gap in the effective utilization of experts in mixture of experts (MoE) models, particularly focusing on the limitations of existing architectures in optimizing expert parameters…
Problem The paper addresses the challenge of fitting longer context traces into GPU memory for agentic large language models (LLMs). This issue is critical as it limits the performance and…
Problem The paper identifies a gap in the effective initialization of linear Vision Transformers (ViTs) using pre-trained weights from Softmax ViTs. This issue is particularly relevant as the performance of…
Problem The paper addresses a significant gap in the evaluation of financial research agents, specifically the need for expert-guided rubrics that can systematically assess their performance. Existing methods lack the…
Problem This work addresses the gap in understanding how explanation-only training influences agent behavior, specifically in the context of software engineering tasks. The authors explore the potential of fine-tuning models…
Problem This work addresses a significant gap in the capability of tool-using agents to report failures with evidence justification. The authors highlight the lack of systematic evaluation in this area,…
Problem The paper addresses the severe exploration problem encountered in training generalist policies with task-agnostic rewards in reinforcement learning (RL). This issue is particularly pronounced when attempting to generalize across…
Problem The paper addresses the low-entropy bias prevalent in large language models, which constrains their effectiveness in facilitating open-ended scientific ideation. This limitation hampers the generation of diverse and innovative…
Problem This work addresses the challenges of calibration, out-of-distribution (OOD) detection, and robustness to distribution shifts in graph neural networks (GNNs). The authors highlight the need for a unified approach…
Problem The paper addresses a critical gap in the literature regarding defenses against distillation attacks, particularly those that do not consider the implications of further training with reinforcement learning. Existing…
Problem This work addresses the gap in reasoning performance within continuous diffusion models, specifically targeting the limitations in their ability to handle complex reasoning tasks. The authors propose a novel…
Problem This work addresses a gap in the literature regarding the impact of progressive disclosure on operational costs and skill-retrieval quality in large language model (LLM)-based agents. The authors conduct…
Problem This work addresses a gap in the literature regarding circuit-based explanations, specifically their inadequacy in accounting for model errors. The authors highlight that existing methods fail to provide insights…
Problem The paper identifies a critical gap in reinforcement learning (RL) systems that utilize verifiable rewards, specifically the issue of imperfect verifiers that can inadvertently reward incorrect responses. This phenomenon,…
Problem This work addresses the inefficiencies in the navigation processes of mobile GUI agents, which often struggle with the complexity and variability of app interfaces. The authors highlight the need…
Problem The paper addresses a critical limitation in GLA Transformers, where fixed-capacity memory matrices operate at a single temporal resolution, leading to representational bottlenecks. This issue restricts the model's ability…
Problem The paper addresses a gap in existing population-based optimization methods regarding the retention of contextual information related to search behavior success or failure. This limitation hinders the ability of…
Problem This work addresses a gap in understanding whether different forms of reasoning in language models, particularly in the context of multi-hop reasoning tasks, rely on the same underlying mechanisms.…
{'Problem': 'The paper addresses a significant gap in the reliability of reward models for instruction-following tasks in image generation. Existing models often struggle to provide consistent and verifiable rewards, particularly…