Problem Non-deterministic responses of large language models (LLMs) complicate failure reproduction in LLM agents. This paper addresses the gap in regression testing methodologies for LLMs, particularly in ensuring reliable failure…
Problem Autonomous underwater vehicles (AUVs) are required to autonomously recover from failures without human intervention. This paper addresses the gap in existing literature regarding effective diagnostic strategies for AUVs, particularly…
Problem Inference engine fingerprinting attacks represent a significant gap in current literature, particularly in the context of model-driven environmental discovery and exploitation. This work addresses the lack of practical exploration…
Recent research by Han et al., Balek et al., and Malberg, Mosca & Groh explores the application of large language models (LLMs) as classifiers, particularly in the context of irony…
Problem This work addresses a gap in the understanding of dependencies between tokens in discrete diffusion processes. The authors highlight that existing literature does not adequately explore how these dependencies…
Problem This work addresses the high computational and GPU memory costs associated with learning visual policies for tasks such as locomotion and manipulation. The authors highlight the inefficiencies in existing…
A recent article from The Economist highlights a pivotal development in artificial intelligence, asserting that AI systems have now outperformed some of the best human forecasters. This claim underscores the…
Problem The paper addresses a gap in existing mitigation approaches for algorithmic collusion specifically in the context of general repeated games. The authors highlight the inadequacy of current methods to…
Problem This preprint addresses the gap in understanding how full-consensus rates as indicators of collective cognition are influenced by participation and the operationalization of final states. The authors highlight that…
Problem Closed-loop AI evaluation can inadvertently support incorrect claims, even when reproducibility is maintained. This paper addresses the need for a robust protocol that ensures claims made by AI systems…
Problem The paper addresses the challenge of poor transferability of predictive maintenance models across different machines and operating conditions, particularly in scenarios with limited labeled data and varying sampling frequencies.…
Problem This work addresses a gap in token efficiency for scaling recursive self-improvement in coding agents. The authors propose a novel approach to enhance the performance of auto-research loops, which…
Problem This work addresses the confounding of optimization objectives and search procedures in tokenization algorithms, a gap in the literature that has implications for the design and evaluation of tokenizers.…
Problem The paper addresses a gap in the capability of existing methods to extract information from preference pairs with small likelihood margins. This is particularly relevant in the context of…
{'Problem': 'Existing video generation methods primarily focus on kinematic trajectories, lacking the integration of force information, which is critical for successful execution in contact-rich manipulation tasks. This gap leads to…
Problem Language agents currently exhibit limitations in interactive environments, particularly in tasks that require long-horizon state tracking and the ability to recover from failures. This paper addresses these gaps by…
Problem The paper addresses a significant gap in the design of user interfaces, which are primarily tailored for human users and often lack clarity for machine readers. This misalignment can…
Problem Emergent coordinated behaviors of AI agents pose significant safety risks, necessitating a mechanistic understanding for effective collective alignment. This work addresses the gap in literature regarding the interpretability of…
Problem The paper addresses a gap in existing vision-language-action (VLA) inference frameworks, which do not effectively exploit embodied workloads and distinct bottlenecks. This limitation hinders the efficiency of VLA models…
Problem This paper addresses the gap in workforce readiness for the adoption of artificial intelligence (AI) in healthcare within low- and middle-income countries, specifically focusing on Nigeria. The study is…
Problem Variations in radiology reporting practices significantly influence the evaluation of AI-based radiology report generation models. This paper addresses the gap in understanding how these variations affect model performance, particularly…
Problem Ambiguity in passive syndrome records complicates the selection of recovery operations in quantum error correction. This paper addresses this gap by proposing a method to secure quantum error correction…
Problem The paper addresses a significant gap in the evaluation of large vision-language models specifically within educational settings, particularly concerning artistic content. Existing benchmarks do not adequately assess the capabilities…
Problem This work addresses a gap in the literature regarding probabilistic explainability specifically for continuous regression and binary classification tasks. The authors highlight the need for effective methods that can…
{'Problem': 'The paper addresses the double descent phenomenon in model test error as a function of the number of parameters, a gap in understanding that remains underexplored in the literature.…
{'Problem': 'The paper addresses a gap in accessible implementations of reinforcement learning (RL) specifically designed for educational purposes. The authors note that existing resources often lack practical, illustrative examples that…
Problem The paper addresses a gap in the capability of coordinating multiple agents in stochastic environments, which is critical for applications in AI where agents must operate under uncertainty. The…
Problem Large parameter counts in Mixture-of-Experts (MoE) language models create significant memory bottlenecks, hindering their deployment and efficiency. This paper addresses this issue by proposing a novel pruning technique that…
Problem The paper addresses a gap in the cost-effectiveness of agent benchmark evaluations compared to conventional large language model (LLM) benchmarks. The authors highlight the inefficiencies in current methodologies, particularly…
Problem The paper addresses a significant gap in the evaluation of privacy for tool-using large language model (LLM) agents, specifically concerning unauthorized exposure during multi-step sessions. Existing methodologies do not…