Problem This work addresses the unresolved question of whether financial news can reliably predict short-term stock movements using zero-shot natural language processing (NLP) techniques. Despite advancements in large language models,…
Problem The paper addresses the inefficiencies in large language models (LLMs) when integrating reusable natural language skills, particularly the high prefill costs and latency associated with repeatedly invoking full procedural…
Problem This work addresses the gap in performance when spoken dialogue models, which typically leverage text-based large language model (LLM) backbones, are conditioned on speech inputs. The authors identify a…
Problem This paper addresses the lack of systematic categorization and analysis of agentic environments for large language models (LLMs), which are critical for the development of interactive systems. Existing literature…
Problem — The paper addresses the lack of standardized resources for enthymeme detection in politically controversial discourse, which is characterized by subjective annotation practices. Existing datasets often eliminate disagreement among…
Problem This paper addresses the gap in the capability of large vision-language models (LVLMs) to perform high-stakes clinical reasoning that is grounded in visual evidence and clinical knowledge. Existing models…
Problem This work addresses the gap in effective hallucination detection methods for instruction-tuned large language models (LLMs). Existing techniques often struggle with generalizability and accuracy, particularly in zero-shot scenarios. The…
Problem Sparse autoencoders (SAEs) are commonly employed for interpreting neural network representations, yet their effectiveness hinges on the reproducibility of learned features across different training runs. This paper addresses the…
Problem The paper addresses the inadequacies in current benchmark evaluations of large language models (LLMs), which often misrepresent a model's knowledge due to reliance on specific formatting requirements. This issue…
Problem The paper addresses the lack of research on detecting sensitive personal information in Japanese pre-training corpora for large language models (LLMs). While there has been substantial work in English…
Problem Transformer-based language models applied to SMILES (Simplified Molecular Input Line Entry System) strings encounter a locality gap due to standard character-level tokenization. This fragmentation disrupts the representation of chemically…
Problem This work addresses the gap in fairness research in natural language processing (NLP) that typically relies on direct access to protected attributes (e.g., gender, race) for debiasing. The authors…
Problem The paper addresses the challenge of training deep search agents with verifiable questions that require substantial evidence for resolution. Existing methods for synthesizing search tasks often increase difficulty through…
Problem The paper addresses the limitations of traditional softmax attention mechanisms in energy-constrained environments, particularly on von Neumann hardware. Softmax requires exponentiation and global reduction, which are computationally expensive and…
Problem The paper addresses the limitations of existing runtime enforcement frameworks that primarily focus on untimed or discrete-time specifications, which are inadequate for reactive systems with complex continuous dynamics. The…
Problem The paper addresses the gap in existing frameworks for social intelligence reasoning, particularly the challenge of effectively utilizing multi-modal data while mitigating the overshadowing of long-tail events by more…
Problem This work addresses a critical gap in the understanding of model behavior during reinforcement learning (RL) training, particularly in the context of training-aware models. The authors highlight the phenomenon…
Problem Current diffusion-based virtual try-on methods primarily focus on 2D inpainting, emphasizing texture preservation at the expense of physical plausibility. This results in generated images that, while visually appealing, often…
Problem This work addresses the gap in applying tabular foundation models to clinical survival analysis, specifically for predicting right-censored time-to-event outcomes. While traditional survival analysis methods have been extensively studied,…
Problem Existing self-consistency methods for large language models (LLMs) primarily focus on exact matching, limiting their applicability to tasks with categorical outputs. This work addresses the gap in self-consistency approaches…
Problem The paper addresses the growing capabilities gap between trusted and untrusted AI models, which may compromise the reliability of traditional monitoring systems. As AI agents become more advanced, existing…
Problem Remaining Useful Life (RUL) prediction is critical for industrial predictive maintenance, yet existing learning-based methods often require extensive feature engineering or large labeled datasets, which can be impractical. This…
Problem The paper addresses the gap in simulating credible rainfall conditions for autonomous-driving perception tests, which is crucial for identifying system boundaries and supporting Safety of the Intended Functionality (SOTIF)…
Problem The paper addresses a significant gap in the literature regarding uncertainty modeling in dynamical systems, an area that has received less attention compared to supervised learning and generative modeling.…
Problem The paper addresses a significant gap in preference-based reinforcement learning (PbRL) methodologies, particularly the mismatch between utility function training and policy optimization. Existing PbRL approaches typically rely on per-step…
Problem This work addresses the challenge of accurately recovering structured Markdown documents from document page images, a task that requires both precise content recovery and faithful structure reconstruction. The authors…
Problem This work addresses the limitations of existing LLM-based agents that utilize linear exploration strategies for file localization in software repositories. The authors argue that such linear approaches are inadequate…
Problem This work addresses the computational inefficiencies in existing algorithms for multinomial logistic bandits (MLogB), specifically the OFUL-MLogB algorithm, which suffers from high time and space complexity due to parameter…
Problem Existing neural operator architectures struggle to effectively model nonlinear time-dependent systems characterized by multi-scale structures, long-range interactions, and stable long-term evolution. This paper addresses these limitations by introducing a…
Problem This work addresses the limitations of in-context learning (ICL) in large language models (LLMs) when applied to structured data, particularly under conditions of distribution mismatch. The authors highlight a…