Problem Current post-training methods like supervised fine-tuning (SFT) and preference optimization typically enforce a singular global assistant behavior in language models. This approach, while effective in enhancing average helpfulness, often…
Problem This work addresses the gap in evaluating large language models (LLMs) on cross-lingual semantic understanding, specifically between Arabic and Hebrew, two closely related Semitic languages. The authors highlight the…
Problem This work addresses the gap in unsupervised methods for detecting hallucinations in neural machine translation (NMT) and abstractive summarization, specifically focusing on the limitations of existing detection techniques. The…
Problem This work addresses the gap in understanding the internal mechanisms of reward models used in reinforcement learning from human feedback (RLHF), particularly the conflicting objectives of helpfulness and harmlessness.…
Problem This work addresses the inadequacy of existing metrics in stance detection for large language models (LLMs), particularly in handling complex examples. Despite the increasing use of prompt-based LLMs for…
Problem The paper addresses the lack of large-scale, domain-specific datasets for stance detection in bioethical debates on social media, particularly on platforms like Reddit. Existing resources do not adequately capture…
Problem Existing legal NLP datasets predominantly focus on single jurisdictions, limiting their applicability for multinational companies that require cross-jurisdictional contract review. This paper addresses this gap by presenting LAUKIN (Legal…
Problem This paper addresses the gap in the literature regarding the resurgence of analog computing, particularly in the context of solving differential and matrix equations. Despite the growing interest in…
Problem The paper addresses the challenge of managing memory in large language model (LLM) agents during long-term interactions, where the accumulation of dialogue data leads to an unbounded memory store…
Problem — The paper addresses the persistent issue of interactive LLM agents failing to retain user preferences across sessions, leading to repeated violations of user corrections. Despite existing memory mechanisms,…
Problem The paper addresses the critical issue of hallucinations in large language model (LLM)-based timeline summarization (TLS), which is underexplored in existing literature. Hallucinations manifest as unfaithful content and omissions…
Problem Existing persona-grounded dialogue systems inadequately model the complex relationships among persona attributes, treating them as a flat set of sentences. This limitation hinders the generation of contextually relevant responses…
Problem This paper addresses the limitations of existing prefix caching mechanisms in retrieval-augmented workloads, particularly in the context of vLLM (variable-length Language Model) engines. Traditional prefix caching requires identical prefixes…
Problem The paper addresses the challenge of excessive pauses in simultaneous speech-to-speech translation, which can lead to unnatural speech flow and increased cognitive load for listeners. While existing methods prioritize…
Problem Existing benchmarks for evaluating search agents, such as BrowseComp, rely on static datasets, which can lead to test-set contamination and parametric memorization. This results in inflated performance metrics, as…
Problem The paper addresses a significant gap in spatio-temporal forecasting, particularly the inability of existing spatio-temporal graph neural networks (STGNNs) to effectively handle the phenomenon of temporal mirage, where similar…
Problem This work addresses the gap in automated design methodologies for approximate multipliers, which are critical for enhancing power efficiency and reducing latency in error-resilient applications like neural networks. The…
Problem The phenomenon of grokking in transformers, where models transition from near-chance to near-perfect performance on modular arithmetic tasks, lacks a comprehensive understanding of its timing, causal structure, and controllability.…
Problem The paper addresses the challenge of mixed categorical-continuous optimization in black-box settings, which is prevalent in various practical applications. Existing methods, particularly evolution strategy-based approaches like CMA-ES, struggle with…
Problem The paper addresses the gap in the integration of natural language processing (NLP) techniques with three-dimensional (3D) molecular modeling, specifically in the context of conformer prediction and molecular generation.…
Problem The paper identifies a critical gap in the application of synthetic datasets for biomedical machine learning, specifically the simulation-to-reality gap that undermines the reliability of synthetic data in predicting…
Problem Current visual-token reduction methods in vision-language models (VLMs) predominantly utilize a rank-and-remove strategy, which permanently discards visual tokens deemed less important. This approach is problematic as the relevance of…
Problem The paper addresses the limitations of existing dialogue systems that struggle with the growing computational costs associated with maintaining extensive dialogue histories. Traditional methods, such as naive truncation or…
Problem This work addresses the gap in understanding how design choices impact the performance of general-purpose large language models (LLMs) when applied to specialized tasks in pathology, particularly with whole-slide…
Problem The paper addresses the gap in force sensitivity for commodity robot arms, which typically lack dedicated force sensors due to cost constraints. This limitation hampers their ability to perform…
Problem The paper addresses the inefficiencies in deploying Vision-Language Models (VLMs) as high-level planners for embodied agents, particularly the challenge of scaling test-time compute. While increasing compute can enhance performance,…
Problem The paper addresses the limitations of existing context distillation methods in Large Language Models (LLMs), particularly the inefficiencies associated with long input sequences. While prior work like Doc-to-LoRA has…
Problem This work addresses the lack of design principles for router matrices in Mixture-of-Experts (MoE) models, which serve as proxies to determine expert activation based on input similarity. The authors…
Problem Existing vision-language-action (VLA) models struggle to effectively ground actions in complex 3D environments, often relying on frozen 3D features or sparse geometric constraints that lack dense spatial signals. This…
Problem The paper addresses the gap in domain-specific research for classical Chinese poetry translation and affective-semantic understanding, which has been largely overlooked in favor of general-domain approaches. The authors highlight…