Problem This paper addresses the limitations of existing machine learning engineering (MLE) agents, particularly their inability to effectively manage inter-branch information isolation, memoryless search processes, and lack of hierarchical control.…
Problem This work addresses the challenge of unstable weight conditioning during the training of large language models (LLMs), which can hinder convergence and performance. The authors identify a gap in…
Problem This work addresses the gap in understanding the distribution of effective interpolating classifiers in high-dimensional settings, particularly in the context of overparameterization. The authors investigate the performance of unit…
Problem This paper addresses the inefficiencies in existing formal theorem proving methodologies, particularly those relying on recursive lemma decomposition, which can lead to looping on dead-end strategies. The authors propose…
Problem The paper addresses the limitations of existing sparse attention mechanisms in large language models (LLMs), particularly in the context of long-context inference. Current methods often face a trade-off between…
Problem This study addresses the "conjunctive handicap" in causal learning, where adults struggle to identify conjunctive causal rules compared to disjunctive ones. Previous research primarily utilized passive observation paradigms, limiting…
Problem The paper addresses the limitations of existing benchmark construction methods for large language models (LLMs) and multimodal language models (MLLMs), which are often labor-intensive, difficult to reuse, and prone…
Problem As autonomous LLM agents increasingly manage sensitive operations without human oversight, there is a critical gap in the ability to communicate access restrictions effectively. Current access control mechanisms either…
Problem The paper addresses the challenge of detecting interactions between multiple AI-driven control functions in next-generation wireless networks, specifically within AI Radio Access Network (AI-RAN) and Open Radio Access Network…
Problem This work addresses the limitations of existing Multiple Instance Learning (MIL) algorithms, particularly in low-label regimes prevalent in real-world applications. Current methods either overfit due to their flexibility or…
Problem This paper addresses the gap in understanding the efficacy of Popperian code-generation skills in large language models (LLMs), particularly in distinguishing whether performance improvements stem from the content of…
Problem This paper addresses the engineering challenges associated with deploying and evaluating sparse attention algorithms for large language models (LLMs). As the demand for longer generation lengths increases, existing methods…
Problem The paper addresses the lack of systematic characterization of agent memory systems, which are crucial for large language model (LLM) agents engaged in long-horizon tasks that require sustained reasoning…
Problem This work addresses the limitations of existing latent reasoning methods in large language models (LLMs), which often compromise the advantages of chain-of-thought (CoT) reasoning. While CoT enhances reasoning by…
Problem This paper addresses the limitations of existing multi-domain audio encoders, particularly the USAD and SPEAR models, which have restricted coverage and evaluation metrics. The authors highlight the need for…
Problem This work addresses the gap in understanding the reliability of large language models (LLMs) in simulating user stances in online discussions. Specifically, it investigates whether LLM-generated stances accurately reflect…
Problem The paper addresses the limitations of conventional Bayesian network construction methods, which typically rely on optimization techniques that may not adequately capture the inherent structural ambiguity in causal relationships.…
Problem This work addresses the limitations of existing methods for translating unseen or low-resource languages, which often rely on continued training or encoding specific language grammars. These approaches tend to…
Problem The paper addresses the limitations of existing diffusion-based methods for generating safety-critical traffic scenarios, which are essential for evaluating autonomous driving systems. These methods suffer from high computational costs…
Problem This work addresses the lack of resources and methodologies for evaluating machine translation in endangered languages, specifically focusing on the Komi-Yazva language, which is extremely low-resource. The authors present…
Problem The paper addresses the gap in optimization techniques for deep learning models that experience test-time feedback (TTF), a phenomenon where the mismatch between training/validation loss and downstream performance metrics…
Problem The paper addresses the challenge of discovering effective skills for data analysis in an unsupervised manner, particularly in the absence of reliable supervision, which is often costly. The authors…
Problem Current medical imaging AI excels in isolated image interpretation but lacks alignment with radiological practices that depend on comparative analysis of prior studies and reference cases. This paper addresses…
Problem This work addresses a significant gap in the evaluation of multi-agent systems (MAS) built on large language models (LLMs), specifically the lack of focus on collaborative competence. While existing…
Problem Current evaluation practices in relational learning predominantly utilize flat leaderboards that average performance across diverse datasets, implicitly assuming a uniform underlying structure. This assumption introduces systematic bias, obscuring geometry-dependent…
Problem This preprint addresses the gap in understanding the multifaceted risks associated with autonomous driving technology, particularly the interplay between technical failures, ethical dilemmas, and regulatory frameworks. While autonomous vehicles…
Problem The paper addresses the gap in the literature regarding the evaluation and training of probabilistic forecasts in the context of right-censored survival data. Conventional scoring rules are not applicable…
Problem The paper addresses the Certified Allocation Problem, which lacks a robust mechanism for redistributing financial burdens from rare adverse events among participants without making any individual worse off. This…
Problem The paper addresses the gap in generating coherent and realistic indoor scenes for robot simulation and interior design, particularly in the context of complex layouts and limited 3D scene…
Problem — The paper addresses the lack of authentic human collaboration data with action-level mental model annotations, which is crucial for developing agents capable of effective collaboration. Current models are…