Problem — This work addresses the gap in the evaluation of dataset distillation (DD) methods in data-centric machine learning, particularly the inconsistency in evaluation protocols and the assumption that DD…
Problem — Current world models struggle with the trade-off between computational depth for accurate long-horizon predictions and the associated costs and error propagation. This paper addresses this gap by proposing…
Problem — This work addresses the limitations of existing looped architectures in deep learning, particularly the signal propagation issues that arise from depth, which can hinder effective learning in tasks…
Problem This work addresses the lack of standardized computational lexicons for Arabic, specifically focusing on the Al-Mawrid Arabic-English dictionary, a significant legacy print resource. The authors highlight the challenges posed…
Problem The paper addresses the challenge of evaluating LLM-empowered personal health agents, which leverage user health metrics to improve healthcare access. Current evaluation methods are limited by the high cost…
Problem — This work addresses a critical gap in the security of LLM-based systems, specifically the vulnerability of agent skill scanners to multimodal hidden instruction attacks. Current defenses primarily focus…
Problem The paper addresses the gap in applying on-policy self-distillation (OPSD) techniques to diffusion large language models (dLLMs), a domain that has not been previously explored. Existing OPSD methods are…
Problem This paper addresses the gap in understanding the adversarial robustness of large language models (LLMs), specifically Anthropic's Fable 5 and Opus 4.8, against automated jailbreak attacks. The study is…
Problem The paper addresses the scarcity of high-quality, long-context training data for large language models (LLMs), particularly in the financial domain. Existing datasets are often proprietary, costly, or limited to…
Problem The paper identifies a significant gap in the capabilities of deep research (DR) systems, which have primarily focused on generating reports and summaries rather than facilitating concrete workflows. Existing…
Problem The paper addresses the lack of publicly available multi-source cybersecurity datasets that include detailed labeling of events according to the MITRE ATT&CK framework. Existing datasets either focus on single…
Problem This work addresses the limitations of finite-dimensional (FD) diffusion policies, which suffer from temporal drift due to discretization artifacts, particularly affecting long-horizon performance in real-world applications. The authors highlight…
Problem — The paper addresses a gap in the understanding of temporal-difference (TD) learning with linear function approximation, specifically the limitations of the classical ordinary differential equation (ODE) framework that…
Problem The paper addresses the lack of a comprehensive quantitative understanding of Illegal, Unreported, and Unregulated (IUU) fishing and related activities, which are critical threats to marine ecosystems and fisheries…
Problem The paper addresses the gap in available datasets for training interactive world models, which require temporally aligned video-action-language trajectories that reflect human gameplay dynamics. Existing datasets either lack executable…
Problem This work addresses the limitations of standard physics-informed neural networks (PINNs) in solving nonlinear partial differential equations (PDEs), particularly the nonconvex nature of gradient-based training that can hinder convergence…
Problem This paper addresses the gap in understanding the verification capabilities of test files generated by AI coding agents in open-source pull requests (PRs). Despite the proliferation of agent-authored PRs—over…
Problem The paper addresses the inadequacy of existing evaluations of open-source Large Language Models (LLMs) for classifying Cyber Threat Intelligence (CTI) using the MITRE ATT&CK framework. Prior evaluations have relied…
Problem The paper addresses a significant gap in the evaluation of AI systems in the legal domain, specifically concerning doctrinal legal reasoning, which is essential for interpreting law. Current benchmarks…
Problem — The paper addresses the limitations of existing 3D editing methods in the context of face re-aging, particularly the inability to maintain consistency across multiple 2D views. Current techniques…
Problem The paper addresses the challenge of creating personalized cardiac electrophysiology (EP) digital twins, which traditionally rely on expert-driven hybrid physics-neural architectures. This approach is limited by the need for…
Problem The paper addresses the limitations of classical structure-from-motion techniques in generating 3D tree maps for the Open Forest Observatory (OFO). These traditional methods are prone to artifacts, lack detail,…
Problem The paper addresses the gap in effective medical question answering using wearable health data, which is characterized by continuous, high-dimensional, and longitudinal data streams. Current language models (LMs) struggle…
Problem — This work addresses the lack of a systematic approach to pricing flash memory endurance in embodied agents, treating it as a depreciating asset. Existing memory systems do not…
Problem This work addresses a critical gap in the evaluation of AI agents' ethical decision-making, specifically regarding animal welfare in agentic contexts. Existing benchmarks primarily assess text-based responses to prompts,…
Problem The paper addresses the lack of high-quality, field-collected datasets for analyzing firearm muzzle blast sounds, particularly for caliber classification. Existing research often relies on audio samples sourced from the…
Problem The paper addresses the limitations of existing end-to-end meta-reinforcement learning (MRL) methods, which often couple task inference with embodiment-specific control. This coupling can obscure non-parametric task semantics, reduce sample…
Problem — This work addresses a critical gap in the evaluation of large language models (LLMs) used in mental health support, specifically the inadequacy of existing benchmarks that focus on…
Problem This work addresses the gap in understanding the unintended regional biases introduced by user metadata in large language models (LLMs). Despite the increasing reliance on user location data for…
Problem Current methodologies for predicting immune biomarkers associated with the tumor immune microenvironment (TIME) are predominantly limited to single image modalities, which restricts their resolution and the effective use of…