Problem The paper addresses the critical gap in the evaluation of T cell receptor (TCR) antigen specificity prediction models, which are essential for advancing T cell biology and immune engineering.…
Problem The paper addresses the high cognitive demands placed on scrub nurses during surgical procedures, particularly when handling unfamiliar instruments. This gap in capability is critical, as existing methods often…
Problem The paper addresses the lack of mechanisms for verifying, debugging, and auditing the behavior of large language model (LLM)-based agents, particularly as they interact with external tools and environments.…
Problem Existing datasets for multi-party dialogue often lack a focus on structured, complex reasoning tasks, particularly in collaborative settings. This paper addresses this gap by presenting DeliChess, a novel dataset…
Problem This work addresses the limitations of existing Vision-Language Models (VLMs) in food analysis, which predominantly rely on supervised fine-tuning (SFT). Such methods often hinder reasoning and generalization capabilities due…
Problem The paper addresses the challenge of deploying Mixture-of-Experts (MoE) architectures, which are constrained by memory requirements due to the necessity of storing all expert weights. Existing mixed-precision quantization methods…
Problem This work addresses the gap in understanding the alignment of large language models (LLMs) with human decision-making mechanisms in risk scenarios, specifically in the context of the St. Petersburg…
Problem — The paper addresses the limitations of existing robotic grasping systems and autonomous vehicle safety mechanisms, particularly their inability to generalize across novel objects and complex driving scenarios. It…
Problem The paper addresses the inefficiency of inference in diffusion large language models (DLLMs), which require numerous denoising steps for high-quality generation. Despite their parallel processing capabilities, the computational cost…
Problem — This work addresses the gap in the literature regarding the evaluation of machine learning engineering (MLE) agents in sensitive domains, particularly concerning fairness and compliance with regulatory standards.…
Problem The paper addresses the lack of large-scale, cross-domain benchmarks for proactive procedural assistance systems, particularly in scenarios where users deviate from expected task sequences. Existing literature does not adequately…
Problem The paper addresses a gap in the literature regarding the operational frameworks that facilitate AI software development agents. While previous surveys have cataloged AI tools for programming, there has…
Problem Current blockwise decoding methods for diffusion language models (DLMs) often utilize fixed block sizes or delimiter-based signals, which can misalign with semantic boundaries in text generation. This paper addresses…
Problem The paper addresses the limitations of traditional system-generated logs, which often follow rigid template formats that impede both automated analysis and human interpretability. The authors highlight the need for…
Problem The paper addresses the gap in effective management of gestational diabetes, particularly in resource-constrained settings, by proposing a digital solution that integrates patient-generated data with clinical workflows. Existing literature…
Problem This work addresses a gap in the understanding of active inference, particularly the relationship between Expected Free Energy (EFE) and Variational Free Energy (VFE) in decision-making processes. Prior research…
Problem The paper addresses the challenge of modeling nonlinear dynamics in nonstationary data streams, a gap in the literature that remains largely unfilled, particularly in real-time applications. Existing methods often…
Problem The paper addresses the critical issue of training data attribution (TDA) in large language models (LLMs), which remains an open problem in the literature. As LLMs are increasingly utilized…
Problem The paper addresses the gap in the literature regarding unsupervised video panoptic segmentation (VPS), a task that integrates object detection, segmentation, and tracking in video sequences without human supervision.…
Problem The paper addresses the emerging challenge posed by Large Language Models (LLMs) to the integrity of crowdsourced data, particularly in natural language processing (NLP) tasks. As LLMs become prevalent,…
Problem This work addresses the gap in understanding and mitigating reward hacking in rubric-based reinforcement learning (RL), particularly when using large language models (LLMs) as judges (LaaJ). Reward hacking occurs…
Problem Current prompt-based and adapter-based tuning methods for vision-language models (VLMs) in medical imaging are limited by their treatment of classes as equally incorrect, which neglects clinically relevant class relationships.…
Problem The paper identifies a significant gap in the AI-generated text detection literature, particularly the lack of a standardized definition of what constitutes harmful AI-generated text. Existing datasets and methodologies…
Problem This work addresses the gap in the literature regarding the need for linear auditability in large language model (LLM) agents, particularly in nontrivial problem domains. The authors highlight the…
Problem The paper addresses the limitations of existing preference optimization methods, particularly in the context of non-chatbot applications. While Direct Preference Optimization (DPO) has shown promise in dialogue systems, its…
Problem The paper addresses the challenge of optimizing urban layouts for climate adaptation, specifically balancing building density with cold-air ventilation. Traditional methods rely on computationally expensive physics-based climate simulations, limiting…
ParetoPilot introduces a zero-surrogate diffusion framework for offline multi-objective optimization, enhancing efficiency and performance without external models.
This paper presents SimuScene, a novel pipeline for simulation-ready 3D scene reconstruction from a single image, integrating physics during the generative process.