Problem The paper addresses the misclassification of harmful memes by vision-language models, specifically highlighting issues related to missing internal evidence and routing problems. This work is particularly relevant as it…
Problem The paper addresses the challenge of ensuring safe and stable resource management in Open Radio Access Networks (O-RAN) when utilizing autonomous AI agents. These agents can exhibit unsafe independence,…
Problem The paper addresses the quality-capacity trade-off in compact acoustic models for speech synthesis. It highlights the limitations of existing models in effectively utilizing context for accurate predictions, particularly in…
Problem The paper addresses a significant gap in the capability of serving systems, specifically the inaccurate estimation of tool call durations. This inaccuracy leads to inefficient management of key-value (KV)…
Problem Conventional language models are limited in their ability to learn from live interaction data, which restricts their adaptability and responsiveness in dynamic environments. This paper addresses this gap by…
Problem This work addresses a gap in understanding how visual language models (VLMs) execute optical character recognition (OCR). The authors investigate the mechanisms within VLMs that facilitate the extraction and…
Problem Compositional Policy Violations (CPVs) represent a significant gap in the evaluation of AI workflows, particularly when individual steps may pass compliance checks while the overall execution fails to adhere…
Problem This work addresses a gap in the evaluation of coding agents, specifically focusing on their performance in inferring features from incomplete applications. The authors highlight the need for a…
Problem The paper addresses a significant gap in the adaptability of large language models (LLMs) due to the separation of query routing processes and agent fine-tuning. This limitation hinders the…
Problem This work addresses the challenge of answering questions based on normative documents while considering various contextual factors such as version, jurisdiction, subject, date, and source text traceability. The authors…
{'Problem': 'The paper addresses the issue of poor quality or noisy annotations in Named Entity Recognition (NER), which adversely impacts model performance. This is particularly critical for low-resource languages where…
Problem The paper addresses a gap in the capability of large language models (LLMs) to perform logically consistent deductive reasoning over extended interactions. This limitation is particularly evident in complex…
OpenAI Economic Research has published findings on the evolving use of AI tools, specifically ChatGPT, in workplace settings. The analysis, which spans from April to July 2026, examines over 1.5…
Problem Existing harnesses and messaging primitives are inadequate for facilitating effective communication and collaboration in agentic societies. This paper identifies a critical gap in capability, where honest and competent agents…
{'Problem': 'The paper addresses the gap in interactive scientific research tools that can improve through user interaction. It highlights the need for agents that not only assist in scientific tasks…
Problem Existing controllable video generation methods either require predefined control schedules or rely on pixel-space signals for object positioning, which do not effectively leverage physical dynamics. This paper addresses these…
Problem Large language models (LLMs) often generate fluent responses that lack robust factual support, leading to potential misinformation. This paper addresses the need for a mechanism that allows LLMs to…
Problem High frame rates in neural audio codecs result in long sequence lengths, leading to increased computational costs. This paper addresses the inefficiencies associated with existing dynamic frame rate methods,…
Problem The paper addresses a significant gap in uncertainty estimation for Vision-Language Navigation (VLN) models, particularly in the context of conformal prediction. The authors highlight that existing methods do not…
Problem The paper addresses a significant gap in the capability of existing models for structured data intelligence, particularly in their ability to generalize across diverse structured datasets. The authors highlight…
Problem This paper addresses a significant gap in explainability techniques for black-box object detectors, particularly in the context of marine mammal research. Existing methods lack the capability to provide interpretable…
Problem This work addresses the significant memory constraints faced when running open-weight models for coding and reasoning tasks on local devices, specifically laptops. The authors highlight the challenges of managing…
Problem The paper addresses the systematic bias and errors that arise in large language model (LLM) distillation due to direct imitation under covariate shift. This issue is particularly pronounced when…
Problem This work addresses a gap in understanding the impact of task decomposition in multi-agent systems on information discovery. The authors highlight that existing literature does not adequately explore how…
Problem The paper identifies a significant gap in the reliable performance of scientific AI agents specifically within the domain of quantum engineering. Despite advancements, there is a lack of robust…
{'Problem': 'The paper addresses a gap in the integration of caregiver wellbeing within the dementia care ecosystem. It highlights the lack of focus on caregiver needs and wellbeing in existing…
Problem The paper addresses the challenge of one-to-many mobile charging in environments with numerous candidate actions, specifically in large dynamic action spaces. The authors highlight the limitations of existing methods…
Problem The paper addresses the challenge of full and long-term occlusion in multi-object tracking systems, which is a critical gap in existing literature. The authors propose a novel framework to…
Problem The paper addresses the inadequacy of the SWE-bench leaderboard in effectively ranking coding agents, highlighting that the current metrics fail to differentiate between the top entries. The authors conduct…
Problem This work addresses a gap in automated tuning and optimization for model serving stacks, which is critical for enhancing the performance of machine learning applications. The authors highlight the…