Problem The paper addresses a significant gap in the literature regarding the early detection of reward hacking in reinforcement learning (RL) systems. Traditional studies focus on observable reward hacking after…
Problem Generating coherent and controllable long-form content remains a significant challenge for Large Language Models (LLMs), particularly in open-ended writing scenarios. Existing reasoning-enhanced models exhibit a severe performance decline, termed…
Problem The paper addresses the lack of robust and reproducible tools for modifying neural network weights, particularly as model sizes increase. Existing workflows often depend on fragile, ad-hoc Python scripts,…
Problem This work addresses the gap in understanding the conditions necessary for local score models to achieve stable size extrapolation in generative modeling. While existing literature acknowledges the importance of…
Problem The paper addresses the challenge of adaptive red teaming in AI, specifically the need for continuous evolution of both attackers and defenders in language models. Previous works have shown…
Problem This work addresses a critical gap in the literature regarding the vulnerability of large language model (LLM)-powered content moderation systems to adversarial attacks that exploit human perceptual cues. The…
Problem Existing generative models like CycleGAN and Pix2Pix struggle with cross-modality semantic alignment in craniofacial reconstruction, particularly when translating from X-ray skull images to optical face images. This paper addresses…
Problem This work addresses the gap in the capability of large language models (LLMs) to handle requests that necessitate refusals, particularly in high-risk scenarios involving crisis or coercion. Traditional refusal…
Problem The paper addresses a significant gap in the observability of delegated execution within agentic AI systems, particularly those utilizing large language models (LLMs). Current audit logs and execution traces…
Problem The proliferation of numeric formats in machine learning hardware, including FP8 (E4M3 and E5M2), BF16, MXFP4, and various microscaling formats, has created a significant gap in vendor-neutral reference materials.…
Problem The paper addresses the challenge of efficiently compiling large language models into CUDA kernels without manual intervention, a gap in the literature regarding automated kernel synthesis. The authors highlight…
Problem — The paper addresses the lack of robust AI-enabled video-oculographic solutions for detecting saccadic signatures in neurological diseases, which are critical for screening and localizing brain abnormalities. Current methods…
Problem The paper addresses the challenge of player-centric ball-action spotting in soccer, specifically for the SoccerNet 2026 competition. The task involves identifying which player performs specific actions at given times…
Problem Current discriminative models for multi-channel speech separation, while effective in reference-based metrics, often fail to deliver optimal human listening quality. This paper addresses this gap by proposing a novel…
Problem — The paper addresses the challenge of autoformalization in mathematical proofs, particularly the lack of reliable systems that can convert natural language proofs into formalized versions. This gap is…
Problem The paper addresses a significant gap in causal discovery within biomedical language models, particularly the inability of existing models (e.g., BioBERT, PubMedBERT, BioM-ELECTRA) to accurately discriminate between unrelated cross-domain…
Problem The paper addresses the limitations of existing machine learning approaches in Alzheimer's disease (AD) prediction, which often rely on static classification or cohort-level risk estimation. These methods struggle with…
Problem Recent anomaly detection methods have demonstrated high accuracy on benchmark datasets like MVTec; however, they often fail under real-world conditions where foundational assumptions—such as consistent object scale, viewpoint, background,…
Problem Existing benchmarks for multimodal large language models (MLLMs) primarily focus on passive evaluations, such as static visual question answering (VQA), or are limited to specific simulation environments. This paper…
Problem The paper addresses the limitations of existing algorithms for contextual queueing bandits, which achieve a queue length regret of $\widetilde{\mathcal{O}}(T^{-1/4})$. This work is particularly relevant as it is a…
Problem This work addresses the gap in silent speech synthesis (SSI) technologies, particularly the integration of surface electromyography (sEMG) and video-based lipreading for continuous speech synthesis. Existing multimodal approaches often…
Problem The paper addresses the gap in efficient sampling methods for interval patterns under user-defined syntactic constraints, a topic that remains underexplored in the literature. Existing methods often rely on…
Problem — This work addresses the nascent concept of MetaAI, which lacks a comprehensive framework for evaluating AI systems that can modify their own design and operation. The authors argue…
Problem This work addresses the unclear effects of built-in reasoning mechanisms in large reasoning models (LRMs) on instruction following performance. While prior research has established that LRMs enhance capabilities in…
Problem The paper addresses the limitations of existing techniques for compressing key-value (KV) caches in long-context language model inference, which are hindered by memory constraints as the KV cache size…
Problem This paper addresses the disconnect between the advancements in machine translation (MT) technology and the concerns of non-AI stakeholders, such as professional translators, language learners, and service providers. Despite…
Problem This work addresses the gap in understanding whether pretrained video foundation models (VFMs) encapsulate intuitive-physics knowledge in their representations. The authors conduct a layerwise probing analysis to evaluate this…
Problem Current multimodal large language models (MLLMs) excel in visual reasoning tasks but lack mechanisms to verify whether their answers are grounded in the correct visual evidence. This issue is…
Problem The paper addresses the limitations of existing inference-time alignment methods for Large Language Models (LLMs), specifically Best-of-$N$ and rejection sampling. These methods are constrained by the quality of the…
Problem The paper addresses the lack of scalable simulation tools for civil court cases, which are more prevalent than criminal cases but have received less attention in existing research. Current…