Problem The paper addresses the inadequacy of current large language models (LLMs) in structured tutoring scenarios, where interactions are typically unstructured and lack a curriculum. Existing models struggle to effectively…
Problem This work addresses the gap in leveraging abundant 2D medical images to improve 3D medical visual question answering (VQA). The authors highlight the lack of effective methods that utilize…
Problem The paper addresses the challenge of interpreting language model representations, a critical aspect for understanding and controlling model behavior. Existing methods, particularly Sparse Autoencoders (SAEs), require extensive training and…
Problem — This study addresses the gap in the literature regarding the reproducibility and applicability of RXNGraphormer, a model designed for reaction prediction tasks. Despite its initial promise, there is…
Problem — The paper addresses the lack of comprehensive evaluation for automatic speech recognition (ASR) systems on code-switched speech, particularly in bilingual contexts. Existing literature primarily focuses on monolingual ASR…
Problem This work addresses the limitations of traditional GPU architectures for machine learning, particularly in applications requiring ultra-low latency (sub-microsecond) and high hardware efficiency. The authors highlight the inadequacy of…
Problem The paper addresses a significant gap in the literature regarding the systematic understanding of cross-modal alignment (CA) and cross-modal prediction (CP) in multimodal representation learning. Despite their prevalence, there…
Problem The paper addresses the limitations of traditional supervised fine-tuning (SFT) methods, which typically maximize the likelihood of observed tokens in a trajectory. This approach can be suboptimal due to…
Problem This work addresses the gap in multimodal models that effectively integrate image understanding, generation, and editing within a unified framework. Prior models often treat these tasks separately, lacking a…
Problem Existing autoregressive video generation methods for World Action Models (WAMs) face challenges in training convergence and accuracy, particularly at high frame rates. Current approaches are limited by their reliance…
Problem Low-light video enhancement (LLVE) is hindered by significant information loss in low-illumination environments. While recent multimodal approaches have improved performance by leveraging auxiliary modalities (e.g., event streams, infrared images),…
Problem This work addresses the limitations of existing test-time prompt learning methods, which are primarily designed for single-dataset scenarios. The authors highlight that real-world applications necessitate the ability to process…
Problem Current diffusion-based lip synchronization models, while achieving high visual quality and audio-visual alignment, are hindered by their reliance on full-sequence bidirectional attention and extensive denoising steps, making them impractical…
Problem The paper addresses the gap in automated journalism by proposing a system that can generate complete news articles from raw data, a task traditionally requiring extensive human effort and…
Problem This work addresses the underexplored area of context design in self-distillation for language models, particularly focusing on how feedback influences the learning process. The authors highlight that while conditioning…
Problem Deployed large reasoning models (LRMs) often exhibit unpredictable behaviors, complicating their application in critical tasks. Existing steering techniques typically manipulate hidden representations based on features derived from already generated…
Problem This paper addresses the gap in understanding the relationship between Gaussian-process upper confidence bound (GP-UCB) methods and decision-estimation-coefficient (DEC) methods within the context of frequentist Reproducing Kernel Hilbert Space…
Problem This paper addresses the limitations of existing distributed training systems that require manual design of parallelism strategies, which can hinder adaptability to new methods. Current frameworks often impose a…
Problem Current full-duplex spoken dialogue models primarily rely on supervised learning via token-level likelihood maximization, which inadequately captures interaction-level dynamics. This limitation leads to interactivity issues such as excessive silence…
Problem This preprint addresses the overestimation of Large Language Models (LLMs) as equivalent to human experts in knowledge economy tasks. The authors highlight that existing benchmarks often rely on performance…
Problem The paper addresses the inefficiencies in decoding-time key-value (KV) cache management in large language models (LLMs) during long chain-of-thought (CoT) reasoning. Existing methods primarily focus on uniform budget distribution…
Problem This paper addresses the limitations of existing forecasting models in handling long-term predictions on irregular geospatial meshes, particularly in the context of physical simulations. Current methods often rely on…
Problem The paper addresses the gap in the literature regarding the integration of stochastic differential equations (SDEs) in generative modeling, particularly the challenge of defining a precise distillation procedure for…
Problem Flow Matching models have shown strong generative capabilities but are hindered by the computational demands of ODE-based iterative sampling during inference. This paper addresses the gap in existing distillation…
Problem The paper addresses a significant gap in the evaluation of multimodal large language models (MLLMs) concerning their ability to generate parametric 3D models from textual and visual inputs. Existing…
Problem The rapid advancement of large language models (LLMs) in biological research raises concerns about biosecurity, particularly as these models can perform tasks traditionally requiring human expertise. This paper addresses…
Problem This work addresses the challenge of learning drifting concepts in the presence of Massart noise, a scenario where the labels of samples are noisy versions of a target concept…
Problem Current virtual try-on methods primarily focus on direct clothing replacement, which limits the diversity and flexibility of try-on outputs. This paper addresses the gap in the literature regarding user…
Problem The authors identify a critical gap in the computational modeling of longitudinal resistance in EGFR-mutant non-small-cell lung cancer (NSCLC) patients undergoing treatment with osimertinib. Despite the well-documented clonal evolution…
Problem This work addresses the gap in data assimilation (DA) methodologies for subsurface flow modeling, particularly the challenge of calibrating model parameters to observed data while maintaining geological realism. The…