Problem Multimodal image fusion integrates information from various modalities to create a coherent fused image. Existing methods predominantly utilize 2D feature grids, which effectively capture local structures but struggle with…
Problem The paper addresses the gap in temporal grounding research for hour-long videos, a domain that has been largely overlooked in favor of short video contexts. The authors argue that…
Problem The paper addresses the limitations of Vision-Language-Action (VLA) models, which often exhibit brittle mappings from natural language instructions to robotic behaviors. This brittleness can lead to inconsistent task execution,…
Problem This paper addresses the gap in multimodal AI capabilities, specifically in video retrieval and grounded generation of textual content based on retrieved videos. The work is presented as a…
Problem This study addresses the gap in forensic image retrieval systems that effectively integrate multimodal data sources, particularly in real-world scenarios. Prior research has primarily focused on optimizing multimodal retrieval…
Problem The paper addresses a critical gap in the evaluation of large language models (LLMs) in medical applications, specifically their ability to maintain accurate medical judgment when faced with misleading…
Problem The paper addresses the lack of a unified theoretical framework for designing interpretable machine learning methods, which has led to a fragmented literature and inconsistent evaluation protocols. The authors…
Problem — The paper addresses the high energy consumption of Transformer architectures in natural language processing (NLP) by proposing a fully spiking neural network (SNN) implementation. While previous works have…
Problem — The paper addresses the challenge of accurately counting living cells in phase-contrast microscopy images, a critical task in biological research workflows, particularly in large-scale genome editing studies. The…
Problem The paper addresses the limitations of existing expressive performance rendering (EPR) models, particularly those that utilize flow matching audio editing techniques. Traditional models are constrained to manipulating synchronized music…
Problem The paper addresses the limitations of existing action-advising methods in Decentralized Training and Decentralized Execution (DTDE) for Multi-Agent Reinforcement Learning (MARL). Current approaches often lead to excessive reliance on…
Problem This work addresses the gap in empirical evaluations of hardware-specific trade-offs in post-training quantization for large text-to-image diffusion transformers, particularly for consumer-grade GPUs that lack FP8 tensor cores. The…
Problem — This work addresses the gap in the understanding of genetic algorithms (GAs) when applied in machine learning contexts, particularly at inference time. Traditional GAs utilize random mutation and…
Problem This work addresses the inefficiency of existing iterative pruning methods, which require multiple training cycles to identify sparse subnetworks that maintain performance. The authors highlight the limitations of the…
Problem The paper addresses the lack of methodologies for discovering multiple models that exhibit similar performance metrics (loss/accuracy) while possessing significantly different context-aware characteristics. This gap is particularly relevant in…
Problem This work addresses the limitations of existing post-training methods for diffusion large language models (dLLMs), which predominantly utilize random masking strategies that fail to leverage the intrinsic dependencies between…
Problem The paper addresses the challenge of eliciting latent knowledge (ELK) from AI systems, particularly focusing on the difficulty of ensuring that these systems report their beliefs honestly about hidden…
Problem — This preprint addresses the underexplored area of how memory augmentation in AI models can lead to performance degradation. While memory systems are often posited to enhance model capabilities…
Problem The paper addresses the vulnerability of Latent Diffusion Models (LDMs) to unauthorized mimicry, a growing concern in visual synthesis applications. Existing defenses rely on injecting deceptive perturbations to mislead…
Problem The paper addresses the inadequacies of current market structures for human-generated content used in AI training, highlighting a gap in the literature regarding effective market design that balances technological…
Problem The paper addresses the challenge of cross-domain day-night re-identification (ReID), which suffers from significant visual discrepancies between daytime and nighttime images. Existing fully supervised methods require extensive manual annotations,…
Problem The paper addresses the computational inefficiency in training deep neural networks for electrocardiogram (ECG) classification, particularly in resource-constrained healthcare environments. Existing methods like Progressive Data Dropout, which reduces training…
Problem Gradient-based adversarial attacks pose a significant threat to deep neural networks (DNNs) by leveraging gradient information to optimize adversarial perturbations. This paper addresses the gap in existing literature regarding…
Problem The paper addresses the gap in the evaluation of large language models (LLMs) in medical applications, specifically highlighting the limitations of traditional multiple-choice question answering (MCQA) methods. The authors…
Problem Current research on bias in large language models (LLMs) has primarily utilized third-person audits, which assess how models represent demographic groups as external subjects. This approach neglects the user's…
Problem Cold-start item recommendation is a significant challenge in recommendation systems, particularly due to the lack of interaction histories for new items. Existing models often rely on item content features…
Problem This paper addresses the inefficiencies in speculative decoding (SD) for large language models (LLMs), particularly the high inference costs associated with the binary accept/reject decisions of existing draft-verify methods.…
Problem The paper addresses the limitations of traditional recurrent neural networks (RNNs), particularly Long Short-Term Memory (LSTM) networks, in modeling complex multivariate time-series data characterized by irregular sampling and heterogeneous…
Problem The paper addresses the gap in understanding the trade-offs between effectiveness and fluency in conditioning methods for Large Language Models (LLMs). While existing literature often evaluates conditioning techniques based…
Problem Masked diffusion language models (dLLMs) have emerged as a competitive alternative to autoregressive models, particularly due to their potential for faster inference through parallel token generation. However, a critical…