This paper presents DiffUNet^2, a conditional diffusion model for bidirectional prediction and interactive visual analytics in scientific data exploration.
This paper introduces High-Confidence Causally Aligned Training (HICAT) to enhance adversarial training by addressing overfitting to spurious correlations.
This paper explores knowledge editing in masked diffusion models, revealing performance discrepancies compared to autoregressive models and proposing a corrective method.
This paper introduces a contrastive learning framework for approximate k-coloring in graphs, enhancing generalization across varying graph sizes and distributions.
This paper introduces the Visual STAte Tracking benchmark (VSTAT) to evaluate visual state tracking capabilities in Multimodal Large Language Models (MLLMs).
This paper presents a model for predicting the diffusion of scientific concepts, focusing on quantum computing, using a co-occurrence network and LightGBM.
This paper introduces Hedge-Bench, a benchmark for evaluating AI agents on complex financial reasoning tasks, addressing gaps in existing evaluation methods.
This paper introduces PatchScene, a diffusion-based framework for large-scale LiDAR scene completion, enhancing geometric accuracy and temporal consistency.
This paper introduces Bootstrap Your Generator (ByG), a novel framework for unpaired visual editing using flow matching, enhancing scalability in generative models.
This paper introduces NetKV, a network-aware scheduling method that optimizes Time to First Token in disaggregated LLM inference by considering network costs.
SparseStreet introduces a novel compression framework for 3D Gaussian Splatting, optimizing street scene simulation by reducing Gaussian primitives while maintaining fidelity.
This paper introduces scTranslation, a benchmark for evaluating single-cell multi-omics modality translation, addressing gaps in systematic evaluation.
CoralBay introduces a self-supervised learning framework for 3D CT imaging, enhancing representation learning for medical tasks through a novel architecture.
This paper introduces Reveal-IG, a novel path-based attribution method that enhances feature importance explanations by utilizing structured probe distributions.
This paper introduces a benchmark for evaluating the reasoning structures of large reasoning models, revealing insights beyond traditional accuracy metrics.
This paper introduces a novel framework for evaluating encoder roles in multi-encoder vision-language models, revealing insights into optimal configurations.
This paper presents a novel approach for generating narrative summaries from multi-modal tracking data for remote family members of older adults using LLMs.
This paper presents a method for visual instruction tuning that enhances multimodal integration in Large Language Models by embedding visual features in intermediate layers.
This paper introduces a training-free mixture-of-agents framework for multi-document summarization, leveraging LLMs and knowledge graphs to enhance performance.
This paper introduces Taiji, a novel framework for optimizing LLM-enhanced recommendation systems by addressing semantic-ID alignment and reward trade-offs.