This paper presents Iteris, an agentic AI system that aids in solving open problems in computational mathematics through numerical experimentation and proof generation.
This paper introduces Speculative Tool Privacy Contracts to mitigate privacy risks from speculative tool calls in language agents, addressing user intent leakage.
This paper introduces a hybrid approach using physics-informed neural networks for adaptive mesh refinement in finite-difference PDE solvers, enhancing efficiency.
This paper introduces MCP-Persona, a benchmark for evaluating LLM agents on personalized applications, addressing gaps in existing evaluation frameworks.
This paper introduces Luar, a reinforcement learning framework that enables reasoning language models to selectively translate non-English inputs for improved multilingual reasoning.
This paper introduces AGENTCL, a framework for evaluating continual learning in language agents through controlled task streams and memory design analysis.
This paper presents Langevin Speculative Dynamics (LSD), a novel method for accelerating molecular dynamics simulations without introducing relative error.
This paper introduces the HLL benchmark to evaluate multimodal agents' ability to navigate CAPTCHA verification, highlighting their limitations in human-like interaction.
This paper introduces a visual program synthesis framework that utilizes input binarization to bridge the sim-to-real gap in semiconductor inspection tasks.
This paper presents a systematic study of error propagation in large language model inference, introducing a fault-injection framework to enhance reliability.
This paper introduces GC-MoE, a novel framework for predicting cell-type-specific gene expression from histopathological images in single-cell spatial transcriptomics.
This paper introduces a local perturbation theory to explain cross-domain interference in multi-domain reinforcement learning and proposes recovery strategies.
This paper introduces PaW, a co-training framework that integrates world modeling into reinforcement learning for language agents, enhancing training efficiency.
This paper provides a theoretical framework for understanding the optimality conditions of Sparse Autoencoders (SAEs) in extracting interpretable features.
This paper introduces TabPrep, a preprocessing pipeline that enhances feature engineering in tabular benchmarks, improving model performance across various architectures.
This paper introduces SPADE-Bench, a benchmark for evaluating spontaneous strategic deception in agents through plan-action divergence in tool-use contexts.