Problem Large language models (LLMs) often struggle to access new information necessary for answering questions that fall outside their pre-training data. This paper addresses the gap in capability regarding how…
Problem The paper addresses structural limitations in large language model (LLM) agents, specifically focusing on issues such as personality drift, non-evolutionary reflection, and the absence of a self-other boundary. These…
Problem This work addresses the challenge of detecting and classifying hallucinated character spans in outputs generated by vision-language models. The authors highlight the need for robust methods to identify inaccuracies…
Problem Existing methods for scientific poster generation rely on transient prompts, which can lead to inconsistencies and drift in requirements across content and layout modules. This paper addresses this gap…
Problem The paper addresses a gap in the understanding of how intrinsic rewards can facilitate adaptation and self-organization in artificial neural systems without relying on external objectives. This is particularly…
Problem The paper identifies a gap in the governance of fast-growing agent-skill registries, particularly in the context of managing the implications of viral agent-skill ecosystems. It highlights the challenges posed…
Recent research indicates that engaging with the Google Gemini chatbot for approximately seven minutes can significantly diminish conspiracy beliefs regarding contemporary crises. This finding suggests that conversational AI may be…
{ "meta": "This paper investigates the robustness of linear probes in medical QA across linguistic register, medical specialty, and corpus shifts.", "body": "## Problem\nThis work addresses the gap in understanding…
Recent research highlights a critical limitation in AI coding assistants, specifically Claude Code and Codex, revealing their inability to accurately perceive time. The study indicates that these AI agents systematically…
LAION has introduced the Big Video Dataset (BVD), a substantial resource for AI research comprising 80 million videos and 10 million hours of runtime. This dataset includes 55 million auto-described…
In recent explorations of large language models (LLMs) for vulnerability research, Jordy Zomer identified a significant limitation: LLMs often lose track of established facts during prolonged investigations, leading to erroneous…
Problem The paper addresses the gap in reproducibility within hybrid Earth system models that integrate AI techniques. As AI becomes increasingly prevalent in environmental modeling, the complexity introduced can hinder…
Problem Current large language models (LLMs) used in molecular design lack proper calibration for uncertainty, which limits their effectiveness in experimental discovery. This paper addresses this gap by proposing a…
A recent study conducted by researchers at the Wharton School highlights the unpredictability of AI shopping agents, suggesting they are not yet reliable for making purchasing decisions on behalf of…
The OpenAI Blog reports on a randomized study involving over 1,000 students that explores the effects of ChatGPT and critical thinking training on academic performance. The research aims to assess…
Problem This work addresses the gap in predicting cellular responses to untested drugs, particularly in the context of cancer therapeutics. Existing models often lack the ability to generalize effectively to…
Problem The paper addresses the gap in existing AI frameworks that lack the ability to autonomously adapt to dynamic environments, drawing inspiration from biological organisms. It highlights the need for…
Problem — The paper addresses the unexamined reliance on generative AI in cognitive tasks, leading to a gap between presented outputs and defensible knowledge. This phenomenon, termed 'epistemic debt,' emerges…
{ "meta": "This paper establishes the W[1]-hardness of computing $Lp$-Lipschitz constants for two-layer ICNNs, resolving an open problem in parameterized complexity.", "body": "## Problem\nThe paper addresses the challenge of computing…
Problem This preprint addresses the gap in understanding how AI professionals make sense of the rapid advancements in artificial intelligence and the societal implications that arise from these developments. The…
The Hugging Face Blog provides an overview of the Granite 4.2 large language models (LLMs) developed by IBM, highlighting their architectural innovations and performance enhancements. The article emphasizes that Granite…
Problem — This preprint addresses the gap in understanding why Large Language Models (LLMs) sometimes fail to exhibit ethical behavior, particularly in scenarios where their objectives of helpfulness and harmlessness…
Problem This work addresses the challenge of compressing large-scale scientific measurements into compact representations while preserving fine-scale details. The authors highlight a gap in existing methods that struggle to maintain…
Problem This work addresses the gap in existing exoskeleton control systems, which often rely on task-specific programming and lack adaptability to varying user needs. The authors highlight the need for…
Problem — This work addresses the gap in efficient data compression methods for scientific datasets, particularly focusing on preserving fine details and signal fidelity. The authors propose a novel approach,…
A recent theoretical study posits that the integration of AI, particularly language models, into scientific research may paradoxically result in a decline in the quality of publications. The research indicates…
A recent study conducted by researchers at Princeton University and UC San Diego investigates the impact of 'skills' on AI agents, revealing that these skills enhance performance primarily through structured…
Recent research has revealed significant shortcomings in existing world models such as Sora and Genie, which primarily focus on simulating physical environments while neglecting human beliefs, desires, and emotions. This…
Researchers from the UK AI Security Institute have employed psychometric methods to critically assess the efficacy of existing safety benchmarks for language models. Their findings indicate that these benchmarks fail…
{ "meta": "This paper analyzes test-time scaling in language models, revealing exploitation as the primary bottleneck in generating high-quality outputs.", "body": "## Problem\nThis work addresses a gap in the literature…