This paper investigates the developmental trajectories of attention-head circuits in 1B-class language models, revealing distinct patterns of emergence and capability formation.
This paper introduces Harness-1, a 20B parameter search agent that utilizes a state-externalizing harness to improve retrieval performance in reinforcement learning.
This paper presents a minimax-optimal algorithm for policy regret in partially observable Markov games, addressing strategic adversaries in sequential decision-making.
This paper introduces SIRI, a self-internalizing reinforcement learning framework that enables LLM agents to autonomously discover and utilize skills without external dependencies.
This paper introduces local preferential Bayesian optimization methods that enhance efficiency in high-dimensional optimization using pairwise feedback.
This paper introduces sampling techniques for empirical pairwise loss estimation, achieving performance comparable to full pairwise evaluations with reduced computational cost.
This paper introduces a dual-encoder architecture with Choquet integral fusion for improved underwater acoustic classification, enhancing efficiency and interpretability.
This paper introduces Distribution Shift Bias Reduction (DSBR) to mitigate prediction bias in medical imaging during test-time adaptation without model collapse.
This paper introduces a novel method for simultaneous optimization of constants and expression structure in GP-GOMEA for enhanced symbolic regression performance.