Notableother

IdeaLens: Detecting AI Ideas in Long-form Writing

Rishanth Rajendhran, Minjoon Choi, Jenna Russell, Ramya Namuduri, Deniz Bölöni-Turgut, Marzena Karpinska, John Wieting, Mohit Iyyer

Published
Oct 5, 2026 — 17:43 UTC

Problem

This paper addresses the challenge of detecting the origin of ideas in long-form documents, specifically distinguishing between human-generated and AI-generated ideas. The authors highlight a gap in existing literature regarding effective methods for idea provenance detection, particularly in the context of long-form writing. The work is presented as a preprint and has not undergone peer review.

Method

The authors propose IdeaLens, a specialized detector designed to identify the provenance of ideas within documents. The architecture leverages a unique representation of documents as outlines, which pair discourse roles with brief descriptions. This approach minimizes word-level overlap, enhancing the model's ability to discern the source of ideas. The training data consists of 1 million FineWeb documents, which have been annotated with silver labels sourced from Pangram. However, the authors do not disclose the specific training compute used in the development of IdeaLens.

Results

The results demonstrate significant improvements in detection capabilities. The AI flag rate for documents drops from 95% to 7% when detailed human plans are included, compared to a baseline of 92% from Pangram 4. Additionally, when detecting human stories derived from AI plans, IdeaLens flags 68% of these as AI-generated, in stark contrast to the 8% flagged by Pangram 4. The model exhibits strong detection rates with low false positive rates across 19 different benchmarks, indicating robust performance in various scenarios.

Limitations

The authors do not report any limitations in their study. However, the absence of disclosed training compute may hinder reproducibility and scalability assessments of the proposed method.

Why it matters

The implications of this work are significant for downstream applications in content verification, academic integrity, and the broader discourse on AI-generated content. By providing a reliable method for detecting the provenance of ideas, IdeaLens could enhance the ability to assess the originality of written works, thereby contributing to the ongoing discussions about the ethical use of AI in creative processes.

Summarised from the paper by the Turing Wire Research Desk. The full paper has the complete methods and results.

Source: arXiv cs.AI