UMass Amherst's IdeaLens Catches AI-Originated Ideas Hidden Behind Human Prose

A new detector called IdeaLens tries to answer who thought up a document's ideas, not just who typed the words.

·
·
·
UMass Amherst's IdeaLens Catches AI-Originated Ideas Hidden Behind Human ProsePRO
Read2 min
TypePaper
SubtopicRed Teaming
  • IdeaLens detects whether a document's ideas came from AI, not just the words
  • Documents are reduced to paraphrased outlines so classifiers cannot cheat on prose style
  • Trained on 1M FineWeb documents with silver labels from the Pangram detector
  • Flags 68% of human-written stories built from AI plans; Pangram flags only 8%
  • Correctly backs off from 95% to 7% when humans supply detailed plans to LLMs
  • Code, models, and a demo are live at ideadetector.ai

IdeaLens targets AI provenance at the idea level

AI-writing detectors usually estimate whether a model produced a document’s wording. Schools, journals, and newsrooms often need to assess the source of its thesis, examples, and reasoning. The IdeaLens paper, from researchers at UMass Amherst and collaborating institutions, introduces a detector designed to classify that intellectual provenance.

Policies care about provenance

Tools such as Pangram and GPTZero return document-level estimates based largely on patterns in the prose, including word choice, syntax, and rhythm. Those signals can identify text drafted directly by a model, but they reveal less about who developed the underlying argument.

Mixed-authorship workflows expose that limitation. A person can brainstorm with ChatGPT and write the result independently, producing human prose from AI-derived ideas. Someone else can create a detailed plan and ask Claude to turn it into polished paragraphs, producing model-written prose from human ideas. Academic and editorial policies may treat those cases differently even when a conventional detector gives the opposite verdict.

How IdeaLens hides the prose

IdeaLens changes the classifier’s input by converting each document into a paraphrased outline. The pipeline works as follows:

  1. The document is divided into meaningful sections or claims.
  2. Each item receives a discourse role, such as thesis, counterexample, or anecdote.
  3. The item’s contribution is summarized in language that minimizes overlap with the source.
  4. The classifier receives the resulting outline without access to the original document.

Pro article

This story is for Pro members

You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.

Trending
  • No trending articles

Comments

avatar

Next Reads