Papers
arxiv:2609.10539

IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications

Published on Sep 9
· Submitted by
YilingMa
on Sep 11
Authors:
,
,
,
,

Abstract

The study introduces a benchmark to evaluate whether research methods are specified clearly enough for implementation, finding that identifying missing details is the primary challenge for language models.

A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.

Community

Paper submitter

Excited to share IdeaAMBIG!
We ask a simple question: when an AI system is given a research idea, can it tell whether the idea is actually specified well enough to implement?
We introduce a benchmark of 660 evidence-grounded specification gaps and evaluate 13 LLMs on detecting, localizing, and resolving implementation-critical defects.
Our most striking result: LLMs are much better at fixing a gap than finding it. On real-world cases, the best model achieves only 9.6% defect recovery, but 80.6% clarification success once the defect is given. Providing the gold resolution raises downstream codification readiness from 14% to 98%.
This points to an important bottleneck for research agents: before asking models to implement an idea, we may first need them to reliably recognize what is missing, ambiguous, or inconsistent.
Would love to hear the community’s thoughts on this!

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.10539
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.10539 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.10539 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.10539 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.