Can AI produce better research ideas? What happens after the brainstorm matters

Steven Watson | EdgeLab | Research and Scholarship

Publication note, 19 September 2026: This review focuses on the original 2024 study. The authors subsequently investigated implementation in The Ideation-Execution Gap (first posted in June 2025). Its abstract reports that 43 researchers executed assigned ideas and that AI proposals lost more of their initial evaluation advantage after implementation. That follow-up directly addresses the distinction discussed here; the proposed collaborative comparison below would extend this work. Read the follow-up.

A research idea can sound brilliant over coffee and become much less convincing when somebody tries to investigate it. Equally, an apparently ordinary question can open an unexpected path once the work begins. This gap between a promising suggestion and a contribution to knowledge is central to how we should assess AI in research.

Chenglei Si, Diyi Yang and Tatsunori Hashimoto’s Can LLMs Generate Novel Research Ideas? offers a useful starting point. Their 2024 study involved 104 researchers, including 49 idea writers and 79 reviewers, with some participating in both roles. Experts assessed anonymised proposals for prompting research in natural language processing. AI-generated proposals received significantly higher novelty ratings than human proposals. Slightly lower feasibility ratings for AI were not statistically significant. Read the study.

The use of expert reviewers and standardised presentation makes this a serious contribution. It tests a consequential capability while leaving a larger question open: what happens when an attractive proposal encounters the demands of actual research?

One methodological qualification matters. The AI system generated 4,000 candidate ideas per topic before filtering and ranking; each human submitted one proposal. The comparison therefore includes the advantages of large-scale generation and selection. Study methods.

That is a legitimate practical advantage. A researcher may care more about whether a workflow produces useful possibilities than about whether every participant had an identical process. But claims about superior creativity require greater care. We should ask what resources, selection procedures and support produced the result.

There is also a distinction between encountering an unfamiliar suggestion and establishing its intellectual value. Imagine a proposal to improve classroom discussion by asking AI to represent several competing viewpoints. It might initially seem imaginative. Investigating it would require asking whose viewpoints are represented, whether pupils recognise the differences, how teachers intervene and what would count as a better discussion. These questions could transform the original idea.

The example is hypothetical, but it illustrates the work that a novelty rating cannot settle. Research develops through clarification, criticism, practical difficulty and revision. Sometimes the most important contribution is discovering why an appealing approach fails.

From an autopoietic ecology perspective, this is where the discussion becomes particularly interesting. AE directs attention to how activities and relationships sustain and transform one another. Applied to research, it invites us to examine the connections among researchers, AI systems, texts, methods, institutions and the situations being investigated.

An AI suggestion acquires significance as people interpret it, challenge its assumptions, connect it to a problem and establish ways of testing it. The model’s contribution can be substantial within that process. Understanding the contribution requires following what happens to it.

This is an interpretation through AE, rather than a conclusion demonstrated by the experiment. It changes the question we might ask next: which forms of human–AI collaboration make inquiry more searching, more accountable and more capable of changing direction?

For EdgeLab, that question suggests a concrete design agenda. A research assistant could help preserve alternative explanations, identify assumptions and record why a proposal was revised or abandoned. It could ask what evidence would count against an attractive claim. It could help participants notice when they agree on terminology while understanding the problem differently.

Such functions would need evaluation. A useful study could compare researchers working independently, researchers using an AI idea generator, and researchers engaged in sustained dialogue with an AI system. Following complete projects would allow assessment of the quality of the questions, the handling of setbacks and the reliability of the resulting claims. Keeping records of revisions would also reveal where the collaboration actually made a difference.

Si and colleagues provide reason to take AI-assisted ideation seriously. The opportunity for research is to turn that capacity into better inquiry. A promising idea earns its place through the work it makes possible—and through the criticism it can withstand.

Paper reviewed: Si, C., Yang, D., and Hashimoto, T. (2024). “Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.” arXiv:2409.04109, version 1, 6 September 2024. This review concerns that version, with a publication note acknowledging the subsequent execution study.

Leave a comment