Modern large language models excel at tasks that require recalling and applying existing human knowledge. But a truly intelligent system should be able to go beyond what it has learned: ask questions no one has asked before, and turn hypotheses into verifiable scientific results. It is precisely this ability — or rather, its absence in current AI — that a new benchmark called ASI-Bench set out to measure.
From knowledge reproduction to creating new knowledge
Researchers have long noticed that modern models are trained to compress and combine information from vast datasets created by humans. However, superintelligence, as the authors understand it, is not just a more powerful "re-teller" but a system capable of independently exploring the unknown and producing new knowledge. Traditional tests almost always boil down to checking how accurately a model reproduces learned information and whether it can follow detailed human instructions. They do not answer the key question: can AI act as an independent scientist?
This situation called for a fundamentally different tool that would test precisely scientific creativity and autonomy — and the first step in this direction was made by ASI-Bench.

How the benchmark works
Developing the new test took over 31,000 person-hours of collaborative work by more than 40 experts. As a result, the benchmark includes 60 project-based research tasks from 11 scientific fields — from natural sciences to engineering disciplines. Each task is structured as a small research project that requires carrying out work from an initial hypothesis to a verifiable result.
A key feature of ASI-Bench is the gradual removal of methodological support. For each task, the authors prepared several levels of guidance: from maximally detailed instructions where every step is spelled out, to a formulation in which the agent must choose its own method and even determine the direction of research. Thus, the test makes it possible to see how far AI is willing to go without human hints.
To eliminate errors and "memorization" of answers, all tasks undergo multi-stage verification: expert review, audit using another AI system, execution in an isolated sandbox, and separate validation of scoring results. This makes the results more reliable compared to conventional tests.

First results: dependence on humans
The authors ran 18 "agent–model" configurations at the state-of-the-art level through all task levels. The results proved telling. With full methodological guidance, the average score was 50.91; when the agent was left only with a method, it dropped to 29.10. When the model had to determine the method on its own, the score fell to 26.62.
Such a sharp decline suggests that even the most advanced systems are not yet ready to conduct end-to-end scientific research without detailed human hints. They are strong at executing clearly defined steps but noticeably struggle when they need to take the initiative. This is an important signal for anyone expecting the imminent emergence of superintelligence: progress in knowledge reproduction does not equal progress in autonomous scientific discovery.
Open platform and next steps
The developers have made ASI-Bench open to the scientific community and invite researchers from around the world to add new tasks. This will allow the set of fields and scenarios to expand, making the benchmark more comprehensive. In addition, the openness of the test makes it possible to track the progress of AI systems over time: today's results will likely become a baseline for comparison with future models.
Despite the still-modest scores, the main takeaway of this research is not disappointment but the emergence of a tool that, for the first time, makes it possible to measure that very "research" quality of AI. What remains is to see whether new models and agent architectures can close the gap between guided execution and independent scientific work.



