Skip to content

Paper & Citation

Citation details and a concise overview of the ASI-Bench paper and benchmark contributions.

Benchmark Overview

Title ASI-Bench: At the Dawn of Artificial Superintelligence
Preprint arXiv:2608.17271
PDF Read the paper
Repository github.com/apexin-ai/ASI-Bench
Website asibench.apexin.ai

Cite ASI-Bench

Please cite the ASI-Bench arXiv preprint:

BibTeX arXiv:2608.17271
@misc{zhou2026asibenchdawnartificialsuperintelligence,
      title={ASI-Bench: At the Dawn of Artificial Superintelligence},
      author={Junwei Zhou and Zhen Sun and Binyu Li and Jiangyu Zhou and Yuexi Pan and Hengyu Wang and Honghe Ren and Xiaohan Jia and Xueyang Zhou and Xiaoyu Cao and Yongchao Chen and Yuanning Feng and Junhao Wu and Cheng Zhang and Sijia Chen and Haoyu Xue and Chengsong You and Huan Wang and Koutian Wu and Peigan Gao and Jiakun Wu and Wenzhe Li and Ergan Shang and Qingyuan Zheng and Jingjing Zhou and Ruixuan Jia and Yan Xu and Hongrui Zhang and Xiao-Han Ma and Zhengxiang Cheng and Yuexing Hao and Liting Mai and Xianglin Ji and Wenjun Zhang and Zhuofan Chen and Yixiao Huang and Chi Wang and Wenyue Hua and Yilun Hao and Yuantao Zhai and Ziyan Zhao and Jingyan Xie},
      year={2026},
      eprint={2608.17271},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.17271},
}

The preprint was released on August 18, 2026.

Key Contributions

The paper presents:

  1. Three dimensions of intelligence — a joint evaluation of general intelligence, innovation, and autonomous execution
  2. Controlled guidance withdrawal — the same research objective, data, outputs, and scoring are held fixed while human methodological guidance is progressively removed
  3. Project-level scientific research — 60 long-horizon tasks across 11 domains requiring method selection, implementation, experimentation, and verifiable artifacts
  4. Rigorous construction and validation — expert review, AI-assisted auditing, sandbox execution, and task-specific scorer validation
  5. A diagnostic result — the largest average drop occurs from B1 to B2, identifying method operationalization as a larger bottleneck than method selection or distractor robustness