Skip to content

AutoDesign: Improving operational systems gives AI agents the ability to excel in academic design

Share
AutoDesign: Improving operational systems gives AI agents the ability to excel in academic design

Listen to this article

Read by Anchor

A research team including Yaxin Luo, Haobin Jiang, Jialv Zou, Xu Huang, Wenhao Yan, Haodong Li, Zhengrong Yue, Jing Li, Xiaofu Chen, Xiaohan Zhao, Jiacheng Liu, Jiacheng Cui, Zhiqiang Shen, and Xiaotong Li presented a new framework calledAutoDesignThe framework focuses on improving the meta harness optimization for agentic design in long-term tasks, aiming to convert multimodal sources into dense, structured informational outputs. The research paper, published in arXiv on 13 August 2026, falls under the fields of computer vision, artificial intelligence, and natural language processing.

The paper is based on a core concept that treats the transformation of multimodal data into operational, long-term agentic structural media as revolving around a system that combines the model and the operational harness (model harness system). The authors assume that an ideal system should align with humans’ pre-existing design concepts and that reusable experience accumulation through experimental exploration can drive a repeatedly self-optimizing process. In contrast, current patterns suffer from staticness and rigidity and lack these adaptive capabilities, limiting their efficiency in handling complex designs.

To address this shortcoming, the framework proposesAutoDesignA mechanism in which the meta harness optimizer directs a code agent to make repeated modifications and improvements to the operational system based on feedback derived from rollout cycles. To evaluate the framework and test its success, the researchers focused on a practical application of converting academic papers into scientific posters, launching a benchmark calledPosterBenchThe main track of this benchmark consists of 100 academic papers spread across five different scientific disciplines, plus a small sampled subset called PosterBench mini that includes 10 papers for evaluating control experiments.

Test results on the main track of the PosterBench benchmark showed AutoDesign outperforming the closed-source commercial system Claude Design by achieving a top score of 78.32 points, a margin of 7.45 points. When the framework was tested across seven control configurations combining different models and code agents, integrating the acquired DesignHarness system continuously improved performance, raising the benchmark’s average score from 54.99 points to 67.39 points, an increase of 12.4%.

In terms of operational efficiency, the system ran within a fully independent long-term loop, executing 253 tool calls and 11 modification rounds in under 40 minutes at a total cost of less than $3 USD. In human evaluation, the outputs reached the quality level of posters presented at academic conferences. A blind human study, in which participants could not identify the source of the posters, confirmed AutoDesign’s superiority and its highest human preference rate among all compared systems.

Don't miss the next story

Subscribe for updates