Skip to content

Shell framework restructures intelligent agent safety through adaptive learning of operational paths

Share
Shell framework restructures intelligent agent safety through adaptive learning of operational paths

Listen to this article

Read by Anchor

A team of researchers consisting of Wanqing Qiu, Qinghua Mao, Yu Li, Jiao Liu, Shen Zhang, Dadi Guo, Yanxu Chu, Qingyu Liu, Litao Yuan, Shi Lin, Shanfeng Zhu, Yanwei Fu, Jing Shao, Xia Huo, and Dongrui Liu published a new paper on arXiv on August 10, 2026, in the field of artificial intelligence, addressing the operational security of large language model agents. The study clarified that the safety of these adaptive systems does not rely solely on the internal model weights, but fundamentally depends on the operational control structure surrounding them, which manages context, memory, tools, permissions, and runtime control. Traditional security mechanisms treat this structure as a fixed element after deployment, limiting its ability to evolve and adapt to newly emerging risks during continuous execution.

The researchers pointed out that the overlap of functions among different components within current operational structures leads to ambiguity in determining security responsibility among those elements, making it extremely difficult to perform targeted local development for specific parts.To address this structural gap, the team proposed a framework called SHE (Security Hierarchy Evolution), or Shell, which is a system that learns evolving security boundaries based on the execution and operation paths of intelligent agents.This framework relies on decomposing the operational structure into four independent elements with explicit security responsibilities, including the system's basic navigator, rule bank, security memory, and tool usage policy, which defines clear functional boundaries that allow for independent development of each component without affecting the rest of the system.

The Shell system relies on a development loop guided by responsibility identification, where this loop converts failures and errors that occur on operation paths into precise structural diagnoses. Based on these diagnoses, the system receives targeted improvements to the specific boundaries of each of the four elements. The developed structures then undergo a precise selection and filtering stage based on integrated verification of security and operational utility at the same time, to confirm the absolute balance between reducing security risks and maintaining the intelligent agent's ability to accomplish tasks efficiently and effectively.

Experiments conducted on the Agent-SafetyBench evaluation scale demonstrated the efficiency of the new framework, achieving a 3.1-fold decrease in the success rate of attacks and failures compared to the static SafeHarness security structure, with a noticeable improvement in utility and intact functions during normal task performance.The results also showed the developed structure's ability to generalize and handle new, unseen risks through the AgentHarm independent harm evaluation bank, as well as its transferability and ability to work across multiple agent models without requiring additional development or training operations.

Don't miss the next story

Subscribe for updates