Skip to content

VAKRA: A new standard reveals declining performance of AI agents in complex reasoning across programming interfaces

Share
VAKRA: A new standard reveals declining performance of AI agents in complex reasoning across programming interfaces

Listen to this article

Read by Anchor

A research team consisting of Aniketa Rajaram Nair, Anupama Murthy, Benjamin Elder, Sewoong Hoo, Ravi Gupta, Abhinav Jain, Pravien Vynatheya, Abdulhamid Adedeji, and Danish Contractor published a scientific study on the Archiv platform on August 12, 2026, introducing a new evaluation standard called VAKRA. This standard is specifically designed to measure and test the capabilities of artificial intelligence agents in multi-step reasoning across application programming interfaces and knowledge retrieval under tool usage policy constraints. The standard aims to fill the existing gap in evaluating artificial intelligence agents directed at enterprise and business environments, where realistic tasks require concurrent linking and reasoning between structured programming interfaces, document sets, and diverse documents, after previous standards tested each of these capabilities separately and independently.

The VAKRA standard includes an operational database of over 8,000 executable application programming interfaces distributed across 62 different fields. The researchers divided the test tasks within the standard into three levels of increasing difficulty to cover various operational interaction scenarios. The first level focuses on testing diverse interaction patterns with programming interfaces, while the second level moves on to evaluating multi-step reasoning across composite structured application programming interfaces. The third level, which represents the highest level of difficulty, tests the agents' ability to perform multi-source reasoning that adheres to tool usage policy constraints formulated in natural language.

To achieve the highest degree of accuracy and objectivity in evaluation, the validity of answers is verified by re-executing the predicted tool calls by models against live and direct application programming interfaces, taking into account the accommodation and acceptance of multiple valid execution paths to reach the desired outcome without imposing a single programming path. The study also relied on the use of a unified and fixed evaluation framework based on the ReAct structure, in order to isolate the model's inherent capabilities and monitor its power in reasoning and thinking in an abstract manner, without being affected by structural improvements or modifications specific to the agent's programming structure.

The comprehensive evaluation included testing the latest state-of-the-art models and open-weight models, and the results revealed a noticeable decline in the efficiency of these models when dealing with composite tasks. While the best laboratory model recorded a success rate of only 70.4% in single-step tasks based on a single endpoint, this rate decreased to between 50% and 51% when evaluating performance on aggregate and composite programming interfaces. Analytical data showed that the models' performance deteriorates and decreases by more than 50% as the depth of reasoning and sequence required to complete the task increases.

The most severe failures occurred when facing questions constrained by usage policies, as the success rate dropped to only 2.4% when dealing with queries that cannot be answered based on the specified data and policies.The analysis of execution traces revealed that the sources of failure do not lie in the technical mechanics of tool invocation, but rather clearly concentrate in the stage of linguistic reasoning, represented by disambiguating entities and precisely linking information across different sources.The researchers made the source code and dataset of the standard available in an open manner to enhance future research in this field.

Don't miss the next story

Subscribe for updates