Verigrey
Research

As LLMs get powerful, what about agent assurance?

By Abhik Roychoudhury, Chief Expert, Verigrey Inc2 min readUpdated

Agent assurance is an area on many peoples’ minds these days – given the role on colonies of autonomous agents are playing achieving high productivity. Such productivity can be seen in different sectors of the economy – software development, banking, financial services, insurance, healthcare, and so on. Organisational workflows are getting transformed by agentification – and agents are being rapidly deployed without full understanding of security, compliance and other risks. Hence agent validation and assurance seems more relatable in the current days.

What about the future – where the Large Language Models (LLMs) continue to improve – will the agentic harness disappear? To do crystal ball gazing along these lines – we can look into few dimensions.

Notion of Correctness

The most important issue to settle is the “intent” or the notion of correctness, the intent in particular. As the models get powerful and complex – we are able to pass off more complex tasks to LLMs via prompts or other mechanisms. Then, capturing the intent and checking that the model acted as per the intent becomes more complicated. Most commercial solutions today conduct this validation again by using another LLM in the LLM-as-judge style. However, what is really needed is a formal verification that the LLM-based system met the desired intent. Verigrey currently provides such a solution for agent trajectories. In future if the agentic harness becomes thinner, we can observe LLM executions and even modify them to achieve desired security / functionality properties. Research has already looked into such possibilities 1.

Future Harnesses

As the LLMs get more powerful we can imagine more powerful activities being taken up by future agents – that is the harnesses do not disappear but achieve more. Today there is increasing movement towards orchestration of agents – rather than working with single agents – agents instructing other agents. As the LLMs get more powerful and achieve the autonomous action sequences of an agent via a single prompt – in future we could have

  • Verification of agent orchestration, that a collection / colony of connected agents indeed achieve a complex task as per intent.

Thus, instead of testing a single agent or a full network of agents – the checking could move to a higher level of abstraction – where whether we are connecting or linking up the agents “correctly as per intent” could be checked.

Verigrey’s future-proof assurance platform can help check such orchestration with little changes. The bigger question thus, is not, whether agent assurance will be needed moving forward - since it will be needed. The bigger question is whether the assurance techniques can match up with the enhancement in LLM capability. Verigrey is uniquely positioned to meet this challenge owing to its innate focus on agent behaviors.

References