← All writing

What makes an AI agent reliable

By Rajkiran Panuganti

An AI agent can write a convincing account of work it has not finished. It can describe the report it meant to save or the change it meant to deploy. The explanation can be clear while the job remains incomplete.

I want the agent's final claim to match something a person can check outside its answer. That sounds obvious. It has consequences for how we design the whole system, starting before we choose a model.

Start with the finished report

Suppose a user asks for a customer report. This is a hypothetical example, but it is a useful design exercise. What would have to be true before we could call the task complete?

Requirement Evidence
The report exists Open the saved document
It covers the requested period Check the dates against the request
It uses the right customer records Trace a sample of entries to the source
Its totals are correct Recalculate the totals
The intended reader can access it Open it through that reader's access path

The last requirement is easy to miss. A report that opens for the agent's account can still be useless to its reader. Checking from the wrong account gives us evidence for a different claim.

Once we write these conditions down, missing parts become visible. The agent needs the source records, permission to save a document, and a way to test access. A stronger model may help it reason about the report. It cannot supply records it cannot read.

This is why I would define the finish line before spending time on prompts. Otherwise we risk optimizing the agent's explanation of success while leaving success itself undefined.

Check the place where the work happened

Two research benchmarks offer useful examples of this approach. In the 2024 OSWorld paper, tasks run in real computer environments and have scripts that evaluate the resulting work. Tau-bench, also introduced in 2024, tests conversations with users, tools, and domain rules. Its evaluation compares the final database state with the intended result.

The environments differ. The design choice I take from them is to check the effect of the action independently of the agent's account of it.

For a website, that means opening the published page and using the flow a visitor uses. A server response alone does not tell us whether the page contains the requested change. For our customer report, it means opening the saved document, checking the data, and testing the reader's access.

Those checks also need care. A test that accepts any document with the right filename would miss an empty report. A test that demands one exact layout might reject a correct report with a different presentation. The test should establish the user's requirements and leave room for valid ways to meet them.

Keep failed attempts in the count

One successful attempt gives us a working example. To understand consistency, we need repeated attempts. Tau-bench's repeated-trial measure asks whether an agent succeeds across multiple runs of the same task. That is a different question from whether it can eventually produce one successful run. Tau-bench paper

In a product evaluation, I would keep the first attempt visible even when a retry succeeds. I would also record what a person had to fix. An agent that completes a task after several retries and manual repairs has a different operating cost from one that finishes without help.

The starting point matters too. If a demonstration uses accounts and files prepared by hand, test the path that prepares them. Otherwise the demo tells us about work done after setup, while the first user may fail during setup.

What happens after the connection drops?

Now interrupt the customer-report task just after the document is saved. The outside service accepted the action, but the agent did not receive the confirmation.

On restart, repeating the last action could create a second report. Assuming it succeeded could leave the user without a report if the save actually failed. The agent needs to check the outside system before deciding which case it is in.

My design preference is to save the goal, completed actions, evidence, and pending work. A restart can then begin with a specific question: does the report already exist, and does it meet the request?

Saved state is a design proposal, not a measured improvement claimed here. It can preserve mistakes as well as progress. Its effect on reliability needs testing, including cases where saved assumptions no longer match the outside system.

Ask for help without losing the task

Sometimes the missing piece is a human decision. Perhaps two customer records conflict, or the requested sharing permission is unavailable. The agent should say which decision is needed and which part of the report depends on it.

It should also preserve the work already completed. A person should be able to answer the question without reconstructing the task from a vague error. Work that does not depend on the answer can continue.

For each run, I would keep a small record of the request, result, evidence, elapsed time, cost, and human help. Reviewing the failures should tell us which checks are missing and where the agent's final answer is too confident.

When your agent says it finished, what can the user open and check?

Related reading

Read more essays on rajkiranpanuganti.com, or subscribe to My thoughts on LLMs.

Read this every weekSame analysis, published as a LinkedIn newsletter. Substance over hype, every Friday.
Subscribe on LinkedIn ↗