AIBIS AI solutions & training

All posts

The agent finished the job. Could you prove how, a month from now?

In a recent experiment an AI agent worked on its own for nearly half an hour and returned a finished result — while the user never saw what code the machine actually ran along the way. The result was good. The more interesting question for us is a different one: once the same logic reaches your quotes, invoices or customer emails, what exactly do you show a month later when someone asks why?

The result alone isn't enough

A human employee leaves a trail without trying. There's an email where the price was discussed, a calendar entry, someone's confirmation in a chat, a signature on a document. When a customer disputes something, we can find those. An agent doesn't do that — it hands over the final output and says nothing about which data it used, which model it ran on, or what intermediate steps it took.

For most small companies the problem isn't that AI makes mistakes. The problem is that the mistake can't be reconstructed afterwards. If a quote went out with the wrong margin, we want to know whether the fault came from the input data, the instructions or the model. Without a trail, the only remaining method is the old one: guess and hope.

Four things worth recording

In our experience you don't need a full audit platform — four items are enough, and they fit into an ordinary database, a spreadsheet or your existing CRM. First, the input: what data and what instruction went in. Second, the tooling: which model, which version, through which provider — the same model accessed through different intermediaries doesn't always behave the same way, so "we used AI" is not an answer. Third, the output in two forms: what the machine produced and what the human changed before it was sent. Fourth, who approved it and when.

That fourth point is the most important and the cheapest. Once a name is attached, the AI output becomes a company document rather than anonymous machine text. It also changes behaviour immediately — people approve more carefully when their name is on it.

Decide in advance what needs a human signature

In the US a defence lawyer was sanctioned after presenting invented citations and witnesses produced by a language model. The lesson isn't "don't use AI". The lesson is that there was no written rule about which outputs have to cross a human desk before leaving the building.

We suggest two simple lines. First: anything that leaves the company under your name — to a customer, a partner or an authority — needs human approval. Second: anything that moves money, changes a database or sends something out at scale needs human approval. Everything else — internal summaries, drafts, idea lists, a first version of code — can run freely. That fits on a single page and is far more useful than a long AI policy nobody reads.

Where we draw the line ourselves

An honest caveat: not everything should be logged. A log is a data store, and if customer personal data or health and financial details leak into it, we've simply created a different problem. So we set a retention period up front, decide which fields are stored masked, and keep access to the log narrow.

The second line: don't build traceability before you have a process. If half the team is still exploring what the models can do at all, building an audit trail is premature — that comes once a process has stuck and starts repeating. During the experimentation phase, one agreement is enough: no AI output goes straight to a customer.

A practical first step this week: write down the three places where AI already does or writes something in your company, and next to each, whether the result could be reconstructed a month from now. Wherever the answer is "no", that's exactly where to start — before the agents get even more independent.

All posts

Next step

Book a free 20-minute call.

Book a free 20-min call