Ben Gilmore

Buy the stack, build the evidence

Buy the software when it does the job. Keep responsibility for finding out whether it helped.

I don't want to build an agent platform if I can buy one that does the job. I've got enough things to maintain.

The bit I won't hand over is deciding whether the work made the service better.

What bothers me is seeing teams rebuild available software, then accept the supplier's definition of success. We do the expensive part and leave the judgement to somebody else.

Buy the service interface, warehouse and connectors when they meet the need. A custom build needs a reason beyond being interesting to make. Count what it will cost to keep running too.

Salesforce and Snowflake — maybe part of the future? Salesforce has service agents, and Snowflake has agents that work with data. Those are vendor descriptions; they don't prove the pair will suit the job.

Maybe is the word there. Both still have to pass the test.

Read the contract as closely as the demo. If you're paying per conversation, a conversation that fixes nothing can still cost you. Don't confuse the billable unit with a customer getting help.

A change can save staff time without reducing queries or conversations. Keep the labour saving and platform bill separate. Otherwise it's too easy to claim a saving in one line and miss a cost in another.

Here's the work I want kept: write down what changed, what it should fix and what happened afterwards. Include the changes that did nothing — especially the ones we were keen on.

A list of recommendations isn't enough. Without the follow-up, we've just found a faster way to produce more crap.

The record can live in bought software. I don't care who wrote it. I care whether we can inspect the evidence and act on it.

That evidence needs to reach beyond the supplier's own reports. Can it show where retrieval missed an existing answer? Can it identify a product decision that keeps sending customers to support?

Some failures need another team to act. They don't stop mattering because they sit outside the reporting tool.

An independent reviewer has to show their working too. Give me retrieval tests and labels checked by people who know the job. Two people agreeing on the same wrong answer is still a problem.

The practical work is keeping a set of approved test answers, checking changes before wider release and going back to measure the result. If a model grades the answers, someone has to check its judgement too.

None of that requires a new platform. It requires an owner who comes back after release, including when the result is awkward.

I could be wrong about how much needs building. If the supplier already does this well and we can retain the evidence, buy that too. A spreadsheet might be plenty for a small team.

The harder limit is authority. If nobody can act on the finding, more measurement won't fix the service. Sort out that route before spending more on the reports.

Take the last four changes to your AI service. How many can you show actually helped?