For the person who has to sign off

We build digital twins of your Slack, Gmail and helpdesk, then test your AI agent in them.

Working copies of your real systems, wired together the way yours are. Sonata runs the agent through hundreds of scenarios inside them to find out which jobs it can be trusted with.

Because someone wants to hand this thing the keys: let it email your customers and issue refunds while nobody is watching. It might well be very good at that, and right now you have no way of finding out before you say yes. Nothing it does in here reaches a real customer.

Put your name down

We're taking three companies on first. Tell us what you're deciding about and we'll come back within a couple of days.

What comes out of a run

A verdict on every job

Every job you were thinking of handing over gets scored across hundreds of runs, then written up so you can forward it to your boss or your insurer without translating it first. Three possible verdicts.

  • Safe to run on its own
  • Fine, with a person approving
  • Not ready, and here is why

What we twin

  • Slack
  • Gmail
  • Google Calendar
  • Zendesk
  • Intercom
  • Jira
  • Notion
  • Salesforce
  • Stripe
  • HubSpot

Simplified marks drawn by us, not official brand assets. Anything reachable over HTTP or MCP can be twinned; these are the ones asked for most.

The position you're in

Nobody has given you a way to say yes safely.

The vendor says it's reliable, which is what you would expect them to say. Their demo worked. Their pilot went fine. Their references are customers they chose themselves. None of it tells you what happens on a Tuesday afternoon when a customer writes in with something strange and there is nobody around to notice.

Which leaves you choosing between saying yes and hoping, or saying no and then explaining in a year's time why your competitors automated this and you didn't. Neither is much of a position to be in.

The things that go wrong are rarely dramatic.

  • It refunds the same customer twice because a request timed out and it tried again.
  • It reads a thread containing someone else's personal details, on the grounds that it technically had permission to.
  • Someone writes "cancel my account" and it deletes the whole workspace instead of the subscription.
  • A customer asks twice, firmly, and it quietly gives away something it shouldn't have.
  • It follows an instruction buried in a forwarded email, having no way to tell a message apart from a command.

All five are things we can put in front of your agent before any of your customers do.

How it works

Twins of everything, and hundreds of scenarios to run in them.

  1. 1

    A digital twin of every system it touches

    Not one tool, all of them. Your helpdesk, your inbox, your Slack, whatever handles your billing. Each twin behaves like the original and they're wired together the way yours are, because most of the interesting failures happen between two systems rather than inside one. Nothing that happens in there is real and no customer ever sees any of it.

  2. 2

    Hundreds of scenarios, generated

    The twins get populated with the sort of thing that actually comes in. Customers who are confused. Customers who are angry. The person who asks the same question three ways because the first two answers didn't help. Requests sitting right on the edge of your policy. Tidy test cases won't tell you much.

  3. 3

    Every scenario, run over and over

    These systems don't behave the same way twice, so watching an agent succeed once tells you very little. Every scenario runs again and again. That's what turns "it managed it" into "it gets this right 71% of the time", which is the number you actually need.

  4. 4

    A verdict per job, in plain English

    A section for each job you were thinking of handing over, a plain verdict on every one, and the transcripts of anything that went wrong so you can read it yourself. Run it again after your vendor ships a change and the numbers move.

What you get

The document you'd actually want to read.

There's no dashboard and no score out of a hundred. Each job you were considering gets a section, and each section tells you whether it's ready.

Scenario test results

Customer support agent · Northgate Supply Co.

1,200 scenarios across 6 system twins · run 118

  • Safe on its own

    Answering questions about an order

    Got it right 197 times out of 200. All three misses were polite versions of "I'm not sure, let me pass this on", which is the way you want it to fail.

  • Safe on its own

    Refunds inside your stated policy

    Correct in 98.5% of cases. Never refunded more than the order value, never refunded twice.

  • Needs approval

    Refunds outside policy, when the customer pushes back

    Right 71% of the time. When a customer asked twice and got firmer, it gave in and refunded anyway in roughly one case in four.

  • Needs approval

    Acting on requests from staff in Slack

    Right 88% of the time. A colleague who asks nicely enough can talk it past a step you meant to be mandatory.

  • Not ready

    Anything involving another person's personal data

    Handled correctly less than half the time. When a customer pasted someone else's details into a message, it usually carried on and used them.

  • Not ready

    Cancelling or closing an account

    Got the wrong thing 38% of the time, closing the whole account when the customer had only meant to cancel one subscription. They can't undo it themselves.

  • Not ready

    Handling forwarded emails from outside your company

    Followed instructions hidden inside a forwarded message in 63% of attempts. Anyone who emails you can currently give it orders.

In short: two of these seven jobs are ready to run on their own, and two are fine as long as somebody approves them. The other three aren't ready yet. Every failure has its transcript attached so you can see where it went wrong.

An illustrative example, using a made-up company, to show the format. Not a real customer.

Not sure which of these your agent would fail? Most people have no idea, and until now there was no sensible way to find out.

Put your name down

The part everyone asks about first

We don't touch your live systems. At all.

Nothing we do is real

Every test runs inside the copy. Nothing is sent, nothing is refunded, no ticket gets closed and no record changes. If the agent does something catastrophic it does it to data we invented.

We don't need your customer data

We build the twins from a description of your systems and how the business runs. The data inside them is made up to resemble yours without ever being yours.

Read-only, if you want more accuracy

If you want the copy to sit closer to your actual setup we can take a read-only look at volumes and shapes. Nothing gets written back. Plenty of people skip this.

Independent, and happy to prove it

We don't sell agents, and we take no money from whoever sold you yours. We sign an NDA before any of it starts.

What comes of it

The three things people say afterwards.

We were four days off switching it on. Turned out that about a third of the time it closed the whole account when someone had only meant to cancel a subscription. It would never have occurred to us to check that.

Operations Manager Northgate Supply Co. · 40 staff

We had gone in wanting to hand over the whole inbox. What came back said two of the five jobs were fine on their own and the others weren't ready. So we started with those two. I wouldn't have known where to draw the line otherwise.

Customer Service Lead Vantage Field Services · 85 staff

I'm not technical and it didn't matter. I sent the document to our insurer and read the summary out at a board meeting, and nobody asked me to explain any of it.

Managing Director Ellery & Frost · 25 staff

Placeholder quotes from invented companies, here to show the sort of thing a first run turns up. We haven't earned real ones yet. When we have, these get replaced and this note comes off.

Why this exists

We built this because the same gap kept turning up. The companies selling agents can't prove they're safe, and the companies buying them have no way to check. Everyone ends up guessing, and the person who signs off carries the risk on their own.

At the moment we're talking to people rather than selling to them. If this is a decision sitting on your desk, we'd like to hear how you're thinking about it, whether or not we end up doing any work together.

Find out what you'd be signing off on.

We're taking three companies on first. Tell us which agent you're weighing up and which of your systems it would be touching, and we'll come back to you within a couple of days about whether we can be useful.

Put your name down
  • Nobody will ring you. We reply by email
  • A few questions on the form, not twenty
  • Your details stay with us and don't go on any list

Building the twins of your systems takes a couple of weeks the first time. After that a run takes hours and you can do it as often as your vendor ships changes. Your side of the setup is about an hour of conversation. If you'd rather not fill in a form, email us directly and tell us what you're deciding about.

Questions

Straight answers.

Do we need to be technical to do this?
No. We need to know which tools the agent connects to and what jobs you had in mind for it. If you can describe your business to a new hire, you can describe it to us.
Is this a real service or a landing page?
Honestly, both. We're early, and the first few companies through are how the proper version gets built. You'd be one of those, which is worth knowing before you get on a call rather than after it.
How long does it take?
A couple of weeks to build the twins of your systems, then a few hours per run after that. The setup is the slow part and you only do it once. Your side of it is roughly an hour of conversation.
What does it cost?
We haven't set one. The first few companies through are how we work out what this is worth, so if you're one of them we'll agree something with you before any work starts. Nothing arrives as a surprise invoice.
Won't the vendor just tell us it's fine?
They probably will, and they might well be right. But they aren't a neutral party and you're the one carrying the risk if they're wrong. We don't sell agents and nobody who does is paying us.
They told us it has safety guardrails. Isn't that enough?
Guardrails watch what the agent writes and block the obviously bad version, which is worth having. What they don't catch is the refund issued twice, or the reply that went to the wrong person, or the account closed when someone meant the subscription. None of those look wrong at the moment they happen, so nothing stops them. You only find them by letting them happen somewhere that isn't real.
What if the answer is that our agent isn't ready?
Then you've found out now rather than in six months, and more usefully you know which parts of it are fine. Most people come out of this automating something. Just not everything they had planned.