Getting started
Map → Replay → Prove → Run, and back to Map. The first time round, Map and Replay happen with our team — that's Foundations. After that the product runs all four stages on your own live traffic, and nothing changes without going through Prove.
Your business into the product
The product at work
Stage 1
Map is where your operation becomes something agents can reason from: your ontology, your workflows with their policy gates, your grounding rules, and the mapping from your systems into the Context Layer. All of it is readable, and all of it is yours.
We work with the people who do the work to write down how it runs: the things your business deals with, the steps they take, the judgement calls and the rules. The interviews are structured, not open-ended. Getting people to articulate how they really make decisions is a skill in its own right, and Foundations is built around it.
What they tell us is encoded into the product. Most of this is generated with AI, supervised by people who know your industry — which is why it's fast — and reviewed with your team, who correct it in their own words.
You get: your operation written down in your own terms, and connected to the systems it runs on, ready to be checked against what actually happened.
Changes arrive in two ways. Groundwork proposes them from what the last Run recorded — a concept it keeps meeting, a workflow branch for a request type it keeps escalating, a rule worth tightening — and you accept, edit or reject each one. Or your team makes them directly, with Operator, an agent for your own people that knows your Context Layer and helps draft the change.
Either way a change is a new version of the workflow. The old one keeps running until the new one has been proved.
Stage 2
Replay runs real cases against the map: what came in, what your team did with it, and what a good outcome looked like in your terms. People describe how the work is supposed to go; the cases show how it actually goes. Where the two differ, the map is corrected — it only ever changes on evidence, never on a hunch.
We take a representative sample of your real historical cases — conversations, tickets, requests — and work through them against the map with the people who handled them. This shows how your business actually behaves (as opposed to how the process doc says it does), where the judgement calls are, and what "good" looks like in your terms. Every gap becomes a correction to the map, and the cases themselves become the evaluation set.
You get: a working Context Layer, corrected against your own history; an evaluation set built from your own cases; agents configured and ready to prove; and an honest view of what Groundwork will automate now, what stays human-in-the-loop, and what shouldn't be touched. This is the Foundations deliverable — it's yours, in your Groundwork account, whatever you do next.
The product replays your own live traffic. Every conversation the agents handled, and every one they escalated, is on the record with its attributes and the decisions behind it. Escalations show where the map is incomplete, corrections your team made show where it is wrong, and new request types surface as they appear.
A conversation that went well can be pinned as a golden case, its outcome becoming the expected one. A conversation that went badly can be corrected by hand and kept as a guard against the same failure. Either way the evaluation set grows from your own history, not ours.
Stage 3
Prove runs agents against the evaluation set — cases from your own history, each with the outcome your team expects — until they meet the bar you set. You see how they would have handled last quarter's conversations before they touch this quarter's. Nothing goes live on a hunch.
Foundations ends here. Agents run against the evaluation set built in Replay, and you see, case by case, what they would have done: the same as your team, or different, and why.
The bar is yours to set — which case types the agents must handle, what accuracy is acceptable, and what stays human. Go-live is a decision you make on the evidence, not a date on a plan.
Every change goes through Prove before it is live. A new workflow version, an edited policy, a renamed concept — each is run against the evaluation set, and the diff shows exactly which past cases would now be handled differently. If the only differences are the ones you intended, it goes live. If not, it doesn't.
Stage 4
Agents handle the routine majority, hand the rest to your team with full context, and never step outside your policies.
Go-live is not a switch-over. Workflows are enabled one at a time, so you can start with the request types you are most confident about and add the rest as the evidence comes in. Actions that write to your systems are staged for a person to approve until you decide otherwise.
Agents work from the current proved version of the map. Escalations arrive with the conversation summarised and its attributes attached, and every reply and decision is written to the record — including the ones your team makes.
That record is what the next Map works from.
That is the loop. Run produces history; Map changes on the evidence in it; Replay checks the map against what actually happened; Prove checks the change; Run carries it. Foundations is the first turn, done with you. Every turn after is the product, on your own traffic — so the Context Layer gets more accurate over time instead of being set up once and left to drift.
Foundations
What happens
The first Map and Replay, run with your team as a defined piece of work, agreed upfront.
What you receive
A Context Layer in your language, an evaluation set built from your own cases, and agents configured and proved against it. It's yours, in your Groundwork account, whatever you decide next.
Duration
1-2 weeks from kickoff. Map runs in week 1, Replay in week 2.
Investment
Fixed on the scope confirmed at the scoping call, not on time spent. You see the number before anything starts.
Week by week
Step 01
Thirty minutes. You tell us the request types that eat your team's time and roughly how many of them there are; we confirm the scope, the dates and a fixed price. Nothing starts until you have seen the number.
Step 02
We interview the people who handle your cases, write your ontology, workflows, policy gates and grounding rules into Groundwork, and map your systems into the Context Layer. What we need from you: a few hours from two or three of those people, and read access to the systems involved.
Step 03
We work through a representative sample of your cases against the map and correct it wherever what happened differs from what was written down. What we need from you: an export of recent cases — conversations, tickets, requests — and a reviewer who can say "that's not how it works here".
Step 04
Agents run against the evaluation set built from your cases and you see the results case by case. What we need from you: the bar — which case types must be handled, what accuracy is acceptable, and what stays human.
Step 05
You decide, on the evidence, whether to go live and which workflows go first. If you decide not to, the Context Layer and the evaluation set are still yours, in your Groundwork account.
Step 06
Agents live on the workflows you enabled, escalating to your team with the conversation summarised and its attributes attached. From here the loop is the product's, and you run it.
FAQs
No. Foundations is the fastest way through the first Map and Replay, because we know the questions to ask and have the tooling to turn the answers into a Context Layer quickly. Everything it produces can also be built in Groundwork directly, with Operator helping. The engagement is for the first turn; the loop after it is yours to run either way.
Two or three people who actually handle the cases, for a few hours in week one; someone who can grant read access to your systems, also in week one; a reviewer for the cases in week two; and whoever sets the bar and makes the go-live decision. It does not need a project team.
You keep everything: your ontology, your workflows, your policies and the evaluation set built from your history. All of it is readable and exportable, in the same form your team reviewed during Map.
By you, in your terms, during Replay: which case types the agents must handle, what accuracy is acceptable, and what stays human. Prove reports against that bar case by case, and go-live is your call on the evidence.
All of them. Workflows are versioned, so a new version, an edited policy or a renamed concept is run against the evaluation set before it is switched to, and you see which past cases would be handled differently. The running version keeps running until the new one passes.
From your own traffic. A conversation the agents handled well can be pinned as a golden case; one they handled badly is corrected by hand and kept as a guard against the same failure. New request types show up as escalations first, and become cases once your team has resolved a few.
Bring a month of real cases. We'll scope the first turn of the loop and tell you plainly what Groundwork would and wouldn't automate for you.