Someone in the room is going to say it. Usually about twenty minutes in, usually the most technical person there, and usually not unkindly: we already have Claude. Why can't we just build this ourselves?

The honest answer is that you can build a lot of it, quickly, and I would rather say that out loud than pretend otherwise. Give a competent engineer an afternoon and an API key and they will hand you a chatbot that plays a skeptical buyer. It will be genuinely impressive. It will hold character, push back, and improvise objections you did not script. Everyone who sees it will be a little bit delighted.

That demo is real. It is also about four percent of the work, and the ninety-six percent is invisible from where you are standing when you see it.

The demo is not the hard part

Roleplay is the thing large language models are naturally best at. You are asking a system trained on human dialogue to produce human dialogue. It is playing to its strengths, and it needs almost no scaffolding to be convincing for ten minutes.

So if you build the demo, you will conclude the problem is easy. That conclusion is the trap, because the demo answers a question nobody in your business actually asked. Nobody needs a convincing roleplay. They need to know whether a rep got better, and by how much, and at what specifically, and whether the number they are looking at means the same thing this month as it did last month.

That is a measurement problem, not a conversation problem. And measurement is where LLMs are at their worst.

Ask it to score, and the ground moves

Here is the thing you find in week three, not week one. Take a single transcript and ask the model to score it against a rubric. Then ask it again. You will get a different number.

Not wildly different. Just enough. A 6 becomes a 7. A framework that read as a weakness now reads as adequate. Nothing about the conversation changed, only the run.

This is fatal in a way that takes a while to become obvious, because the whole value of a score is comparison. You are asking whether this rep improved since last quarter, whether this branch is behind that one, whether the coaching is working. Every one of those questions is a subtraction between two numbers. If the numbers wobble on their own, the subtraction is noise, and you have built a very expensive random number generator with a nice interface.

The failure is quiet, which is what makes it dangerous. It does not error. It does not crash. It produces a plausible number every time, and it will be months before someone notices that the trend line never meant anything.

What it actually takes to hold a score still

Making a score stable is not a prompt engineering problem you solve once. It is a rubric with defined bands, tie-break rules for the ambiguous middle, explicit handling for the cases that break the rubric, and, above all, a regression suite that runs every time you touch any of it.

Parlare's is 880 tests as of August 2026, currently passing at zero failures. It runs against every change to the scoring engine. When a scenario is added, when a framework rule is tuned, when a model version moves underneath us, that suite is what tells us whether a score means today what it meant last week.

I want to be precise about why that number matters, because it is easy to read as vanity. It is not a measure of how clever the product is. It is a measure of how many ways we have already found for the scoring to drift, each one caught once and then pinned down so it cannot come back. Every test in there exists because something moved that should not have.

Nobody builds that for an internal demo. Not because they are not capable, but because the demo does not need it and the need only becomes visible after you have shipped to real users and started making decisions with the output. By then you have a score people trust in a system that does not deserve it yet.

Then the rest of it arrives

Suppose you solve consistency. You have a stable score. Now the questions start, and they are the ordinary questions any real user asks in the first fortnight.

What should I practice next? Parlare looks at the rep's lowest-scoring framework across their history, requires at least three scored sessions before it will act on the signal, ignores anything already above 7.5, and excludes the last three things they practiced so they are not handed the same drill twice. New users with no history fall back to the weakness they named at onboarding, and that seed decays automatically once real measurement exists. That is one feature. It is a page of rules, all of them arguable, each one the result of a decision someone had to make and defend.

Is this getting harder as I get better? The prospect's difficulty scales with the rep's measured proficiency on the framework that scenario trains, clamped so that one strong session does not immediately unlock the hardest counterpart. A practice tool where the difficulty never moves stops being practice after a fortnight.

What is my team actually doing? Who practiced, what they practiced, how they scored, whether they are trending up. Assignment, so a leader can put a scenario in front of a specific person. Completion, so they can see it happened. This is the part that decides whether the tool gets used at all, and it has nothing to do with AI.

Where are the scenarios? Parlare ships twenty-one for financial services alone, each tagged to a line of business, each written so a banker recognizes the situation rather than a training vendor's idea of one. Mortgage renewal with rate shock and a competitor offer on the table. Intergenerational transfer-out risk when the adult child moves the account. Renegotiating a facility the company has outgrown. Writing those is not engineering work. It is domain work, and it is the part that takes years rather than weeks.

Month four

The internal build usually does not fail. It just stops.

The model provider deprecates the version everything was tuned against, and the scores shift underneath a system with no regression suite to catch it. A new product launches and needs scenarios, and the person who knows how to write them is on something else now. Someone asks for the manager view and it is a quarter of work nobody scoped. The engineer who built it moves teams, and what they left behind is a prompt file with no tests and no one who fully understands why any particular line is in it.

None of this is a failure of competence. It is what happens to internal tools that were scoped as a demo and then quietly asked to behave like a product. The build was never the expensive part. The owning is.

When you should build it

There is a version of this where building is the right call, and I would rather name it than pretend there isn't.

If conversation quality is genuinely core to how your business competes, if you have a team who can own it for years rather than an engineer with some spare cycles, and if you want the rubric to encode something proprietary that no vendor could sell you, then build it. You will end up with something better than anything you could buy, because it will be yours.

What does not work is the middle path: building it because the demo looked easy, then discovering the maintenance bill after everyone has moved on to the next thing. That is the version that costs the most and delivers the least, and it is by far the most common.

So the question worth asking in the room is not can we build this. You can. It is: who owns the scoring rubric in eighteen months, and what happens to every number in the system on the day they leave?

If that question has a confident answer, build it. If it does not, you were never really choosing between building and buying. You were choosing between buying and starting.

Sample · Scored practice session
5.4 /10
Sales · "Mortgage renewal — rate shock and a competitor offer"
Practice sim  ·  AI client  ·  Financial Services / Retail Banking
Hook Accuracy
7/10  ·  client named the cost of switching in their own words
Curiosity Quotient
4/10  ·  asked broadly, never followed up on the answer
MAP Specificity
3/10  ·  closed on "let me send you some numbers"
Conversation Elevator
Floor 3  ·  advised before the client finished surfacing the concern
Where to sharpen
You got the Hook — the client said out loud what leaving would cost them, which is the hard part and you earned it. Then you spent it. The next four turns were you talking, and the close was "let me send you some numbers." A specific next step with a date attached converts that admission into a renewal. Run it again and hold the silence after the Hook lands.

That is the output an internal build has to reproduce, and reproduce identically the next time the same transcript goes through. Not the roleplay above it, which is the easy half. The three numbers, the named behaviour behind each one, the specific corrective move, and the guarantee that a 3 means the same thing today as it did last quarter. When we changed the rubric in August 2026 we re-baselined deliberately and dated it, because a scoring change you cannot point to is indistinguishable from drift.

"Practice, not training" is not a slogan. It is the difference between a team that can explain what good looks like and a team that performs it when the deal or the relationship is on the line. The frameworks are not the hard part. Everyone has the frameworks. The hard part is making them automatic, and the only thing that has ever done that is deliberate practice with feedback specific enough to act on. That is the whole product.