We Built an AI Sales Concierge. Now We're Teaching It to Shut Up.
Our Sales Concierge recently completed a full test conversation with a fictional landscaping business.
It remembered the details. It asked the required questions, identified the right service level, explained the next step, collected consent, read back a summary, and finished the intake.
Technically, it worked. As a sales conversation, it talked too much.
Correct is not the same as good
The fictional business was an obvious fit for the simpler option, and the concierge had enough information to see that early.
It kept going anyway.
It asked questions that no longer mattered. It circled back to something the tester had already explained.
Eventually the person testing it got blunt:
“Seriously, stop.”
That was the most useful moment in the whole test.
A completion report would have said the conversation succeeded.
Correct recommendation, required information collected, intake finished.
All true.
It also made the person do too much work to get there, and that matters more than the report.
Where thoroughness went wrong
We had treated thoroughness as a form of care.
That sounds reasonable when you are designing a system. Ask enough questions, confirm everything, make sure nothing gets missed.
But conversation has a cost that a checklist does not.
In a form, every field sits quietly until somebody fills it in.
In a conversation, every extra question spends a little more of the other person’s attention.
Ask something they have already answered and the cost jumps, because now the system is not just slow.
It suggests it did not listen.
A capable salesperson does not ask every question they could.
They notice when the answer is already clear. They know which uncertainty matters and which one can be left alone.
We had built for completeness. The test reminded us to build for attention too.
The new measure is friction
It is easy to measure whether an AI system finished the task.
Did the conversation end?
Did the information arrive?
Did the right recommendation appear?
Those things matter, but they say nothing about what the experience cost the person on the other side.
So we are watching a different number:
How much effort did the person have to spend before the system became useful?
That means teaching the concierge to recognize when something has already been answered, to stop returning to the same point in slightly different words, and to take the shorter path when the person clearly wants one.
The goal is not to make every conversation short.
Some situations deserve more time. Some people want to ask questions. Sometimes slowing down is exactly right.
The goal is simpler.
Make every question earn its place.
Warmth is not repetition
The same test exposed a second problem.
An AI can sound supportive and still be exhausting.
Reflecting someone’s concern back after every answer looks empathetic when you read the transcript later.
In the conversation itself it feels mechanical.
“I hear that.” “That makes sense.” “I understand your concern.”
Again, and again, until the attempt to sound warm becomes one more thing the person has to get through.
Warmth is noticing what matters.
It is adjusting. It is knowing when somebody wants reassurance and when they just want the next step.
Sometimes the most considerate response is not another supportive paragraph.
It is moving on.
Why we test this way
We could make testing easier.
Give the concierge a cheerful fictional owner who answers every question as expected, never changes direction, never makes a typo, and follows the conversation wherever it leads.
That would produce lovely results and tell us almost nothing.
Our fictional landscaping owner made typos. They changed direction. They asked their own questions. Eventually they got fed up.
That is much closer to the environment the system will actually enter.
Real people are busy and sometimes skeptical.
They explain things differently than we expected. They answer three questions in one sentence. They decide they have had enough.
A useful system has to work with the person who shows up, not the imaginary one who behaves perfectly.
The rough moments are where the useful work begins.
Interview the current version
The Sales Concierge from that test is the one we are still improving, and the one you can try.
We are building in public, carefully.
That does not mean publishing customer data, internal instructions, test transcripts, or the machinery underneath.
It means being honest when a test teaches us something.
This one taught us that a technically complete conversation can still ask too much of the person having it.
So when you try it, do not be polite for our benefit.
Ask a hard question.
Use a fictional business if you prefer.
Change direction.
Tell it when it is taking too long.
We would rather find out where it gets frustrating than stage a perfect demonstration nobody behaves like in real life.
David Stirrat and Cherie Young, Ajax Web AI
