CASE STUDY · SALES TRAINING
AI Roleplay Training: somewhere safe to fumble.
A voice AI platform where sales reps run live roleplay calls against AI buyers that stall, object, and push back — then get scored automatically, with feedback that hears tone, not just words.
- Client
- Mary Wylde, MasteryTrainer
- Industry
- Sales training
- Live at
- masterytrainer.ai
- Built by
- Coder Crew
3
AI providers, one voice loop
8
Competency scoring categories
9
Emotions detected live
80+
API endpoints
Rookies train on live prospects.
Practice happens on real deals
The first hundred objections a rep ever handles belong to actual prospects, and the losses are invisible on any report.
Human roleplay doesn't behave like buyers
Colleagues are scripted, scheduled, and too polite. Real buyers stall, object, go quiet, and push back.
Feedback is one manager's gut
Coaching depends on whoever listened in, when they had time. Subjective, inconsistent, and blind to tone.
Nobody sees who's improving
Practice happens, or doesn't, invisibly. Managers can't see who's rehearsing or where the whole team stumbles.
Simulations feel simulated
A two-second lag kills the illusion of a live call, and with it the training value of the whole exercise.
Five problems, five design answers.
Before
Training on real deals
After
Unlimited reps, zero deals at risk. Live, spoken roleplay against AI buyers — always available, no scheduling, no prospect on the other end.
Before
Too-polite partners
After
AI buyers that behave like buyers. GPT-4 personas across buyer and seller scenarios and three difficulty levels — they stall, object, go quiet, and push back.
Before
Gut-feel coaching
After
Deterministic scoring that hears tone. An eight-category GPT-4 rubric run at temperature zero, enriched by nine-state emotion detection — the same call always scores the same.
Before
Invisible practice
After
A manager's view of the whole team. Analytics, scorecards, and a full LMS show who's practicing, who's improving, and where the team keeps stumbling.
Before
The lag problem
After
A streaming voice loop. Deepgram to GPT-4 to ElevenLabs over WebSockets, engineered so the conversation feels live — because that's the whole product.
A voice loop built to feel like a person.
The hard part isn't the AI — it's the silence. A lag kills the illusion of a live buyer, so the entire pipeline streams.
Rep speaks
Browser microphone, audio streamed over WebSockets
Deepgram transcribes
Real-time speech-to-text, shown as a live transcript
GPT-4 plays the buyer
In-persona objections and stalls, aware of the rep's detected emotional state
ElevenLabs answers aloud
The buyer's voice, streamed straight back into the call
Rep speaks
Browser microphone, audio streamed over WebSockets
Deepgram transcribes
Real-time speech-to-text, shown as a live transcript
GPT-4 plays the buyer
In-persona objections and stalls, aware of the rep's detected emotional state
ElevenLabs answers aloud
The buyer's voice, streamed straight back into the call
When the call ends, the same GPT-4 scores it across eight categories at temperature zero, so evaluations are repeatable. Emotion detection runs on the live transcript across nine states — surfacing coaching prompts to the rep mid-call and feeding the buyer's behavior.
A complete SaaS, not a demo.
19+ screens across the two roles, 11 database entities, 131+ source files.
Live voice roleplay
Spoken scenarios across three difficulty levels, live transcript on screen
GPT-4 evaluation engine
Eight-category rubric with detailed coaching feedback
Emotion detection
Nine classifications with confidence, running mid-call
Audio upload analysis
Real recorded field calls scored by the same engine
Analytics & scorecards
Radar charts, score progression, scenario-level insights
Learning management system
Course → lesson → task, with roleplay as a task type
Content management
Scenarios, videos, and documents in one searchable place
Manager dashboard
Who's practicing, who's improving, where the team stumbles
Dual-role architecture
Trainee and manager, with role-isolated routes and recording access controls
Three-provider pipeline
Deepgram, GPT-4, and ElevenLabs orchestrated in one loop
Service-layer backend
Async FastAPI, 80+ endpoints across the whole platform
In the product.



The expensive classroom is now optional.
Reps rehearse objections on AI buyers instead of live prospects
Feedback stopped being one manager's gut — it's an eight-category rubric that also hears tone
Real field calls feed the same engine through audio upload analysis
Managers finally see practice: who's rehearsing, who's improving, who's stumbling
The whole loop streams — speech, reasoning, and voice — so it feels live
How it's built.
| Frontend | React 18 · TypeScript · Vite · Tailwind CSS · Material UI |
| Backend | FastAPI (async, service-layer) · Python · SQLAlchemy |
| Voice pipeline | Deepgram STT → GPT-4 persona → ElevenLabs TTS, over WebSockets |
| AI evaluation | GPT-4: eight-category rubric at temperature zero · nine-state emotion detection |
| Database | SQLAlchemy ORM: SQLite (dev), PostgreSQL (prod), 11 entities |
| Security | JWT + bcrypt · role-isolated routes · recording access controls |
| Analytics | Chart.js radar and trend charts |
Why Coder Crew.
Real-time voice AI is an engineering problem before it's an AI problem — streaming, latency, and orchestration across three providers, delivered as one accountable build.
“Before this, new reps were learning on real calls, and we had no consistent way to see what they were struggling with. Now they can practice whenever they want, get useful feedback straight away, and managers can actually track where the team is improving.”
— Mary Wylde, MasteryTrainer
Have a product that needs shipping — or a team still training the expensive way?
Either way, we start the same place: understand first, then build. The audit is one week, one fixed price, credited toward the work.
