The Problem
I have two virtual assistants who handle the front line of my cleaning business. They answer inbound calls, follow up on missed calls, qualify leads, send quotes, book jobs, and handle customer issues. They generate most of my revenue.
I had no idea what they actually said on the phone. I could see how many calls they took, how fast they responded, and how many bookings they created. But I couldn't tell whether they were quoting prices over text instead of calling back, whether they were missing upsell opportunities, whether they sounded warm or robotic, or whether a customer who didn't book had a fixable objection.
Listening to even 10% of the calls myself would have eaten my entire week. I needed something that could listen to all of them and tell me what mattered.
What I Built
I built a nightly pipeline that pulls every call transcript from GoHighLevel, sends each one to Claude Opus 4.6 along with the related SMS thread, and gets back a structured score on six dimensions plus specific coaching notes that reference real moments from the conversation.
At 11 PM Eastern, the system fetches every transcript GHL recorded that day, batches them by VA and customer thread, runs the analysis, and stores the results in SQLite. A daily Telegram message tells me how each VA scored. A weekly synthesis runs every Sunday and produces a coaching report ranking the top three things each VA should work on.
The key insight is that the AI references specific moments in the conversation. Instead of "your follow-through could be better," it says "on the call with Carla at 2:14 PM you quoted prices via text instead of calling her back. She booked with a competitor 40 minutes later."
The 6 Scoring Dimensions
System Architecture
The two-step sync architecture exists because background threads silently failed in production. Lesson learned the hard way.
A Real Example
Here's the kind of coaching note Opus produces. This is paraphrased from a real one:
The Numbers
Cost Comparison
The traditional way to coach a sales team is to have a senior person listen to call recordings and write notes. That doesn't scale and it doesn't happen.
| Sales Coach | Manual Review | This System | |
|---|---|---|---|
| Monthly cost | $3,000-8,000 | 20-30 hours of your time | ~$30 in API |
| Coverage | 5-10% of calls | 5-10% of calls | 100% of calls |
| Specific moment references | Sometimes | Sometimes | Always |
| Trends over time | Manual | Manual | Automatic |
| Time to first report | 1-2 weeks | Weekly | Daily |
Tech Stack
What It Took
The hardest bug took eight hours to debug. Background threads in Python were silently failing on Railway when fetching transcripts from the GHL API. Same token, same URL, same messageId. The synchronous version worked. The threaded version returned empty. I never found the root cause.
The fix was to stop fighting it and refactor the architecture: do all the transcript fetching synchronously in the main thread first, then run the analysis only after every fetch completes. Slower, but reliable.
The other big lesson was max_tokens. On busy days the response was getting truncated mid-JSON because the default limit was too low. Bumping it to 16,384 fixed the silent failures on high-volume days.
Want this for your team?
Works for VAs, SDRs, customer success reps, support agents, anyone who talks to customers. If you record calls, this can listen to them.
Book a Strategy CallFree, 30 minutes. No pitch deck.