The problem
Replies to a webinar campaign were landing in Smartlead and being read, interpreted, matched to a campaign and routed by hand. The platform’s own categorisation was not good enough to drive the workflow off, and every reply that looked positive needed a person to decide what it actually meant.
The first live pull of twenty replies tells you what the job really is. Twelve out-of-office. Three bounces and departure notices. Four interested. One meeting request.
So the dominant task is not spotting subtle enthusiasm. It is filtering automated noise without throwing away the four replies that matter.
The system
Once a day, on a schedule. Pull unread replies for the active campaign. Classify each one with Claude. Write everything to Airtable. Register the genuine positives for the event. Mark the threads read so the next run does not see them again.
Nine categories, three buckets. The model picks one of nine: explicit yes, soft yes, information request, soft decline, hard no, out of office, wrong person or bounce, referral, unclear. Those collapse into Positive, Unsure and Negative.
Four categories route to Unsure. Only two reach Positive.
The confidence gate is applied last and unconditionally. Anything below 0.70, or anything the model flags for review, becomes Unsure regardless of which category it landed in. A confident soft-yes registers. The same soft-yes at 0.61 goes to a person.
Dry run is the default, permanently. --commit has to be typed. Sixty-five tests, none of which touch the network.
You can run the routing logic yourself.
The numbers
- 9 categories mapping to 3 buckets, four of them landing in Unsure
- 0.70 confidence floor, applied after the mapping, not inside it
- 65 tests, all offline, no network required to run the suite
- On a 64-reply sample it disagreed with the vendor’s own labels 11 times, including two the platform tagged Out Of Office that a human read as interested
- Roughly $0.60 in model spend per 64-reply run with extended thinking on, about $220 a year at that daily volume
The decisions
The push decision ignores the model’s own answer
The classifier returns a push_to_goldcast boolean. The code throws it away and re-derives the decision from the final bucket instead.
That looks redundant. It is the opposite. It means there is exactly one code path to registration, and no way for a model quirk, a prompt edit or a schema change to route around the confidence floor. The gate cannot be bypassed because nothing else is allowed to decide.
“Unsure” goes to a person, not to “no”
The easy build forces a binary. I refused it.
An ambiguous reply is not a rejection, and treating it as one throws away the exact replies most worth a human read. referral sits in Unsure for the same reason: a human handoff deserves a human, not a suppression rule.
At the time this felt like admitting the classifier was incomplete. It is the decision I would defend hardest now.
The idempotency check runs before the model call, not after
The service asks Airtable “have I already seen this message?” before it calls Claude.
Put that check after the call and a re-run costs zero rows but full price in tokens. Put it before and a re-run costs nothing at all. That single ordering choice is what makes the whole thing safe to re-run after any failure, which in turn means a crash needs no cleanup and a failed mark-as-read needs no retry.
Three phases, because marking a reply read is destructive
The obvious structure marks each reply read inside the loop that processes it. Two failures fall out of that.
The unread list is filtered and sorted server-side, so marking items read while paginating removes them from the result set, every later offset shifts, and replies get silently skipped. And if the process dies halfway, work already marked read may never have been written down.
So: phase one fetches everything and mutates nothing. Phase two classifies and writes, one reply at a time, with errors contained so a single bad reply cannot stop the batch. Phase three marks read, only once Airtable is durable.
A routing fault I caused, found and fixed
The first architecture resolved positive replies against a single event. That is correct with one campaign running and wrong the moment there are two.
I found it in testing. A 13 August event ID was being applied to a 19 August campaign, and seventeen test contacts registered against the wrong event. No real contact was affected.
I traced it, corrected the records, and rebuilt the workflow to resolve the campaign to its event before any classification runs.
The failure was not in the AI layer. Every one of those replies was classified correctly. The failure was in the routing underneath it, which was the least interesting part of the system and the part I had spent the least time on.
Handover
Built, tested and documented. Handed over mid-productionisation at the end of the engagement, with a launch runbook and a separate document covering the human-review loop.