
Six weeks to a deployed multi-agent system
A LangGraph-based multi-agent system for consulting sales experts, built solo in 6 weeks. Time-to-draft fell from an expert-reported 2+ hours to about 2 minutes in evaluated runs, and the outcomes were presented to senior stakeholders.
The problem
Interviews with 5 sales experts mapped the prospecting workflow end to end and revealed two distinct halves. Identifying the right client is real expertise. It relies on the expert's network and on factors no system can predict. Everything after that, from spotting an intent signal to packaging the message, mostly comes down to applying rules consistently. That half ran two to ten hours per prospect, and it is where expertise gives way to information gathering.
The mandate was open: build something with AI that helps. Defining that something was up to me. I had to decide which parts of the workflow to automate and which to leave to the expert, then prove it worked. I owned the product end to end and built it in 6 weeks.
Where is the line between what agents should automate and what has to stay personal for the outreach to keep working?
Constraints
| Timeline | 6 weeks from product definition to deployment, as a solo builder. |
|---|---|
| Users | Consulting sales experts with strong personal styles and no tolerance for generic AI output. |
| Technical / domain | Drafts had to be grounded in real account context; a wrong claim in outreach costs credibility with the prospect. |
| Business need | Results had to be presentable to senior stakeholders as evidence the approach works. |

Key decisions
Solo and on a six-week clock, every architectural choice was also a scoping choice. Five decisions determined what the six weeks produced.
A single prompt could produce a draft, but it couldn't research an account and write in an expert's voice with any reliability. Those are different jobs with different failure modes. Splitting the work into specialized agents orchestrated with LangGraph made each step testable and debuggable on its own, which is what made six weeks feasible.
Full automation produces outreach that reads like automation, and experts won't send what doesn't sound like them. The system automates research and structure; the expert's voice and judgment stay in the loop. That balance was the core product decision, and it's what separated this from a template generator.
Review had to happen where a mistake would be expensive, before anything reaches a prospect. The expert reviews and edits at the draft stage, where two minutes of their attention catches what agents miss, instead of auditing every intermediate step or rubber-stamping a finished send.
A 2-minute draft is worthless if it's a bad draft. The system was evaluated on both deterministic and non-deterministic criteria. Citation retrieval accuracy and hallucinated metric references were measured directly, and the final agent's output was judged against an LLM-as-Judge setup built with DeepEval.
Multi-agent systems fail opaquely. When a draft comes out wrong, you need to know which agent went wrong and what the run cost. I built a live console that streams each agent's logs in real time, plus a per-rep cost report generated on every run. Observability was a feature, not a debug tool. An expert reviewing a draft can see exactly which agent did what, and a stakeholder can see what each run costs.
How I measured success
The headline metric was time-to-draft: an expert-reported 2+ hour baseline against timed evaluation runs, counting only drafts that passed every quality gate. Speed without reliability saves nothing, so quality ran through its own evaluation layer.
The evaluation suite
The suite covers six dimensions: trigger precision, retrieval quality, citation match, voice, compliance, and the judge itself. The deterministic dimensions are measured directly. Every citation must resolve to a real source, and no draft may cite a metric missing from the retrieved data. Dimensions that resist hard rules, like voice and compliance, are scored by the LLM-as-Judge built with DeepEval, which reasons before it scores.
The judge gets audited too. A quality gate nobody checks becomes the weakest link in the pipeline, so judge accuracy is its own eval dimension rather than an assumption. That's what lets the headline number carry weight. A draft that ships has passed the deterministic checks and cleared a judge with a measured track record.

The solution
The Sales Prospecting Intelligence System (SPIS) is a LangGraph-based multi-agent pipeline that runs the rule-based half of prospecting end to end. A scheduled monitor scans the watchlist and live news. A trigger evaluator acts as the gate, rejecting stale or irrelevant events before they cost anyone attention. When a trigger passes, a draft compiler runs three sub-agents in sequence. The first frames the message for an executive reader. The second scrubs risky phrasing. The third matches the draft to the expert's own writing. Every draft is grounded through RAG over the consultancy's case studies, not the model's memory.
Quality gets enforced before a human ever sees the result. An LLM judge scores each draft with chain-of-thought reasoning and sends failing drafts back to the exact step at fault. A draft gets up to five passes and ships only at 8/10 or above. Gemini 2.5 Flash runs the pipeline agents, and Pro runs the judge. The expert stays in the loop at both ends, defining the watchlist and approving the send. Research and assembly are automated; voice and judgment are not.
What it refuses to do matters just as much. It is not fully autonomous, because that creates trust and compliance risks in consulting sales. It is not a volume accelerator either, because this business runs on relationships rather than reach. It shipped as a working deployment inside the firm and was evaluated end to end on synthetic account data. Live outreach was out of scope for the engagement. The outcomes were presented to senior stakeholders as evidence for how the consultancy could apply agentic AI to its own sales practice.

Outcomes
The ~2 minute figure counts only drafts that passed every evaluation gate.