Many revenue teams have moved past a single large-prompt LLM wrapper that researches, qualifies and closes in one call. They now split the work across specialized agents built with frameworks such as CrewAI, LangGraph and Microsoft’s Agent Framework. That gain in specialization brings two engineering problems: state drift, where critical facts get lost or distorted between agent handoffs, and context fragmentation, where each tool and system sees only part of the deal.
“State drift” is this article’s shorthand, not a standard term. The underlying failures are well documented. A study of multi-agent LLM systems found that they often add little over single-agent setups. It also cataloged where they break.
This article covers three things:
1. Where multi-agent revenue pipelines fail.
2. How to place human-in-the-loop (HITL) checkpoints at the points where failures cost the most.
3. How to route tool calls from the Model Context Protocol (MCP) into a Slack queue where an Account Executive (AE) can claim the deal.
Control planes such as InstaChime package this pattern. The code below is illustrative, and every threshold, endpoint and name in it is a placeholder.
Part 1: Why Multi-Agent Handoffs Fail
The Anatomy of a Multi-Agent Revenue Pipeline
A typical pipeline splits work into specialized roles:
- Research agent: gathers firmographics, technographics, intent signals and recent news about the account.
- Qualification agent: scores the lead against a framework such as BANT or MEDDPICC.
- Objection-handling agent: answers technical, security and pricing questions in the conversation.
```
+----------------+ +---------------------+ +--------------------+
| Research Agent | --> | Qualification Agent | --> | Objection Agent |
| (enrichment) | | (BANT / MEDDPICC) | | (pricing / tech) |
+----------------+ +---------------------+ +--------------------+
|
risk threshold or high-intent signal
v
+----------------------+
| HITL control layer |
+----------------------+
|
v
+----------------------+
| AE Slack queue |
+----------------------+
```
What the Research Says About Failure
The paper “Why Do Multi-Agent LLM Systems Fail?” (Cemri et al., NeurIPS 2025) is the best empirical grounding for this problem. The authors annotated more than 1,600 execution traces from seven multi-agent frameworks. They defined a taxonomy, MAST, of 14 failure modes in three categories:
- specification and system-design issues
- inter-agent misalignment
- task-verification failures
One of those failure modes is agents failing to ask for clarification when they should. Several others involve information getting lost or ignored between agents. Each maps onto a revenue scenario.
#### 1. Context loss at handoffs
When Agent A passes work to Agent B, long histories are usually summarized or truncated. Details that matter commercially can disappear in that step: an exact budget ceiling, a data-residency requirement, a negotiated contract term. The next agent then works from an incomplete picture.
#### 2. Errors that propagate
If the qualification agent misreads a buyer’s timeline, the objection-handling agent inherits that mistake. Nothing in a plain sequential pipeline forces anyone to verify the earlier agent’s output. Task verification is one of the three MAST categories for this reason.
#### 3. Loops and latency
Conversational and delegating frameworks can bounce a question between agents. A buyer who asks about an enterprise SLA or HIPAA compliance may wait through several internal exchanges before getting an answer. Cap turns and add a timeout path that ends in a human.
```
one agent's output before the next agent relies on it.
from crewai import Agent, Crew, Process, Task
researcher = Agent(
role="Lead Researcher",
goal="Gather account intelligence and key buyer pain points",
backstory="Senior B2B researcher.",
)
qualifier = Agent(
role="MEDDPICC Qualifier",
goal="Score the opportunity and summarize buying intent",
backstory="Enterprise sales qualifier.",
)
research_task = Task(
description="Research {account} and summarize pains, stack and recent news.",
expected_output="A one-page account brief.",
agent=researcher,
)
qualification_task = Task(
description="Qualify the lead using the research brief.",
expected_output="A MEDDPICC summary with a score and open questions.",
agent=qualifier,
)
crew = Crew(
agents=[researcher, qualifier],
tasks=[research_task, qualification_task],
process=Process.sequential,
)
```
Where Human Checkpoints Belong
A human review at every step cancels the speed advantage of automation. A review at no step lets errors reach buyers. The practical middle is to pause on specific conditions:
- High-value intent: large seat counts, security or compliance documentation requests, custom-pricing requests.
- Loop detection: more than N turns, or the same two agents trading messages repeatedly.
- Verification failures: a validator step flags a contradiction with CRM data.
- Irreversible actions: sending a contract, committing to a price or SLA, scheduling an executive meeting.
```
[ Agent pipeline ]
|
( policy evaluation )
|
+------------+-------------+
| |
[ within bounds ] [ rule triggered ]
| |
continue agents pause + snapshot state
route card to AE in Slack
```
When a rule fires, the system should do four things:
1. Pause the workflow.
2. Persist a snapshot of the state: conversation, tool-call log, scoring metadata and the proposed next action.
3. Present that snapshot in a place the AE already works.
4. Wait for an explicit decision: take over, approve the agent’s plan, or correct the context and reroute.
What Frameworks Already Provide
You do not have to build the pause-and-resume mechanics from scratch.
- CrewAI supports HITL in two ways. Flows have an `@human_feedback` decorator, available from version 1.8.0. Crews have a webhook-based approach in which a task pauses in a “Pending Human Input” state until feedback is submitted. CrewAI Enterprise adds routing, escalation policies and SLA settings for Flow HITL, including automatic fallback responses when no human replies in time.
- LangGraph provides `interrupt()` with a checkpointer, so a graph can stop at a node and resume with a human-supplied value.
- AutoGen is now in maintenance mode. Microsoft moved it there in October 2025, and its README directs new users to the Microsoft Agent Framework, which reached 1.0 general availability in April 2026. A community fork, AG2, continues separately. Use Agent Framework for new builds and plan a migration if you run AutoGen today.
A CrewAI Flow with a human review step:
```
from crewai.flow.flow import Flow, start, listen
from crewai.flow.human_feedback import human_feedback
class QualificationFlow(Flow):
@start()
def qualify(self):
Call your qualification crew here and return its summary.
return "Account: Acme Corp | Seats: 450 | Score: 0.94 | Ask: custom VPC"
@listen(qualify)
@human_feedback(message="Review this qualification before outreach continues:")
def review(self, summary):
return summary
```
And a LangGraph pause that carries a state snapshot to a reviewer:
```
from langgraph.types import interrupt, Command
def ae_review(state: dict) -> dict:
Pauses the graph (requires a checkpointer) and surfaces the payload.
decision = interrupt({
"account": state["account"],
"proposed_action": state["proposed_action"],
"transcript": state["messages"][-10:],
})
Resumed later with Command(resume={"action": "approve" | "takeover" | "edit", ...})
return {"ae_decision": decision}
```
Part 2: Routing MCP Tool Calls to an AE Queue
MCP in a Revenue Stack
The Model Context Protocol is an open standard for connecting AI models to tools and data. Anthropic created it, and in December 2025 donated it to the Linux Foundation, where it is a founding project of the Agentic AI Foundation alongside Block’s goose and OpenAI’s AGENTS.md. In a revenue stack, an agent with an MCP client can discover and call tools such as CRM lookups, enrichment, billing queries and calendar availability.
```
+-------------------------------------------------------------+
| MCP host: agent / LLM + MCP client |
+----------------------------+--------------------------------+
| MCP (JSON-RPC)
v
+-------------------------------------------------------------+
| MCP servers |
| CRM | enrichment | billing | calendar | quoting |
+-------------------------------------------------------------+
```
Most tool calls are routine. A few signal real buying intent: a request for an enterprise quote, a security questionnaire, or a custom proposal. Those are the calls worth routing to a person.
What the MCP Specification Says About Humans
The MCP tools specification recommends that applications keep a human in the loop who can deny tool invocations. It suggests visual indicators when tools run and confirmation prompts for operations. This is a recommendation, not a mandate. The protocol deliberately leaves the approval interface and workflow to implementers.
The most recent specification revision (2026-07-28) also lets a server respond to `tools/call` with an `InputRequiredResult`, which signals that more input is needed before the call can complete. That mechanism suits human approval and gives you a protocol-level place for the pattern described here. Check the current spec for exact semantics before building on it.
Design: Intercept at a Tool Boundary
Put the policy check where tool calls pass through a single point: an MCP server wrapping a sensitive tool, or a gateway in front of several servers.
```
Buyer Agent / MCP client Gateway or server Slack (AE queue)
| | | |
|-- "500 seats?" --->| | |
| |-- tools/call --------->| |
| | request_quote |-- evaluate policy |
| | |-- post claim card ---->|
| | |<-- AE clicks "Claim" --|
| |<-- result: handoff ----| |
|<============== AE continues the conversation ========================|
```
#### Step 1: Define the policy
The rules should be simple, explicit and auditable. These thresholds are examples:
```
HIGH_INTENT_TOOLS = {"request_enterprise_quote", "request_security_docs", "book_executive_demo"}
SEAT_THRESHOLD = 100 # illustrative
INTENT_THRESHOLD = 0.85 # illustrative; produced by your own scoring model
def needs_human(tool: str, args: dict) -> bool:
return (
tool in HIGH_INTENT_TOOLS
and (args.get("seat_count", 0) > SEAT_THRESHOLD
or args.get("intent_score", 0) > INTENT_THRESHOLD)
)
```
#### Step 2: Post an actionable Slack card
Slack’s Block Kit supports interactive buttons in messages. Incoming webhooks can post Block Kit layouts, but they cannot receive button clicks. To handle clicks, create a Slack app with interactivity enabled and either a public Request URL or Socket Mode. Slack expects your app to acknowledge each interaction within 3 seconds. Do the slow work, such as resuming the workflow, asynchronously.
```
from slack_sdk import WebClient
slack = WebClient(token="xoxb-...") # load from a secret manager
def post_claim_card(channel: str, session_id: str, summary: str):
slack.chat_postMessage(
channel=channel,
text="High-intent request needs an AE", # notification fallback
blocks=[
{"type": "section",
"text": {"type": "mrkdwn",
"text": f"*High-intent request detected*\n{summary}"}},
{"type": "actions", "elements": [
{"type": "button", "action_id": "claim_deal", "style": "primary",
"text": {"type": "plain_text", "text": "Claim deal"},
"value": session_id},
{"type": "button", "action_id": "allow_agent",
"text": {"type": "plain_text", "text": "Let the agent continue"},
"value": session_id},
]},
],
)
```
```
+-----------------------------------------------------------+
| High-intent request detected |
| Account: Acme Corp | 450 seats | Enterprise tier |
| Ask: dedicated VPC, 24/7 phone SLA |
| Agent state: paused, waiting for a decision |
+-----------------------------------------------------------+
| [ Claim deal ] [ Let the agent continue ] [ View thread ]|
+-----------------------------------------------------------+
```
#### Step 3: Handle the click and resume
```
from slack_bolt import App
app = App(token="xoxb-...", signing_secret="...")
@app.action("claim_deal")
def on_claim(ack, body, client):
ack() # must happen within 3 seconds
session_id = body["actions"][0]["value"]
ae_id = body["user"]["id"]
Record the claim, notify the pending tool call, and hand the thread to the AE.
resolve_pending(session_id, {"status": "claimed", "ae": ae_id})
```
#### Step 4: Gate the tool itself
This sketch uses the official Python MCP SDK’s `FastMCP` helper. The server pauses the tool call until an AE decides or a timeout expires.
```
import asyncio
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("revenue-tools")
PENDING: dict[str, asyncio.Future] = {} # use a durable store in production
def resolve_pending(session_id: str, decision: dict):
fut = PENDING.get(session_id)
if fut and not fut.done():
fut.set_result(decision)
@mcp.tool()
async def request_enterprise_quote(session_id: str, account: str,
seat_count: int, intent_score: float = 0.0) -> str:
"""Generate an enterprise quote, routing high-intent requests to an AE first."""
args = {"seat_count": seat_count, "intent_score": intent_score}
if needs_human("request_enterprise_quote", args):
fut = asyncio.get_running_loop().create_future()
PENDING[session_id] = fut
post_claim_card("#ae-queue", session_id,
f"{account}: {seat_count} seats, intent {intent_score:.2f}")
try:
decision = await asyncio.wait_for(fut, timeout=180) # SLA window
except asyncio.TimeoutError:
return "An account executive has been notified and will follow up shortly."
finally:
PENDING.pop(session_id, None)
if decision["status"] == "claimed":
return "Your request has been passed to a senior team member who will join shortly."
return f"Draft quote prepared for {account} ({seat_count} seats)."
```
Treat this as a starting point. A production version needs these additions:
- Slack request-signature verification, which Bolt handles for you.
- Idempotent claims, so two AEs cannot claim the same deal.
- Persistent state, so a restart does not lose pending approvals.
- An audit log of every decision.
- Authentication and per-tool authorization on the MCP side.
Comparing Approaches
| Concern | Unmanaged multi-agent pipeline | Single LLM with tool calling | Pipeline with HITL control layer |
|---|---|---|---|
| Handoff errors | Can compound across agents; no built-in check | Fewer handoffs, but limited scope | Policy rules and review points catch selected cases |
| Escalation | Usually discovered afterward in logs | Often a static fallback | Explicit, rule-based, routed to a named queue |
| Tool interface | Framework-specific functions | Provider-specific function schemas | MCP tools, usable across clients |
| Human role | Post-hoc debugging | Hard-coded approvals | Approve, take over, or correct, with state attached |
| Main cost | Silent failures | Limited capability | Added latency and operational overhead |
The last row matters. A human checkpoint adds delay and requires staffing. It pays off only when the rules are narrow enough that AEs see few, valuable cards.
Measuring Whether It Works
Avoid adopting benchmark claims from vendors, including this one. Measure your own pipeline against a baseline:
- Time to claim: median and 90th-percentile time from card posted to AE claim.
- Claim rate and timeout rate: a high timeout rate means the queue is under-staffed or the rules fire too often.
- Override rate: how often AEs reject or correct the agent’s proposed action. This is your best signal of agent quality.
- Loop and timeout incidents: conversations ending in a buyer drop-off after exceeding turn limits.
- Conversion and cycle time for escalated versus non-escalated opportunities, with a control group where possible.
Practical Guidance
- Start narrow. Gate one or two irreversible or high-value actions, then widen based on override data.
- Keep policies deterministic. Rules like “seat count over N” or “more than N turns” are testable. Model-confidence signals are less reliable, so treat them as supplements.
- Snapshot what the human needs. Include the transcript, tool-call log, CRM facts and the agent’s proposed next step, so the AE doesn’t rebuild context.
- Design for timeouts. Define what the buyer sees if no AE responds, and make that path graceful.
- Secure the boundary. Verify Slack signatures, scope tool permissions, and treat tool inputs and outputs as untrusted, since prompt injection is a known risk for agents that call tools.
- Plan for framework churn. AutoGen’s move to maintenance mode shows that framework choices age quickly. Keep your policy and handoff layer independent of any one agent framework.
Conclusion
Multi-agent pipelines can speed up research and qualification, but research on multi-agent systems shows that failures cluster around specification, coordination between agents, and missing verification. A narrow, rule-based human checkpoint at high-value or irreversible moments addresses those failures directly. MCP gives tool calls a common interface that such a checkpoint can sit in front of. Slack gives AEs a place where they already work. The aim is a pipeline where agents handle volume and people handle the decisions that decide deals.
Sources
- Cemri et al., “Why Do Multi-Agent LLM Systems Fail?”, NeurIPS 2025
- Model Context Protocol specification: Tools
- CrewAI documentation: Human-in-the-Loop and Human Feedback in Flows
- microsoft/autogen on GitHub (maintenance-mode notice)
- IT Brief, “Anthropic donates MCP to new Agentic AI Foundation”
- Slack API documentation: Handling user interaction
