
THE AGENTIC SALES BRIEF / SYSTEM 05

The first useful AI-outbound experiment should not test how many messages a system can send.
It should test whether one bounded AI-assisted step improves outreach for an approved audience without increasing customer, compliance, or operational risk.
This field experiment uses 50 approved accounts, one control group, one treatment group, one changed variable, and a human final-send decision. It measures message quality, qualified conversations, reviewer effort, and risk together.
Fifty accounts is a practical learning batch. It is not statistical proof. The result can justify another test, expose a weak control, or tell you to stop. It cannot establish a universal benchmark for AI personalisation.
SYSTEM DEFINITION — ONE CHANGE, ONE DECISION
Keep the audience, offer, sender, channel, call to action, follow-up policy, and observation window stable. Change only the preparation step you want to evaluate.
The signal to isolate
Choose one narrow question:
Does a source-linked account hypothesis help a rep produce a more relevant first message than the team's current approved preparation process?
This is testable. “Does AI improve outbound?” is not.
The treatment is also intentionally modest. It does not ask an agent to find contacts, write a full sequence, select an offer, and send messages. It adds one research output before the normal human-controlled drafting and sending process.
If your inputs are not yet reliable enough to produce a cited account hypothesis, build the account research agent that cites every claim before running this test.
The system on one page

One changed variable, one exact approval and one denominator-based decision.
The flow is deliberately simple:
50 approved accounts
↓
documented assignment
↓
control: current process treatment: one source-linked hypothesis
↓ ↓
same human review and exact final-message approval
↓
quality + reviewer effort + qualified conversations + risk
↓
expand / revise / stop
Every arrow needs an owner, a record, and a failure path. The experiment is not ready if the team cannot explain what happens when an entity match is uncertain, a source contradicts the proposed claim, or the sending tool returns an ambiguous status.
Assign the owners
One sales leader or RevOps operator owns the experiment. The account owner or assigned rep owns each final message.
The experiment owner is responsible for:
audience and exclusion rules
lawful and policy-compliant operating conditions
random or otherwise documented group assignment
reviewer instructions
logs, measures, and the observation window
pause, revise, expand, and shutdown decisions.
The rep is responsible for the truth and appropriateness of the message they send. The system cannot inherit that responsibility.
Write down who can pause the test. “Anyone can raise a concern” is not the same as naming the person who has the authority to stop pending sends.
Approve the inputs and tools
Use the minimum set needed for the experiment:
approved CRM account and contact fields
the current suppression and opt-out register
the account's official public sources
an approved offer and claims reference
a review workspace that preserves source links and edits
the normal approved sending channel, controlled by the rep.
The research component needs read access. It does not need permission to change opportunities, create contacts, or send email. The review layer may prepare one exact message version. The sending connection should accept only the version the rep approved and retain a message ID that prevents an uncertain retry from becoming a duplicate.
If enrichment, browsing, or model tools process data outside the team's existing environment, assess their access, retention, training, and contractual terms before the pilot. Do not add a tool simply because it can produce more personalisation fields.
1. Build the approved cohort
Select 50 accounts that meet the same written criteria. Use company-level criteria relevant to the offer, such as business model, region, team structure, or a verifiable operational signal.
Before assignment:
remove current customers, open opportunities, and active conversations unless the test explicitly covers them
apply suppression, opt-out, and do-not-contact records
exclude ambiguous entities and accounts with inadequate evidence
confirm that the contact and channel are permitted under applicable laws, contracts, and platform rules
minimise the personal data used.
The US Federal Trade Commission states that CAN-SPAM applies to commercial email, including business-to-business email. UK Information Commissioner's Office guidance distinguishes rules and data-protection considerations by subscriber and processing context. GDPR principles include purpose limitation, data minimisation, and accuracy, and Article 21 covers objections to direct marketing.
These sources support controls, not case-specific clearance. Markets and channels differ. “B2B” is not a universal exemption. Verify the requirements that apply before any real send.
Treat audience approval as an input, not an outcome. If eligibility is uncertain, exclude the account before assignment. Do not let the treatment group carry the burden of questionable data.
2. Split control and treatment
Assign 25 accounts to the control group and 25 to the treatment group.
Where practical, balance obvious factors such as segment, company size, or territory before random assignment. Record the method. Do not move an account between groups after seeing its draft or outcome.
Control
The rep follows the current approved preparation process.
Treatment
The rep receives one AI-assisted, source-linked account hypothesis before drafting the message.
Keep these stable across both groups:
audience criteria
offer and approved claims
sender profile
channel
call to action
follow-up policy
observation window
outcome-coding rules.
The experiment tests the research aid, not an entire new sales motion. If the treatment has different accounts, copy, timing, offer, and follow-up, a result cannot tell you which change mattered.
3. Define the treatment output
For each treatment account, the system may:
read approved CRM fields
retrieve approved company sources
draft one short account hypothesis
attach the source, retrieval date, and evidence location
label the statement
supported,contradicted, ornot verified.
It may not:
choose additional accounts or contacts
infer sensitive personal characteristics
fabricate a trigger, relationship, or business event
create a performance or customer claim without approved evidence
decide the offer or price
send a message
continue when the entity or evidence is ambiguous.
A source link is necessary but not sufficient. The reviewer checks that the source belongs to the right entity, is current enough for the claim, and supports the exact wording. A page that mentions the company does not automatically support the proposed relevance angle.
The treatment output should remain useful even when the rep decides not to use it. “Not verified” is a valid result. It prevents a weak hypothesis from being polished into a confident claim.
4. Put a human at the final-send gate
The rep reviews the treatment output and the complete resulting message.
Record one decision:
Approved: the evidence supports the wording and the message is appropriate.
Edited: the rep changes a material claim, relevance link, offer framing, or message structure.
Rejected: the hypothesis or message is unsafe or not useful.
No send: the account or contact should not be approached.
HUMAN GATE — THE REVIEWER OWNS THE CONSEQUENCE
Approval applies to one recipient and one exact message version. Approving a research hypothesis does not approve a later message that the reviewer has not seen.
Use the same basic quality standard for the control group. Human review is not the treatment variable.
The reviewer checks the recipient, source-backed relevance, offer, claims, subject, required identification or opt-out mechanism, and final text. The system records the decision and the exact version. If the message changes materially after approval, it returns to review.
5. Measure four layers
Audience integrity
assigned accounts that met every criterion
exclusions and reasons
incorrect entity or contact matches
group imbalances discovered after assignment.
Message quality
drafts approved without material correction
edits to factual or relevance claims
hypotheses rejected or left
not verifiedreviewer time per message.
Commercial response
delivered messages, when the channel provides reliable status
positive replies under a pre-written coding rule
qualified conversations
accepted next steps.
Risk and trust
opt-outs and complaints
wrong-person or wrong-company incidents
unsupported claims caught before send
policy, consent, or suppression exceptions
duplicate or unapproved sends.
Do not hide the denominator. Report results per assigned account, then separately explain exclusions, non-sends, unavailable delivery states, and missing observations.
Opens are not the primary measure. Tracking can be incomplete, and an open does not prove relevance or commercial intent. A qualified conversation should meet a rule you wrote before the pilot—for example, the account confirms the problem is relevant and agrees to a defined next step.
PROOF STANDARD — DENOMINATOR BEFORE NARRATIVE
With 25 accounts per group, one reply can create a large-looking percentage difference. Read the counts together with corrections, reviewer effort, reply quality, exclusions, and risk incidents.
6. Predefine the interpretation
Write the decision rules before sending.
An example:
Expand to another 50-account batch if treatment quality is at least as good as control, no material risk incident occurs, reviewer time remains acceptable, and the qualified-conversation signal is directionally better.
Revise if the hypothesis is useful but the system often selects weak evidence, matches the wrong entity, or requires heavy corrections.
Stop if there is a material privacy, suppression, fabrication, wrong-entity, duplicate-send, or unapproved-send incident; or if treatment messages are consistently less relevant.
Define “acceptable” before the run. You might set a maximum reviewer-time budget, a zero-tolerance risk event, and a minimum quality condition. Do not invent those thresholds after seeing results.
Do not declare a winner from one extra reply. Pilot and experimental-design guidance supports prespecifying objectives and measures. Statistical guidance warns against treating a single threshold or summary statistic as complete evidence. The purpose of this batch is to decide what to test next.
7. Keep the run log
For each assigned account, preserve:
experiment_id
account_id
group
eligibility_check
exclusion_reason, if any
treatment_source_records, if applicable
draft_version
review_decision
material_edits
send_status
outcome_code
risk_or_policy_flags
Keep access limited and retention proportionate. Do not copy entire CRM records or unnecessary personal data into the experiment log.
The log lets you distinguish a weak model output from a targeting, offer, deliverability, or reviewer problem. It also makes uncertain tool state visible. If the sender says “request timed out,” the log should show whether the message is pending, confirmed sent, or blocked for reconciliation.
8. Handle failures and exceptions
Failure | System response | Human fallback |
|---|---|---|
Entity match is ambiguous | Stop treatment generation | Research manually or exclude |
Source does not support the hypothesis | Mark | Use a neutral message or no send |
Contact is suppressed or opted out | Block the send | No outreach |
Personal-data source is not approved | Do not use it | Use authorised company data only |
Rep materially changes treatment | Log the edit | Analyse separately |
Mail or CRM tool fails | Stop; preserve status | Use the normal manual process if safe |
Duplicate-send risk appears | Block all pending sends | Resolve state before resuming |
Never compensate for a low send count by adding unapproved accounts during the run. Record the shortfall and its reason.
9. Protect the experiment from false learning
Common errors include:
putting the best accounts in treatment
allowing treatment reps more time or coaching
measuring replies without classifying quality
excluding failed treatment outputs only after seeing them
changing the offer during the test
treating non-delivery as rejection
expanding before risk and corrections are reviewed.
Record deviations. A messy but honest pilot is more valuable than a clean-looking dashboard built from selective data.
Also separate an observed result from an explanation. If treatment produces more qualified conversations, the source-linked hypothesis may have helped—but the pilot may also contain segment imbalance, reviewer differences, or timing effects. State what happened, then list the plausible explanations and the next test that would reduce uncertainty.
10. Shut it down safely
Before launch, confirm that the owner can:
pause the experiment and all pending messages
revoke the research and sending credentials
preserve assigned groups and review logs
identify messages that were approved, sent, or still pending
apply new suppression information immediately
return the team to the current approved process.
One material unapproved-send, fabricated-claim, or suppression failure should trigger an immediate stop and review.
The experiment card
Complete this before the first account:
Field | Your decision |
|---|---|
Question | |
Account criteria | |
Exclusions | |
One treatment variable | |
Human reviewer | |
Primary quality measure | |
Qualified-conversation rule | |
Reviewer-time limit | |
Risk stop rule | |
Observation window | |
Expand / revise / stop decision |
This is one part of the operating system. Use Start Here for the category map, and connect the outcome to the discovery-call-to-CRM and follow-up system once a qualified conversation begins.
Build the control layer before the volume layer
The point of this pilot is not to prove that a model can write. It is to prove that your team can run a controlled sales experiment, see the evidence behind a proposed claim, keep a person in charge, and stop safely when the system crosses a boundary.
The free Agentic Sales System Starter Kit gives you the experiment card, review gates, permission map, failure and fallback sheet, and evidence record used in this design.
Run one bounded test with one variable, one human gate and an honest scorecard.
Sources and further reading
This article provides general operational and compliance information, not legal advice.
Build systems your team trusts—and your pipeline can prove.

