This website uses cookies

Read our Privacy policy and Terms of use for more information.

SYSTEM BUILD


THE AGENTIC SALES BRIEF / SYSTEM 05

The first useful AI-outbound experiment should not test how many messages a system can send.

It should test whether one bounded AI-assisted step improves outreach for an approved audience without increasing customer, compliance, or operational risk.

This field experiment uses 50 approved accounts, one control group, one treatment group, one changed variable, and a human final-send decision. It measures message quality, qualified conversations, reviewer effort, and risk together.

Fifty accounts is a practical learning batch. It is not statistical proof. The result can justify another test, expose a weak control, or tell you to stop. It cannot establish a universal benchmark for AI personalisation.

SYSTEM DEFINITION — ONE CHANGE, ONE DECISION
Keep the audience, offer, sender, channel, call to action, follow-up policy, and observation window stable. Change only the preparation step you want to evaluate.

The signal to isolate

Choose one narrow question:

Does a source-linked account hypothesis help a rep produce a more relevant first message than the team's current approved preparation process?

This is testable. “Does AI improve outbound?” is not.

The treatment is also intentionally modest. It does not ask an agent to find contacts, write a full sequence, select an offer, and send messages. It adds one research output before the normal human-controlled drafting and sending process.

If your inputs are not yet reliable enough to produce a cited account hypothesis, build the account research agent that cites every claim before running this test.

The system on one page

One changed variable, one exact approval and one denominator-based decision.

The flow is deliberately simple:

50 approved accounts
        ↓
documented assignment
        ↓
control: current process     treatment: one source-linked hypothesis
        ↓                                  ↓
same human review and exact final-message approval
        ↓
quality + reviewer effort + qualified conversations + risk
        ↓
expand / revise / stop

Every arrow needs an owner, a record, and a failure path. The experiment is not ready if the team cannot explain what happens when an entity match is uncertain, a source contradicts the proposed claim, or the sending tool returns an ambiguous status.

Assign the owners

One sales leader or RevOps operator owns the experiment. The account owner or assigned rep owns each final message.

The experiment owner is responsible for:

  • audience and exclusion rules

  • lawful and policy-compliant operating conditions

  • random or otherwise documented group assignment

  • reviewer instructions

  • logs, measures, and the observation window

  • pause, revise, expand, and shutdown decisions.

The rep is responsible for the truth and appropriateness of the message they send. The system cannot inherit that responsibility.

Write down who can pause the test. “Anyone can raise a concern” is not the same as naming the person who has the authority to stop pending sends.

Approve the inputs and tools

Use the minimum set needed for the experiment:

  • approved CRM account and contact fields

  • the current suppression and opt-out register

  • the account's official public sources

  • an approved offer and claims reference

  • a review workspace that preserves source links and edits

  • the normal approved sending channel, controlled by the rep.

The research component needs read access. It does not need permission to change opportunities, create contacts, or send email. The review layer may prepare one exact message version. The sending connection should accept only the version the rep approved and retain a message ID that prevents an uncertain retry from becoming a duplicate.

If enrichment, browsing, or model tools process data outside the team's existing environment, assess their access, retention, training, and contractual terms before the pilot. Do not add a tool simply because it can produce more personalisation fields.

1. Build the approved cohort

Select 50 accounts that meet the same written criteria. Use company-level criteria relevant to the offer, such as business model, region, team structure, or a verifiable operational signal.

Before assignment:

  • remove current customers, open opportunities, and active conversations unless the test explicitly covers them

  • apply suppression, opt-out, and do-not-contact records

  • exclude ambiguous entities and accounts with inadequate evidence

  • confirm that the contact and channel are permitted under applicable laws, contracts, and platform rules

  • minimise the personal data used.

The US Federal Trade Commission states that CAN-SPAM applies to commercial email, including business-to-business email. UK Information Commissioner's Office guidance distinguishes rules and data-protection considerations by subscriber and processing context. GDPR principles include purpose limitation, data minimisation, and accuracy, and Article 21 covers objections to direct marketing.

These sources support controls, not case-specific clearance. Markets and channels differ. “B2B” is not a universal exemption. Verify the requirements that apply before any real send.

Treat audience approval as an input, not an outcome. If eligibility is uncertain, exclude the account before assignment. Do not let the treatment group carry the burden of questionable data.

2. Split control and treatment

Assign 25 accounts to the control group and 25 to the treatment group.

Where practical, balance obvious factors such as segment, company size, or territory before random assignment. Record the method. Do not move an account between groups after seeing its draft or outcome.

Control

The rep follows the current approved preparation process.

Treatment

The rep receives one AI-assisted, source-linked account hypothesis before drafting the message.

Keep these stable across both groups:

  • audience criteria

  • offer and approved claims

  • sender profile

  • channel

  • call to action

  • follow-up policy

  • observation window

  • outcome-coding rules.

The experiment tests the research aid, not an entire new sales motion. If the treatment has different accounts, copy, timing, offer, and follow-up, a result cannot tell you which change mattered.

3. Define the treatment output

For each treatment account, the system may:

  • read approved CRM fields

  • retrieve approved company sources

  • draft one short account hypothesis

  • attach the source, retrieval date, and evidence location

  • label the statement supported, contradicted, or not verified.

It may not:

  • choose additional accounts or contacts

  • infer sensitive personal characteristics

  • fabricate a trigger, relationship, or business event

  • create a performance or customer claim without approved evidence

  • decide the offer or price

  • send a message

  • continue when the entity or evidence is ambiguous.

A source link is necessary but not sufficient. The reviewer checks that the source belongs to the right entity, is current enough for the claim, and supports the exact wording. A page that mentions the company does not automatically support the proposed relevance angle.

The treatment output should remain useful even when the rep decides not to use it. “Not verified” is a valid result. It prevents a weak hypothesis from being polished into a confident claim.

4. Put a human at the final-send gate

The rep reviews the treatment output and the complete resulting message.

Record one decision:

  • Approved: the evidence supports the wording and the message is appropriate.

  • Edited: the rep changes a material claim, relevance link, offer framing, or message structure.

  • Rejected: the hypothesis or message is unsafe or not useful.

  • No send: the account or contact should not be approached.

HUMAN GATE — THE REVIEWER OWNS THE CONSEQUENCE
Approval applies to one recipient and one exact message version. Approving a research hypothesis does not approve a later message that the reviewer has not seen.

Use the same basic quality standard for the control group. Human review is not the treatment variable.

The reviewer checks the recipient, source-backed relevance, offer, claims, subject, required identification or opt-out mechanism, and final text. The system records the decision and the exact version. If the message changes materially after approval, it returns to review.

5. Measure four layers

Audience integrity

  • assigned accounts that met every criterion

  • exclusions and reasons

  • incorrect entity or contact matches

  • group imbalances discovered after assignment.

Message quality

  • drafts approved without material correction

  • edits to factual or relevance claims

  • hypotheses rejected or left not verified

  • reviewer time per message.

Commercial response

  • delivered messages, when the channel provides reliable status

  • positive replies under a pre-written coding rule

  • qualified conversations

  • accepted next steps.

Risk and trust

  • opt-outs and complaints

  • wrong-person or wrong-company incidents

  • unsupported claims caught before send

  • policy, consent, or suppression exceptions

  • duplicate or unapproved sends.

Do not hide the denominator. Report results per assigned account, then separately explain exclusions, non-sends, unavailable delivery states, and missing observations.

Opens are not the primary measure. Tracking can be incomplete, and an open does not prove relevance or commercial intent. A qualified conversation should meet a rule you wrote before the pilot—for example, the account confirms the problem is relevant and agrees to a defined next step.

PROOF STANDARD — DENOMINATOR BEFORE NARRATIVE
With 25 accounts per group, one reply can create a large-looking percentage difference. Read the counts together with corrections, reviewer effort, reply quality, exclusions, and risk incidents.

6. Predefine the interpretation

Write the decision rules before sending.

An example:

  • Expand to another 50-account batch if treatment quality is at least as good as control, no material risk incident occurs, reviewer time remains acceptable, and the qualified-conversation signal is directionally better.

  • Revise if the hypothesis is useful but the system often selects weak evidence, matches the wrong entity, or requires heavy corrections.

  • Stop if there is a material privacy, suppression, fabrication, wrong-entity, duplicate-send, or unapproved-send incident; or if treatment messages are consistently less relevant.

Define “acceptable” before the run. You might set a maximum reviewer-time budget, a zero-tolerance risk event, and a minimum quality condition. Do not invent those thresholds after seeing results.

Do not declare a winner from one extra reply. Pilot and experimental-design guidance supports prespecifying objectives and measures. Statistical guidance warns against treating a single threshold or summary statistic as complete evidence. The purpose of this batch is to decide what to test next.

7. Keep the run log

For each assigned account, preserve:

experiment_id
account_id
group
eligibility_check
exclusion_reason, if any
treatment_source_records, if applicable
draft_version
review_decision
material_edits
send_status
outcome_code
risk_or_policy_flags

Keep access limited and retention proportionate. Do not copy entire CRM records or unnecessary personal data into the experiment log.

The log lets you distinguish a weak model output from a targeting, offer, deliverability, or reviewer problem. It also makes uncertain tool state visible. If the sender says “request timed out,” the log should show whether the message is pending, confirmed sent, or blocked for reconciliation.

8. Handle failures and exceptions

Failure

System response

Human fallback

Entity match is ambiguous

Stop treatment generation

Research manually or exclude

Source does not support the hypothesis

Mark not verified; do not phrase it as fact

Use a neutral message or no send

Contact is suppressed or opted out

Block the send

No outreach

Personal-data source is not approved

Do not use it

Use authorised company data only

Rep materially changes treatment

Log the edit

Analyse separately

Mail or CRM tool fails

Stop; preserve status

Use the normal manual process if safe

Duplicate-send risk appears

Block all pending sends

Resolve state before resuming

Never compensate for a low send count by adding unapproved accounts during the run. Record the shortfall and its reason.

9. Protect the experiment from false learning

Common errors include:

  • putting the best accounts in treatment

  • allowing treatment reps more time or coaching

  • measuring replies without classifying quality

  • excluding failed treatment outputs only after seeing them

  • changing the offer during the test

  • treating non-delivery as rejection

  • expanding before risk and corrections are reviewed.

Record deviations. A messy but honest pilot is more valuable than a clean-looking dashboard built from selective data.

Also separate an observed result from an explanation. If treatment produces more qualified conversations, the source-linked hypothesis may have helped—but the pilot may also contain segment imbalance, reviewer differences, or timing effects. State what happened, then list the plausible explanations and the next test that would reduce uncertainty.

10. Shut it down safely

Before launch, confirm that the owner can:

  1. pause the experiment and all pending messages

  2. revoke the research and sending credentials

  3. preserve assigned groups and review logs

  4. identify messages that were approved, sent, or still pending

  5. apply new suppression information immediately

  6. return the team to the current approved process.

One material unapproved-send, fabricated-claim, or suppression failure should trigger an immediate stop and review.

The experiment card

Complete this before the first account:

Field

Your decision

Question

Account criteria

Exclusions

One treatment variable

Human reviewer

Primary quality measure

Qualified-conversation rule

Reviewer-time limit

Risk stop rule

Observation window

Expand / revise / stop decision

This is one part of the operating system. Use Start Here for the category map, and connect the outcome to the discovery-call-to-CRM and follow-up system once a qualified conversation begins.

Build the control layer before the volume layer

The point of this pilot is not to prove that a model can write. It is to prove that your team can run a controlled sales experiment, see the evidence behind a proposed claim, keep a person in charge, and stop safely when the system crosses a boundary.

The free Agentic Sales System Starter Kit gives you the experiment card, review gates, permission map, failure and fallback sheet, and evidence record used in this design.

Run one bounded test with one variable, one human gate and an honest scorecard.

Sources and further reading

This article provides general operational and compliance information, not legal advice.

Build systems your team trusts—and your pipeline can prove.