A/B testing

Stop guessing which script converts. Prove it.

Add an A/B test node, connect 2–4 variants, and ReplySetter splits new conversations between them. Each conversation keeps its variant, wins are checked every turn and in your CRM, and results are only marked proven when the statistics say so.

New conversations are split between variant A and variant B; B converts more often and is proven the winner at 97% confidenceNEW CONVERSATIONS50 / 50A/B testA · ControlOpens with a questionbookedbookedB · ChallengerOpens with a complimentbookedbookedGOAL · BOOKED18.0%A · 412 conv26.9%B · 405 convLIFT+49%CONFIDENCE97% · needs 95%B proven
2–4variants per test
95%confidence before a winner is proven
30conversations minimum per variant
60 daysof CRM win checks
Deep dive

A closer look at A/B testing

01 · Split

Sticky variants on real conversations

A conversation is assigned a variant at random the first time it reaches the test and keeps it, even if its flow is reset. The node says nothing to the contact. Duplicate the control's step for a new variant in one click, then change the character, prompt or offer.

  • 2–4 variants with custom shares, e.g. 50/50
  • Sticky assignment per conversation
  • Copy the control step for a new variant
  • Test runs are split but never counted
Jordan is assigned variant B at random; after the flow is reset Jordan still gets B, while Sam gets AVariant AStep: Warm openerVariant BStep: Direct offerJordanJordan · random → BSamSam · random → A3 days later · flow resetJordanJordan · sticky → B againTest & Debug runs are split so you can try them, but never counted in results
02 · Goals

Win on bookings, tags, replies, lead score or sentiment

Choose a win: an appointment booked, a step or Stop reached, a tag applied or a reply received. Or compare averages: after each contact reply the conversation is scored 0–100 for lead quality or sentiment. On GoHighLevel, bookings and tags are also checked in the CRM every 15 minutes for 60 days.

  • Binary goals: booking, step, tag or reply
  • Score goals: lead score or sentiment 0–100
  • Optional “what makes a good lead” guide
  • CRM checks catch bookings made by link or by your team
With a lead-score goal, each conversation is scored 0 to 100 after every reply; bookings made in the CRM also count as winsGOALBooked an appointmentReached a step or StopGot a tagRepliedHighest avg lead scoreHighest avg sentiment0100556778lead score · after each replySounds interestingWhat would it cost for 3?Can we start next week?CRM check every 15 minutes, for 60 daysBookings made by link or by your team, and tags added by workflows, count toobooked via link: win
03 · Statistics

A winner only when it's really a winner

Conversions use a two-proportion z-test and scores a Welch test. A challenger is proven when every variant has at least 30 conversations and it beats the control at 95% confidence, made stricter with more variants. Set an end date, and on that date new conversations go to the proven winner.

  • Two-proportion z-test and Welch t-test
  • Bonferroni: 97.5% with three variants, 98.3% with four
  • Sample size estimate to detect a 20% lift
  • Make the winner main with one click
As sample size grows the two variants' distributions separate until confidence passes the Bonferroni-corrected thresholdA · controlB · challengern = 30 per variantn = 140 per variantn = 410 per variantCONFIDENCE B BEATS A95% (97.5% with 3 variants, 98.3% with 4)Early lookProven: B winsTwo-proportion z-test · Welch t-test for scoresMin. 30 conversations per variant
Under the hood

The details

Results

Conversations, wins, conversion rate or average score, lift and confidence per variant.

Early look

Results stay flagged as an early look until the end date.

Restart

Begin a fresh run whenever you change a variant.

Clean-up

Make a variant main and delete the steps only the others used.

Off by default

With A/B testing off, every conversation takes the first variant.

API & MCP

Fetch results programmatically with get_ab_results.

Use cases

Who it's for

Opening lines

Does a question or a compliment get more replies in the first SMS?

Offers

Free consultation vs. discounted first session, measured on bookings.

Personas

Formal vs. casual characters compared on average lead score.

FAQ

A/B testing: common questions

Will contacts see different variants on each message?

No. Assignment is sticky: a conversation keeps its variant for good.

Do test sessions pollute results?

No. Test & Debug runs are split so you can try them, but never counted.

How do I know when to stop?

Set an end date. Results show an early look until then, and the panel estimates the sample size you need.

What happens to conversations after I pick a winner?

New ones go to the winner. Conversations on a deleted step restart from Start.

Get started

Put it to work on your own inbox

Connect an inbox, describe the job, test it in a sandbox and go live with exactly the autonomy you're comfortable with.