Test & Debug

See exactly why your agent said what it said, then fix it

Every reply comes with a complete record: which step and character answered, how scenarios scored, what the agent was thinking, every tool call and the exact prompts sent to the model. Test safely in sandboxes, rate a reply, and let the AI propose the fix.

Clicking a reply opens the debugger, which cycles through Why, Thinking, Tools and Prompts for that turnTest · Happy path (SMS)Hi! Is there parking?Yes! Free parking right behindthe studio. Want me to bookyour first class?Yes please!Type as the contact…Debugger · turn 2WhyThinkingToolsPromptsStepQualify leadCharacterAva · claude-sonnet-5Exit takennone (objective pending)Learnedwants: parking infoSCENARIO SCORESPricing question1/10Wants a human0/10Opt-out0/10STATE{ step: "qualify", objectives: { budget: null, goal: "beginner classes" } }The contact is asking about parking. That's afactual question about the studio, so I shouldcheck the knowledge base rather than guess.Found: free parking behind the building.Budget objective is still open, but it's tooearly to ask. Answer, then offer to book thefirst class to keep momentum.812 reasoning tokenssearch_knowledge240 ms{ query: "parking at studio" }→ 2 passages · 0.91, 0.74save_objective12 ms{ goal: "beginner classes" }→ savedsend_reply8 ms{ text: "Yes! Free parking…" }→ queued (test: simulated)systemYou are Ava, a friendly concierge…## Step: Qualify leadObjectives: budget (required), goal…Tools: search_knowledge, …user<message>Hi! Is there parking?</message>assistantYes! Free parking right behind…in 3,412 · out 92 · 1.4 s · $0.0118
4debugger views per reply
0messages sent from tests
1 clickto replay a real thread in a test
Side by sideold vs new reply
Deep dive

A closer look at Test & Debug

01 · Sandbox

Test conversations that can't embarrass you

Test sessions are saved per agent. Pick the channel, play the contact, reset and try again, and keep several side by side. Nothing reaches a real inbox and every action is simulated, while the builder's test tab runs your unsaved canvas.

  • Saved, renameable test sessions per agent
  • Pick the channel to test SMS vs email behaviour
  • Actions simulated, nothing sent
  • Test the unsaved canvas before you publish
A saved test session on the SMS channel where a booking tool call is simulated and nothing is sent to a real inboxTEST SESSIONSHappy pathAngry customerPrice objectionWants a human+ New sessionCHANNELSMSEmailWhatsAppCan I book a trial for Sat?Sure! 10am or 1pm?10ambook_appointmentGym Intro · Sat 10:00simulatedYou're booked for Sat at 10am!Nothing reaches a real inbox or calendar
02 · Debugger

Why, Thinking, Tools and Prompts

Click any reply. Why shows the step and character that answered, scenario scores, exits, what was learned and the full state. Thinking shows the reasoning per step. Tools lists every call and result. Prompts shows the exact requests to the model with tokens and timing.

  • Why: step, character, scenario scores, exits and state
  • Thinking: reasoning for each step
  • Tools: every call with its result
  • Prompts: system prompt, input, output, tokens and time
  • A run timeline with how long every event took
A run timeline lists every event of a turn with a waterfall of how long each tookRUN TIMELINE · Jordan Lee$0.0041 · 2,980 tokensInbound messageSMS · 14:02:11Scenario scoring3 scenarios · max 2/10Step: Book callMax · haiku-4.5Thinking214 tokenscheck_availabilityGym Intro · Thubook_appointmentThu 15:30 · confirmedReply sent1 message · 2.9 s total0 s2.9 s
03 · Fix loop

Rate it, diagnose it, apply it, re-run it

Thumbs-down a reply and say what went wrong. Diagnose reads the whole trace and proposes concrete edits to instructions, exits, objectives, scenario triggers or knowledge, shown as before and after. Apply them, then re-run the turn to compare old and new replies side by side.

  • Rate replies and describe the problem
  • AI-proposed, concrete before/after edits
  • Apply to the agent or onto the canvas
  • Re-run the turn and compare side by side
A reply is rated down, the AI diagnoses the trace and proposes an instruction change, which is applied and the turn re-run side by side1 · RATE“Should have offered to bookonce they said they're ready.”2 · DIAGNOSEReading the trace…The step has no exit for “ready tobook”, so the agent kept qualifying.1 edit proposed3 · APPLY · Qualify lead › instructions- Collect budget and timeline before anything else.+ Collect budget and timeline, but if they say they're+ ready, take the “Ready to book” exit right away.Applyapplied4 · RE-RUN THIS TURNBEFOREGreat! And roughly what budgetdid you have in mind?AFTERAmazing! I have Thu 3pm or Fri10am, which works best?exit: Ready to book
Under the hood

The details

Real conversations too

Live threads store the same debug data, available from the run timeline.

Debug in test

Copy a real thread into a test session and replay its last message.

Fallback attempts

Every model attempt, including failed ones, is shown with its error.

Costs

Token usage and cost for each model call.

A/B scoring

Lead score and sentiment scoring appear as their own debugger line.

Diagnose over MCP

Run diagnose and re-run from Claude Code or any MCP client.

Use cases

Who it's for

Pre-launch QA

Walk through every branch of a flow with simulated contacts before going live.

Incident review

Replay the exact conversation that went wrong and see which step misfired.

Prompt iteration

Tune instructions with side-by-side re-runs instead of guesswork.

FAQ

Test & Debug: common questions

Do test conversations send real messages?

No. Nothing is sent to real inboxes and actions are simulated.

Can I debug a real conversation?

Yes. Real runs store the same trace, and Debug in test replays a thread in a sandbox.

What does Diagnose change?

It proposes edits; nothing changes until you apply them.

Can I see the raw prompt?

Yes. The Prompts view shows the exact system prompt, input, output, thinking, tokens and time.

Get started

Put it to work on your own inbox

Connect an inbox, describe the job, test it in a sandbox and go live with exactly the autonomy you're comfortable with.