See exactly why your agent said what it said, then fix it
Every reply comes with a complete record: which step and character answered, how scenarios scored, what the agent was thinking, every tool call and the exact prompts sent to the model. Test safely in sandboxes, rate a reply, and let the AI propose the fix.
A closer look at Test & Debug
Test conversations that can't embarrass you
Test sessions are saved per agent. Pick the channel, play the contact, reset and try again, and keep several side by side. Nothing reaches a real inbox and every action is simulated, while the builder's test tab runs your unsaved canvas.
- Saved, renameable test sessions per agent
- Pick the channel to test SMS vs email behaviour
- Actions simulated, nothing sent
- Test the unsaved canvas before you publish
Why, Thinking, Tools and Prompts
Click any reply. Why shows the step and character that answered, scenario scores, exits, what was learned and the full state. Thinking shows the reasoning per step. Tools lists every call and result. Prompts shows the exact requests to the model with tokens and timing.
- Why: step, character, scenario scores, exits and state
- Thinking: reasoning for each step
- Tools: every call with its result
- Prompts: system prompt, input, output, tokens and time
- A run timeline with how long every event took
Rate it, diagnose it, apply it, re-run it
Thumbs-down a reply and say what went wrong. Diagnose reads the whole trace and proposes concrete edits to instructions, exits, objectives, scenario triggers or knowledge, shown as before and after. Apply them, then re-run the turn to compare old and new replies side by side.
- Rate replies and describe the problem
- AI-proposed, concrete before/after edits
- Apply to the agent or onto the canvas
- Re-run the turn and compare side by side
The details
Real conversations too
Live threads store the same debug data, available from the run timeline.
Debug in test
Copy a real thread into a test session and replay its last message.
Fallback attempts
Every model attempt, including failed ones, is shown with its error.
Costs
Token usage and cost for each model call.
A/B scoring
Lead score and sentiment scoring appear as their own debugger line.
Diagnose over MCP
Run diagnose and re-run from Claude Code or any MCP client.
Who it's for
Pre-launch QA
Walk through every branch of a flow with simulated contacts before going live.
Incident review
Replay the exact conversation that went wrong and see which step misfired.
Prompt iteration
Tune instructions with side-by-side re-runs instead of guesswork.
Test & Debug: common questions
Do test conversations send real messages?
No. Nothing is sent to real inboxes and actions are simulated.
Can I debug a real conversation?
Yes. Real runs store the same trace, and Debug in test replays a thread in a sandbox.
What does Diagnose change?
It proposes edits; nothing changes until you apply them.
Can I see the raw prompt?
Yes. The Prompts view shows the exact system prompt, input, output, thinking, tokens and time.
Works hand in hand with
Visual job flows
Design agents on a canvas: steps with objectives, AI conditions, switches, scenarios and exits.
Explore BuildAny model + fallbacks
Anthropic, OpenAI or 350+ OpenRouter models, with five cross-provider backups and cost tracking.
Explore GrowA/B testing
Split conversations across 2–4 variants and prove a winner on bookings, tags, lead score or sentiment.
ExplorePut it to work on your own inbox
Connect an inbox, describe the job, test it in a sandbox and go live with exactly the autonomy you're comfortable with.