While the call lasts, the panel shows the last lines of what’s been said — both your agent and the visitor themselves — labeled with who’s speaking (“You” / “Assistant”). It’s the ONLY channel for captions: never a separate speech recognition done in the browser itself. If your agent doesn’t publish transcription yet (an older version of the voice module), the panel stays with just the state (“Listening…”, “Speaking…”) — it never invents a caption.
