Turn-taking
Turn-taking: VAD + Smart Turn
jargo does natural turn-taking: the bot waits for a real end-of-turn instead of any pause, and the user can interrupt it mid-sentence (barge-in). Two local ONNX models drive this:
- Silero VAD for voice activity detection: is the user speaking at all.
- Smart Turn v3 for end-of-turn detection: has the user actually finished, or just paused mid-thought.
Both models are embedded in the binary (go:embed), so there is nothing to
download or locate at run time except the ONNX Runtime itself.
ONNX Runtime setup
The models run on the ONNX Runtime
, bound through
purego
and loaded at run time, so it
needs no C toolchain at build time, and the default CGO_ENABLED=0 build works.
The runtime shared library is not bundled; download a
build for your platform from the
releases page
and point
jargo at it:
# Linux
export JARGO_ONNXRUNTIME_LIB=/path/to/libonnxruntime.so
# macOS
export JARGO_ONNXRUNTIME_LIB=/path/to/libonnxruntime.dylib
# Windows
set JARGO_ONNXRUNTIME_LIB=C:\path\to\onnxruntime.dllIf JARGO_ONNXRUNTIME_LIB is unset, jargo looks for the library by its
conventional name on the loader’s default search path
(libonnxruntime.so/.dylib/onnxruntime.dll). When the runtime cannot be
loaded, the voice bot still runs. It falls back to STT endpointing for
turn-taking and loses barge-in.
How it fits the pipeline
Turn-taking is a small subsystem split across two processors:
- A
vadproc.Processorjust after the input transport runs the VAD on incoming audio (resampled to 16 kHz mono) and emitsVADUserStartedSpeakingFrame/VADUserStoppedSpeakingFrameplus a periodicUserSpeakingFrame. - A
turns.UserTurnProcessorafter the STT consumes those VAD frames and the transcripts, runs pluggable start and stop strategies, and emits the turn decisions:UserStartedSpeakingFrame+InterruptionFrameto open a turn (the interruption flushes in-progress bot audio for barge-in), andUserStoppedSpeakingFrameto close it. It also owns aUserIdleController(re-engage a silent user) and optional mute strategies.
The default strategies are VAD-or-transcription to start and Smart Turn v3 to stop, so a pause Smart Turn rates incomplete does not end the turn.
vd, _ := vad.NewSilero()
tr, _ := turn.NewSmartTurnV3()
vadProc := vadproc.New(vadproc.Config{VAD: vd})
turnsProc := turns.NewUserTurnProcessor(turns.Config{
Strategies: turns.UserTurnStrategies{
Start: turns.DefaultStartStrategies(), // VAD + transcription
Stop: []turns.StopStrategy{turns.NewTurnAnalyzerStop(turns.TurnAnalyzerConfig{Analyzer: tr})},
},
// Optional: re-engage a caller who goes silent.
IdleTimeout: 10 * time.Second,
OnIdle: func(ctx context.Context, c *turns.UserIdleController) error {
return c.Push(ctx, frames.NewTTSSpeakFrame("Are you still there?"), processor.Downstream)
},
})
pipe := pipeline.New(
t.Input(),
vadProc,
stt,
turnsProc,
agg.User(), // built with aggregators.WithTurnTaking()
llm,
tts,
rtvi.NewProcessor(),
t.Output(),
agg.Assistant(),
)With aggregators.WithTurnTaking(), the LLM runs when the turn is reported
complete and a finalized transcript is in hand, so Smart Turn, not STT
endpointing, decides when the bot responds. See
examples/voicebot
for the full wiring, and
Interruptions
for what barge-in does to the
pipeline.
Strategies, idle, mute, and LLM completion
The start/stop chains are pluggable (turns.StartStrategy / turns.StopStrategy):
VAD, transcription, min-words and wake-phrase starts; Smart-Turn, speech-timeout
and external stops, plus a deferred wrapper. turns.FilterIncompleteUserTurnStrategies
together with a turns.CompletionFilter placed after the LLM add the optional
✓/○/◐ LLM turn-completion gate, where the model itself judges whether the user’s
turn is semantically complete (prepend turns.CompletionInstructions to the
system prompt). Mute strategies (turns.NewAlwaysUserMute, …) suppress user
input while the bot speaks or a tool call runs.
Implementation notes
- VAD gating is confidence-only: jargo trusts Silero’s neural confidence rather than adding a separate volume threshold.
- Feature extraction for Smart Turn is a pure-Go reimplementation of Whisper’s
log-mel features; it and both models are validated to within
1e-3of the reference Python implementation by unit tests. - Smart Turn runs inside the
TurnAnalyzerStopstrategy, driven by the turn controller (~tens of ms per end-of-turn). See benchmarks for the performance picture and the planned FFT optimization.