The voice conversation system was sending every partial transcription to Ollama, causing:
- Multiple LLM requests for a single sentence
- Slow response times (Ollama processing each word separately)
- Confusing conversation flow
- Wasted compute resources
User says: "Hello, how are you today?"
Old behavior:
- Whisper outputs: "Hello" → Ollama processes → TTS speaks
- Whisper outputs: "how" → Ollama processes → TTS speaks
- Whisper outputs: "are you" → Ollama processes → TTS speaks
- Whisper outputs: "today" → Ollama processes → TTS speaks
Result: 4 separate LLM calls, overlapping TTS, chaos!
Added speech accumulation with silence detection in the GPT agent:
- Voice Input → Whisper transcribes continuously
- GPT Agent → Accumulates transcriptions in buffer
- Silence Detection → After 2 seconds of silence, process complete sentence
- Single LLM Call → Send accumulated text to Ollama
- TTS Response → Speak the complete response
User says: "Hello, how are you today?"
New behavior:
- Whisper outputs: "Hello" → buffered
- Whisper outputs: "how" → buffered
- Whisper outputs: "are you" → buffered
- Whisper outputs: "today" → buffered
- [2 seconds silence]
- Buffer contains: "Hello how are you today"
- Single Ollama call → Single TTS response
Result: 1 LLM call, clean conversation!
self.speech_buffer = "" # Accumulates partial transcriptions
self.last_speech_time = None # Tracks when last speech received
self.silence_threshold = 2.0 # Seconds of silence before processing
self.is_processing = False # Prevents new input during LLM call- Runs every 500ms
- Checks if 2 seconds passed since last speech
- Processes accumulated buffer when silence detected
- Clears buffer after processing
is_processingflag prevents new speech during LLM call- Avoids interrupting ongoing conversation
- Ignores partial transcriptions while waiting for response
You can adjust the silence threshold in the GPT agent:
self.silence_threshold = 2.0 # Increase for slower speakers
# Decrease for faster response- Single LLM Call: One request per complete sentence
- Faster Response: No wasted processing on partial words
- Natural Flow: Wait for complete thought before responding
- Better Context: LLM sees full sentence, not fragments
- Resource Efficient: Fewer API calls, less compute
To test the fix:
# Rebuild and restart
docker-compose down
docker-compose up --build
# In another terminal, monitor topics
ros2 topic echo /voice_text # See partial transcriptions
ros2 topic echo /aether_response # See complete responses
# Speak a sentence and wait 2 seconds
# You should see ONE response after silence- Voice Activity Detection (VAD): More sophisticated silence detection
- Configurable Threshold: ROS parameter for silence duration
- Visual Feedback: LED or UI showing "listening" vs "processing"
- Wake Word: Only activate on "Hey Aether" or similar
- Interrupt Handling: Allow user to interrupt long responses