Skip to content

Latest commit

 

History

History
110 lines (85 loc) · 3.39 KB

File metadata and controls

110 lines (85 loc) · 3.39 KB

Voice Flow Fix - Speech Accumulation

Problem

The voice conversation system was sending every partial transcription to Ollama, causing:

  • Multiple LLM requests for a single sentence
  • Slow response times (Ollama processing each word separately)
  • Confusing conversation flow
  • Wasted compute resources

Example of the Problem:

User says: "Hello, how are you today?"

Old behavior:
- Whisper outputs: "Hello" → Ollama processes → TTS speaks
- Whisper outputs: "how" → Ollama processes → TTS speaks  
- Whisper outputs: "are you" → Ollama processes → TTS speaks
- Whisper outputs: "today" → Ollama processes → TTS speaks

Result: 4 separate LLM calls, overlapping TTS, chaos!

Solution

Added speech accumulation with silence detection in the GPT agent:

New Flow:

  1. Voice Input → Whisper transcribes continuously
  2. GPT Agent → Accumulates transcriptions in buffer
  3. Silence Detection → After 2 seconds of silence, process complete sentence
  4. Single LLM Call → Send accumulated text to Ollama
  5. TTS Response → Speak the complete response

Example with Fix:

User says: "Hello, how are you today?"

New behavior:
- Whisper outputs: "Hello" → buffered
- Whisper outputs: "how" → buffered
- Whisper outputs: "are you" → buffered
- Whisper outputs: "today" → buffered
- [2 seconds silence]
- Buffer contains: "Hello how are you today"
- Single Ollama call → Single TTS response

Result: 1 LLM call, clean conversation!

Implementation Details

Speech Buffer

self.speech_buffer = ""           # Accumulates partial transcriptions
self.last_speech_time = None      # Tracks when last speech received
self.silence_threshold = 2.0      # Seconds of silence before processing
self.is_processing = False        # Prevents new input during LLM call

Silence Detection Timer

  • Runs every 500ms
  • Checks if 2 seconds passed since last speech
  • Processes accumulated buffer when silence detected
  • Clears buffer after processing

Processing Lock

  • is_processing flag prevents new speech during LLM call
  • Avoids interrupting ongoing conversation
  • Ignores partial transcriptions while waiting for response

Configuration

You can adjust the silence threshold in the GPT agent:

self.silence_threshold = 2.0  # Increase for slower speakers
                              # Decrease for faster response

Benefits

  1. Single LLM Call: One request per complete sentence
  2. Faster Response: No wasted processing on partial words
  3. Natural Flow: Wait for complete thought before responding
  4. Better Context: LLM sees full sentence, not fragments
  5. Resource Efficient: Fewer API calls, less compute

Testing

To test the fix:

# Rebuild and restart
docker-compose down
docker-compose up --build

# In another terminal, monitor topics
ros2 topic echo /voice_text        # See partial transcriptions
ros2 topic echo /aether_response   # See complete responses

# Speak a sentence and wait 2 seconds
# You should see ONE response after silence

Future Improvements

  1. Voice Activity Detection (VAD): More sophisticated silence detection
  2. Configurable Threshold: ROS parameter for silence duration
  3. Visual Feedback: LED or UI showing "listening" vs "processing"
  4. Wake Word: Only activate on "Hey Aether" or similar
  5. Interrupt Handling: Allow user to interrupt long responses