Skip to content

Repository files navigation

SAM

SAM - Smart Assistant Module

A local, always-on voice assistant that lives on your desktop

Python 3.11+ License: MIT Powered by Ollama UI: WebView2 / PyQt6 Platform: Windows

SAM orb states: idle, listening, thinking, speaking

The orb's four states: idle breathing, level-reactive listening, a sweeping thinking arc, speaking.



Setup Guide · Architecture · Roadmap · Changelog


What SAM actually is

SAM is a Windows background app that sits on your desktop as a small circular orb: out of the way until you need it, gone from the window stack, then instantly on top the moment you speak, type, or click it.

Say the wake word or press a hotkey -> SAM records you -> transcribes locally with faster-whisper -> either runs a matching OS command directly, or streams a reply from a local Ollama model -> speaks the answer back.

It is not a cloud assistant with a local UI bolted on. Voice capture, transcription, and the language model itself (via Ollama) all run on your machine. SAM's only unprompted network call is to localhost:11434 (your local Ollama server). A cloud fallback (Claude) is opt-in.


What's new in v0.4.9

Clipboard Awareness & Quick Actions Say "explain this", "translate this", or "özetle" and SAM reads the copied text via commands/clipboard.py and feeds it to the LLM as context. The typed-input box shows a detachable 📋 Attached badge whenever it finds clipboard text.

Instant Language Switcher "Türkçe konuş" / "Switch to English" / "auto language" lock STT and TTS to a language on the spot; also a one-click 🌐 Language menu in the tray.

Cyberpunk SFX Engine Procedurally synthesized micro sound effects (audio/sounds.py) for wake, success, and warning cues - no bundled audio files, toggle and volume live in Settings.

HUD Toast Notifications A translucent badge (ui/toast.py) appears beside the orb for zero-LLM feedback - volume changes, Spotify track skips, language switches, errors - and auto-fades after 1.6s.

See CHANGELOG.md for the full version history.


How it works

flowchart LR
    WW["Wake word"] --> ROUTE
    HK["Ctrl+Space"] --> ROUTE
    TX["Typed input"] -.skips recording.-> STT
    ROUTE(( )) --> REC["Recorder: VAD"]
    REC --> STT["STT: faster-whisper\n(live partial + final)"]
    STT --> INSTANT{"Predefined\nphrase?"}
    INSTANT -- yes --> TTS["TTS: edge-tts / pyttsx3"]
    INSTANT -- no --> CMD{"Matches a\ncommand pattern?"}
    CMD -- yes --> SYS["OS action\n(no shell, ever)"]
    CMD -- no --> LLM["Ollama / Claude\n(streaming)"]
    SYS --> TTS
    LLM --> TTS
    TTS --> IDLE["back to idle"]

    style SYS fill:#0d3b32,stroke:#00D4AA,color:#e8e8e8
    style LLM fill:#0d2a3b,stroke:#00BFFF,color:#e8e8e8
    style TTS fill:#1a1a24,stroke:#38F2D8,color:#e8e8e8
    style INSTANT fill:#1a1a24,stroke:#00D4AA,color:#e8e8e8
Loading

Everything is orchestrated by AppController (core/app.py) through PyQt signals: no component calls another directly. See docs/ARCHITECTURE.md for the full picture, including the state machine, threading model, and the z-order mechanics behind "out of your way until called."

Stage What runs
Wake word openwakeword (ONNX), continuous, low CPU
Recording RMS-based voice activity detection
Transcription faster-whisper (CTranslate2, int8): small model for live captions, full model for the final pass
Instant responses Dictionary lookup (commands/instant.py): never touches the LLM
Instant commands Regex router -> os.startfile / ctypes / list-form subprocess (never a shell)
Conversation Local Ollama, or Claude if configured
Speech out edge-tts (online voices) or pyttsx3 (fully offline)
Overlay Always-on orb + fading caption + typed-input box

Installation

Tip

Most users should use the installer. See the Setup Guide for a complete walkthrough.

Installer (recommended)From source (development)
SAM-Setup-0.4.9.exe

Per-user install, no admin needed. Optionally installs Ollama, pre-pulls the model, pre-downloads the speech model, and adds a startup entry.

git clone https://github.com/sametgurtuna/SAM.git
cd SAM
python -m venv venv
.\venv\Scripts\Activate.ps1
pip install -r requirements.txt

ollama pull qwen2.5:3b
python main.py

No build step for source use: it is a script you run directly. config.yaml is created from config.example.yaml on first run if it does not already exist.


Using SAM

Method Result
Say "Hey Sam" (default wake word) Orb lights up, starts listening
Press Ctrl+Space Same, no wake word needed
Press Ctrl+Shift+Space Opens a text box under the orb instead
Click the orb Same as the text hotkey
Ctrl + drag the orb Moves it: position is remembered
Right-click the tray icon Settings, mute wake word, clear memory, quit

Command reference

Matches here execute directly: no LLM round-trip, response in milliseconds. Everything else falls through to the local LLM.

CategoryExamples
Launch / close apps"open spotify" · "launch notepad" · "close discord"
Volume"volume up" · "set volume to 50" · "mute"
Media"play" · "pause" · "next track"
Spotify search ¹"play blinding lights on spotify"
Window / session"minimize all" · "lock screen"
Power ²"shutdown computer" · "restart computer"
Web"go to github.com" · "search for local weather"
Screenshot"take a screenshot"

¹ Needs Spotify Client ID/Secret: Settings -> Integrations, or the SPOTIFY_CLIENT_ID / SPOTIFY_CLIENT_SECRET environment variables. The OAuth token is cached in %LOCALAPPDATA%\SAM\, never in the project folder.

² Two-step by design. "shutdown computer" only arms the action and starts a 30-second countdown; say "confirm" to execute it, "cancel" to abort. The phrase must name an explicit object (computer/pc/machine/laptop), so "restart chrome" goes to the app handler, never the power handler.

Shells cannot be opened by voice or text, on purpose - see Security section below.


Configuration

Every key in config.yaml has a default in core/config.py, so a missing or partial file never breaks the app. Edit it via Settings in the tray menu (Appearance-page cosmetics apply live; audio/llm settings apply on restart), or directly via config.example.yaml.

Instant responses live in their own file. In an installed build SAM seeds a writable copy at %APPDATA%\SAM\knowledge\instant_responses.yaml on first launch and reads that one, so your edits survive updates. Settings -> Responses includes Edit Responses, Show Folder and Reload (applies changes without restarting SAM).

hotkey:
  trigger: ctrl+space          # hold to speak
  text_input: ctrl+shift+space # open the typed-input box

ui:
  orb:
    layer: auto        # auto (bottom until called, then on top) | topmost | normal
    click_through: true
    idle_animation: true
    idle_fps: 12        # SAM runs 24/7 - low idle cost
    active_fps: 60

wake_word:
  model: assets/models/hey_sam.onnx
  threshold: 0.5        # lower = triggers more easily

stt:
  model: small           # tiny | base | small | medium | large-v3
  device: cpu             # or cuda, with a working CUDA + cuDNN setup

llm:
  ollama:
    model: qwen2.5:3b
    autostart: true       # SAM starts "ollama serve" itself if it isn't running
    stop_on_exit: false   # never kill a server you already had running

Writing a custom command

A regex in commands/router.py plus a handler in commands/system.py (or a new module) that returns a spoken confirmation string:

# commands/router.py - inside _build_patterns()
patterns.append((
    re.compile(r"\b(what'?s|check) (my )?cpu (temp|temperature)\b", re.IGNORECASE),
    lambda m: system.get_cpu_temperature()
))
# commands/system.py
def get_cpu_temperature() -> str:
    """Handlers never raise - the router already wraps calls in try/except."""
    try:
        # Real OS side effects use list-form subprocess or os.startfile (never shell=True)
        ...
        return f"Your CPU is at {celsius:.1f} degrees."
    except Exception:
        return "Sorry, I couldn't read the CPU temperature."

Troubleshooting

Symptom Likely cause Fix
No LLM engine found in the log Ollama is not installed, or the model is not pulled Check the log for details; run ollama pull qwen2.5:3b
Wake word does not trigger Threshold too high, or wrong mic Lower wake_word.threshold (try 0.35); check your default input device
Whisper transcribes noise on silence Whisper hallucination behavior on background noise Raise audio.silence_threshold
Ctrl+Space does nothing Hotkey hook needs elevated access on some windows Run SAM as Administrator, or check logs/sam.log for a hotkey error
A second orb / doubled hotkeys Two SAM processes running SAM allows one instance only (named mutex) - check the tray before starting another
Settings will not save Installed build config lives in %APPDATA%\SAM\config.yaml Edit that file, or use the Settings window

Logs: logs/sam.log from source, %APPDATA%\SAM\logs\sam.log when installed.


Project layout

SAM/
├── assets/                    icon, activation chime, wake word model
├── audio/                     wake word, recorder (VAD), STT, TTS, procedural SFX
├── commands/                  regex router + OS side effects + clipboard reader
├── core/
│   ├── app.py                   AppController: state machine, wires everything together
│   ├── config.py                DEFAULTS + config.yaml loader/saver
│   ├── paths.py                 dev vs. frozen-exe path resolution, single-instance lock
│   ├── code_parser.py           extracts code blocks from LLM replies to the Desktop
│   └── installer_steps.py       SAM.exe --install-models (used by installer)
├── llm/                        LLMEngine base + Ollama/Claude engines + router + OllamaService
├── ui/
│   ├── web/                     Modern Cyberpunk Dark Webview interface (HTML5/CSS3/JS)
│   ├── web_settings.py          pywebview settings host with native Win32 single instance
│   ├── orb.py · caption.py      The always-on overlay
│   ├── toast.py                 Ephemeral HUD toast for zero-LLM feedback
│   ├── win32.py                 click-through, z-order, foreground-focus helpers
│   └── tray.py                  System tray integration
├── installer/SAM.iss          Inno Setup script
├── tools/make_icon.py         regenerates assets/icon.ico from the orb design
├── SAM.spec                   PyInstaller build spec
├── config.example.yaml        committed template (config.yaml is gitignored)
├── docs/ARCHITECTURE.md
├── setup.md
└── main.py

Security & Privacy

  • No shell, ever, from voice or text. faster-whisper occasionally hallucinates short phrases from silence or background noise. If a hallucination contained a command prompt request, opening a terminal unexpectedly is unsafe. cmd, powershell, wt, bash, and related executables are hard-blocked in commands/system.py, regardless of request origin.
  • Destructive actions are two-step. shutdown/restart only arm; a separate "confirm" executes within a 30-second window.
  • Transcripts never reach a shell. All OS actions use os.startfile() or list-form subprocess (never shell=True).
  • No telemetry. SAM's only self-initiated network calls are to your local Ollama server, Spotify (only if configured), and edge-tts (only if used instead of the offline pyttsx3). Claude is opt-in only.
  • Audio is not written to disk: processed in memory, discarded after transcription.
  • Secrets stay out of the repo. config.yaml is gitignored; API keys are read from the environment first. OAuth caches live in %LOCALAPPDATA%\SAM\.

Contributing

  1. Fork and branch: git checkout -b feature/your-feature
  2. Follow project standards: English identifiers/docstrings, config access through config.get(...), no shell=True.
  3. Verify manually with python main.py and inspect logs/sam.log.
  4. Open a pull request describing changes and verification steps.

License

MIT - see LICENSE.

Keep your data local.

About

A local, privacy-first Windows voice assistant powered by Ollama. No cloud, no telemetry — your voice never leaves your machine.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages