A voice-first 2D real-time digital human that role-plays as an NUS (National University of Singapore) campus assistant. Ask about Computing programmes, faculties, campus locations — in English or 中文 — and the digital human answers in voice with mouth-sync, grounded in scraped NUS Computing pages via RAG.
Built on top of 李锟's ADH fork, with the AI Agent / RAG / ASR layers swapped for a free-tier Microsoft + Groq stack.
Looking for the non-technical pitch? See PRODUCT_BRIEF.md — one-page overview suitable for sharing with stakeholders. Full build journal, including pitfalls and rationale, lives in BUILD_LOG.md.
This repo today is a working proof-of-concept; the longer-term direction is:
- Kiosk deployment — a touchscreen in NUS lobbies / open-house / library, not a personal browser tab. That implies session timeouts, privacy-by-default logging, wake-word / tap-to-start UX, and audio tuning for noisy spaces.
- Backed by NUS AI Know — the current self-scraped RAG (
comp.nus.edu.sg, 24 chunks) and GitHub Models gpt-4o-mini are stopgaps. Plan is to swap both for NUS's internal AI Know RAG + chat API once integration is approved. The OutsideAgent contract is the only seam that needs to change. - Smarter agent loop — currently a one-shot retrieve-then-answer. Ideas on the table: tool-using agent (web search, NUS calendar, bus times), conversation memory beyond raw turns, persona-aware Live2D customization for NUS branding.
If you've seen better frameworks for this (e.g. live2d alternatives, end-to-end voice agents like Pipecat, Vapi, etc.), please open an issue — we're not married to ADH.
Full roadmap with priorities and open questions: ROADMAP.md.
- 🎙️ Voice in / voice out — speak naturally; VAD auto-detects start/stop, no buttons.
- 🌐 Bilingual — Whisper auto-detects English and Chinese (works mid-sentence too).
- 🧠 Multi-turn — last 8 turns of memory; pronouns like "there" / "它" resolve correctly.
- 📚 RAG-grounded — answers cite scraped content from
comp.nus.edu.sgrather than guessing. - 🛡️ Honest fallback — refuses to invent fees, dates, deans; redirects to
nus.edu.sg. - 💰 $0 demo cost — runs entirely on free tiers (GitHub Models + Groq + EdgeTTS).
A Windows browser hits a Next.js frontend (port 3000) running in WSL2 Ubuntu. The frontend captures mic audio, sends it to a FastAPI backend (port 8002, also in WSL2). The backend transcribes via Groq Whisper, then forwards the text to a custom Python agent (adh_ai_agent.nus_agent). That agent embeds the query via GitHub Models text-embedding-3-small, retrieves top-3 chunks from a local nus_rag.npz index, calls GitHub Models gpt-4o-mini with the chunks as context, and streams the response. EdgeTTS synthesizes the audio. The Live2D frontend syncs mouth movement to the audio.
Browser ──► Next.js (3000) ──► FastAPI (8002) ──► Groq Whisper (ASR)
├──► GitHub Models (embed + chat)
└──► EdgeTTS (TTS)
- Windows 10/11 with WSL2 enabled
- Ubuntu 22.04 WSL distro (
wsl --install -d Ubuntu-22.04) - A GitHub account with a Personal Access Token, scope:
Models: Read(github.com/settings/tokens) - A Groq account with an API key (console.groq.com/keys) — free tier is plenty
- Chrome or Edge browser (Web Speech / Whisper need media APIs Firefox lacks)
- ~1 GB free disk for WSL distro +
pnpm installartifacts
Don't bring up an NUS Pulse Secure VPN while bootstrapping — it kills WSL2 outbound networking.
# Inside WSL Ubuntu-22.04 as root
apt update && apt install -y python3-pip python3-venv python3-dev ffmpeg git curl
curl -LsSf https://astral.sh/uv/install.sh | sh
curl -fsSL https://deb.nodesource.com/setup_lts.x | bash - && apt install -y nodejs
npm install -g pnpmmkdir -p /root/work && cd /root/work
git clone https://github.com/freecoinx/awesome-digital-human-live2d.git
git clone https://github.com/<your-github>/nus-digital-human.git # this repoThe canonical sources live in agent/. Follow agent/README.md for the full step-by-step (8 manual steps total: bootstrap an adh_ai_agent Python project, copy nus_agent.py into it, drop whisperASR.py/whisperAPI.yaml into ADH's engine directory, patch ADH's __init__.py and config.yaml, and uv pip install -e ../adh_ai_agent).
Automating this with a single bootstrap.sh is on the ROADMAP. PRs welcome.
export GITHUB_TOKEN=ghp_xxx # your GitHub PAT
export GROQ_API_KEY=gsk_xxx # your Groq keycd /root/work/adh_ai_agent
uv run python scripts/build_rag_index.pyFrom Windows side, double-click start_demo.cmd. Or from WSL:
bash /mnt/c/Users/<you>/path/to/nus-digital-human/scripts/start_all.shThen open http://localhost:3000/sentio in an incognito window (so you get the auto-defaults).
See demo_questions.md for a curated 7-tier question bank (easy retrieval → multi-turn → mixed language → guard rails → reset → voice edge cases) plus a 4-minute live-demo recipe.
nus-digital-human/
├── README.md ← this file (developer onboarding)
├── PRODUCT_BRIEF.md ← English one-pager for non-technical stakeholders
├── ROADMAP.md ← priorities + open questions for NUS AI Know team
├── BUILD_LOG.md ← journal of what was built + 15 documented pitfalls
├── demo_questions.md ← curated Q&A + 4-minute live-demo recipe
├── start_demo.cmd ← Windows double-click launcher (calls start_all.sh)
├── stop_demo.cmd
├── agent/ ← canonical code installed into ADH at runtime
│ ├── README.md ← step-by-step install instructions
│ ├── nus_agent.py ← the AI agent (history + RAG + reset)
│ ├── whisperASR.py ← Groq Whisper engine for ADH
│ └── whisperAPI.yaml ← ADH config for the Whisper engine
└── scripts/ ← bash helpers (run inside WSL)
├── start_all.sh ← bring up backend + frontend (requires env tokens)
├── stop_all.sh
├── health.sh ← quick reachability + agent smoke test
├── build_rag_index.py ← scrape NUS pages → embed → save .npz
├── restart_backend_only.sh
├── rebuild_frontend.sh
└── test_*.sh ← smoke tests
- nus.edu.sg main pages are JS-rendered → BeautifulSoup gets nothing → not in RAG (only
comp.nus.edu.sgis). - History is a single in-memory global — multi-window users will share / collide. Restart backend resets it.
- Live2D character is the default Hiyori, not NUS-themed. Customizing requires a designer + Photoshop + Live2D Editor.
- "Speak end → first audio" latency is 3–5s on free tiers. See BUILD_LOG.md for breakdown.
PRs and issues welcome. Especially looking for help on:
- Better speech / agent frameworks — if you've used something that beats ADH + GitHub Models for kiosk-style voice agents, propose it. We're open to a rewrite if the gain is real.
- Live2D model for NUS persona — current default is the stock Hiyori character. Need a designer who can produce a Live2D model that fits NUS branding (Lion / staff / student persona).
- Knowledge base coverage — the RAG only covers
comp.nus.edu.sgbecausenus.edu.sgmain pages are JS-rendered. A Patchright/Playwright-based scraper (or first-class NUS AI Know integration) would broaden this. - Kiosk-form-factor UX — wake-word detection, idle-timeout reset, error-state UI, telemetry.
Before opening a non-trivial PR, please skim BUILD_LOG.md — it documents the 15 sharp edges we've already hit, plus the rationale for the current stack. Don't reintroduce solved problems.
- Awesome Digital Human Live2D — Live2D + FastAPI + Next.js scaffolding by @wan-h.
- 李锟的 ADH 教程 — the OutsideAgent pattern and overall architecture inspiration.
- Groq, GitHub Models, EdgeTTS — free-tier stack.
MIT (see LICENSE if added).