Voice Agents AI plumbing tool every team rebuilds from scratch β until now. Drop in a call recording. Get back everything: transcript, quality scores, sentiment shifts, latency breakdown, cost attribution, compliance flags. Normalized. Structured. Queryable. Works with any provider, any stack. One integration. Zero plumbing. Ship faster.
import audiotrace
report = audiotrace.analyze(
audio = "call_recording.wav",
metadata = {"agent_version": "v2.1", "provider": "vapi"}
)
print(report.quality.overall_score) # 0.87
print(report.sentiment.caller_frustration) # False
print(report.latency.llm_first_token_ms) # 420
print(report.events.drop_off) # False
print(report.cost.total_usd) # 0.063πΊ Watch the demo:
Every team building voice agents faces the same problem: raw audio is a black box. You can listen to recordings manually, or you can build your own signal extraction pipeline from scratch β but no open-source framework normalizes the full call into a structured, queryable object.
AudioTrace exists to be that shared layer. It handles the hard parts so you can focus on what you're building:
- Transcription with speaker diarization
- Silence gaps, interruptions, speaking pace, and pitch analysis
- Per-turn sentiment tracking and frustration detection
- Per-stage latency breakdown (STT β LLM β TTS β telephony)
- Unified cost calculation across any provider mix
- Compliance flag detection (PII leakage, consent gaps)
pip install audiotrace
# With specific provider adapter
pip install audiotrace[vapi]
pip install audiotrace[retell]
pip install audiotrace[twilio]
# Full install
pip install audiotrace[all]docker build -f docker/Dockerfile -t audiotrace .
docker run -it audiotraceRequirements: Python 3.9+, FFmpeg installed on system
import audiotrace
report = audiotrace.analyze(
audio = "call.wav",
metadata = {
"call_id": "abc123",
"agent_version": "v2.1",
"provider": "vapi",
"campaign": "healthcare_intake"
}
)
# Media
print(report.media.duration_ms) # int
print(report.media.codec) # str
# Transcript
print(report.transcript.full_text)
for turn in report.transcript.turns:
print(f"{turn.speaker}: {turn.text}")
# Quality
print(report.quality.overall_score) # float 0.0β1.0
print(report.quality.interruptions) # int
print(report.quality.silence_gaps) # List[Gap]
print(report.quality.speaking_pace_wpm) # float
# Sentiment
print(report.sentiment.overall) # float -1.0 to 1.0
print(report.sentiment.shift_points) # List[int] β turn indices
print(report.sentiment.caller_frustration)# bool
# Latency
print(report.latency.stt_ms) # int
print(report.latency.llm_first_token_ms) # int
print(report.latency.tts_ms) # int
print(report.latency.total_ms) # int
# Cost
print(report.cost.stt_usd) # float
print(report.cost.llm_usd) # float
print(report.cost.total_usd) # float
# Events
print(report.events.outcome) # "completed" | "dropped" | "failed"
print(report.events.drop_off_turn) # int | None
print(report.events.compliance_flags) # List[str]from audiotrace.adapters import VapiAdapter
adapter = VapiAdapter(api_key="...")
call = adapter.fetch_call(call_id="abc123")
report = audiotrace.analyze(call.audio, call.metadata)CallReport
βββ media
β βββ duration_ms: int
β βββ sample_rate_hz: int
β βββ channels: int
β βββ codec: str
β βββ file_size_bytes: int
β βββ file_format: str
β βββ bitrate_kbps: float
βββ transcript
β βββ full_text: str
β βββ turns: List[Turn] # speaker Β· text Β· start_ms Β· end_ms Β· confidence Β· words[]
β βββ language: str
β βββ diarization_confidence: float | None # pitch-fallback speaker separability (0-1); None if not measured
βββ quality
β βββ overall_score: float
β βββ interruptions: int
β βββ silence_gaps: List[Gap]
β βββ speaking_pace_wpm: float
β βββ pitch_variance: float
β βββ turn_length_avg_ms: float
βββ sentiment
β βββ by_turn: List[float]
β βββ overall: float
β βββ shift_points: List[int]
β βββ caller_frustration: bool
βββ latency
β βββ stt_ms: int
β βββ llm_first_token_ms: int
β βββ llm_full_response_ms: int
β βββ tts_ms: int
β βββ total_ms: int
β βββ waterfall: List[LatencySpan]
βββ cost
β βββ stt_usd: float
β βββ llm_usd: float
β βββ tts_usd: float
β βββ telephony_usd: float
β βββ total_usd: float
βββ events
βββ outcome: str
βββ drop_off: bool
βββ drop_off_turn: int | None
βββ intent_detected: str
βββ failure_type: str | None
βββ compliance_flags: List[str]
Provider adapters are TBD β not yet implemented. The integrations below are planned; today you pass a local audio file path to
analyze()directly. The adapter example above is illustrative of the intended API.
| Provider | Adapter | Status |
|---|---|---|
| Vapi | audiotrace[vapi] |
TBD |
| Retell | audiotrace[retell] |
TBD |
| Twilio | audiotrace[twilio] |
TBD |
| ElevenLabs | audiotrace[elevenlabs] |
TBD |
| Deepgram | audiotrace[deepgram] |
TBD |
| Custom webhook | CustomAdapter |
TBD |
AudioTrace builds on top of best-in-class audio libraries so you don't have to:
Raw audio file
β
βΌ
FFmpeg β format normalization, turn splitting
β
βββ Whisper β transcription
βββ pyannote β speaker diarization
βββ Librosa β silence gaps, pace, pitch, energy
βββ Transformers β sentiment, intent detection
β
βΌ
CallReport (Pydantic)
Part of the Lang ecosystem TBD
AudioTrace is the open-source foundation that powers two commercial products:
| Product | What it does | Built on |
|---|---|---|
| LangTrace | Live call observability & analytics dashboards | AudioTrace |
| LangGate | Pre-deploy simulation & CI/CD quality gate | AudioTrace |
AudioTrace is free and MIT-licensed. The commercial products are optional hosted layers on top.
Want it run for you today? We're taking a few paid design-partner pilots β "early access + we run it for you." See COMMERCIAL.md.
For quick testing or interactive analysis, you can use the provided runner script. It automatically handles virtual environment setup and dependency validation.
# Analyze default golden data fixture
./scripts/run.sh
# Concise per-section summary tables instead of the raw JSON
./scripts/run.sh --summary
# Playback, inferring speakers by pitch (no pyannote token needed)
./scripts/run.sh --playback --skip-pyannote
# Analyze a specific file
./scripts/run.sh path/to/your/audio.wavBefore submitting changes, ensure everything passes the local validation suite (formatting, linting, type-checking, and tests):
./scripts/test_local.sh testTreat a handful of representative recordings as golden fixtures, commit a baseline, and fail the build when a prompt/model/voice change makes the agent measurably worse β slower, colder, less compliant.
# 1. Commit a baseline from your golden calls (one time, and after intentional changes)
audiotrace baseline tests/calls -o baseline.json
# 2. Gate every change against it β exits non-zero on regression, writes per-call reports
audiotrace check tests/calls -b baseline.json --report audiotrace-reportA metric only fails the build when it drifts past its tolerance (quality β0.05, sentiment β0.10, latency +15%, cost +20%; frustration / drop-off / compliance have zero slack). New recordings not in the baseline are skipped, not failed.
Drop the gate into CI in a few lines. It installs AudioTrace, runs the check, and uploads the HTML report as an artifact even when the build fails:
# .github/workflows/voice-quality.yml
name: Voice quality
on: [pull_request]
jobs:
audiotrace:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: dimastatz/audiotrace@v1.2.1
with:
calls: tests/calls
baseline: baseline.jsonContributions are welcome β especially new provider adapters, persona definitions for simulation, and compliance rule sets.
git clone https://github.com/audiotrace/audiotrace
cd audiotrace
./scripts/test_local.sh test # Run all checks (formatting, lint, types, tests)See CONTRIBUTING.md for guidelines.
MIT β see LICENSE


