Skip to content

Commit eca48eb

Browse files
committed
Merge remote-tracking branch 'origin/main' into claude/serene-heisenberg-7onpfp
# Conflicts: # aai_cli/commands/stream/_exec.py # aai_cli/streaming/session.py # tests/test_stream_exec.py
2 parents 508a17f + 410b650 commit eca48eb

38 files changed

Lines changed: 1546 additions & 200 deletions

‎README.md‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -4,7 +4,7 @@
44
[![License](https://img.shields.io/badge/license-MIT-D6402E)](https://github.com/AssemblyAI/cli/blob/main/LICENSE)
55
[![Docs](https://img.shields.io/badge/docs-assemblyai-D6402E)](https://www.assemblyai.com/docs)
66

7-
The AssemblyAI CLI (`assembly`) brings speech AI directly into your terminal: transcribe files, URLs, and YouTube/podcast pages, stream live audio, talk to a two-way voice agent, prompt the LLM Gateway, benchmark speech models, and scaffold ready-to-deploy starter apps.
7+
The AssemblyAI CLI (`assembly`) brings speech AI directly into your terminal: transcribe files, URLs, YouTube/podcast pages, and whole podcast RSS feeds, stream live audio, talk to a two-way voice agent, prompt the LLM Gateway, benchmark speech models, and scaffold ready-to-deploy starter apps.
88

99
<p align="center">
1010
<img src="assets/welcome.png" alt="The assembly CLI welcome screen, listing command groups for transcription, streaming, voice agents, app scaffolding, and account management" width="820">
@@ -44,7 +44,7 @@ That's it. Run `assembly onboard` for a guided tour, or see [Installation](#-ins
4444

4545
| Command | What it does |
4646
| :--- | :--- |
47-
| `assembly transcribe` | Transcribe files, URLs, YouTube/podcast pages, directories, globs, or bucket storage (`s3://`, `gs://`, `az://`) — with speaker labels, PII redaction, summarization, SRT/VTT captions, and resumable batch runs |
47+
| `assembly transcribe` | Transcribe files, URLs, YouTube/podcast pages, podcast RSS feeds, directories, globs, or bucket storage (`s3://`, `gs://`, `az://`) — with speaker labels, PII redaction, summarization, SRT/VTT captions, and resumable batch runs |
4848
| `assembly stream` | Real-time transcription from your microphone, a file, or a URL — on macOS it can capture system audio too |
4949
| `assembly dictate` | Push-to-talk dictation: press Enter to record, Enter again for instant text (Sync STT API, up to 120 s per utterance) |
5050
| `assembly agent` | Full-duplex spoken conversation with a voice agent, right in your terminal |
@@ -285,11 +285,13 @@ assembly transcribe video.mp4 -o srt # captions
285285
assembly transcribe call.mp3 --speaker-labels --summarization --json
286286
```
287287

288-
Transcribe in batches — a directory, a glob, or a piped list, resumable on re-run:
288+
Transcribe in batches — a directory, a glob, a piped list, or a whole podcast
289+
RSS feed (every episode becomes one source), resumable on re-run:
289290

290291
```sh
291292
assembly transcribe ./recordings
292293
assembly transcribe "s3://bucket/calls/*.mp3" # needs: pip install s3fs
294+
assembly transcribe "https://feeds.simplecast.com/54nAGcIl" # every episode in the feed
293295
find . -name "*.wav" | assembly transcribe --from-stdin
294296
```
295297

‎REFERENCE.md‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -107,12 +107,15 @@ writes:
107107

108108
- `<stem>.txt` — the transcript, one finalized turn per line (flushed live).
109109
- `<stem>.wav` — the recorded audio, 16-bit mono PCM. Suppress it with
110-
`--no-save-audio` to keep only the text.
110+
`--no-save-audio` to keep only the text. Under `--system-audio` the two channels
111+
can't share a file, so each gets its own `<stem>-you.wav` / `<stem>-system.wav`.
111112
- `<stem>.md` — written when `--llm "…"` is also passed: the final answer of the
112113
live prompt chain, captured as a note next to the transcript.
113114
- `<stem>.aai.json` — a metadata sidecar so a list/browse UI needs no transcript
114115
parsing: `{"title", "date", "duration_seconds", "speakers", "turns",
115-
"transcript", "audio", "note"}` (`audio`/`note` are `null` when not written).
116+
"transcript", "audio", "note"}`. `audio` is the list of WAV file names (empty
117+
under `--no-save-audio`, two entries under `--system-audio`); `note` is `null`
118+
when no `--llm` note was written.
116119

117120
`--name "Title"` slugs an explicit title into the stem; `--auto-name` instead
118121
derives that title from the transcript via the LLM Gateway once the stream ends,

‎aai_cli/app/transcribe/feed.py‎

Lines changed: 123 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,123 @@
1+
"""Podcast RSS/Atom feed expansion for ``assembly transcribe``.
2+
3+
A feed URL names a whole show, so transcribing it means transcribing every
4+
episode. ``feed_episode_urls`` fetches the URL and, when ``feedparser`` recognizes
5+
it as an RSS or Atom feed, returns its episode enclosure URLs (in feed order —
6+
newest first) for the batch path to transcribe, one resumable sidecar per episode.
7+
The enclosures are direct media URLs the API fetches itself, so — unlike a YouTube
8+
or podcast *page*, which yt-dlp downloads first — no local download step is needed.
9+
10+
Detection is deliberately narrow so a direct media URL or ordinary web page still
11+
falls through to the single-source path untouched (and is never fetched twice):
12+
only an http(s) URL whose path is feed-shaped — no extension, or one of
13+
``.xml``/``.rss``/``.atom`` — and that no dedicated yt-dlp extractor already claims
14+
is sniffed, the response body is bounded, and only content ``feedparser`` parses as
15+
a real feed with at least one enclosure is treated as a feed. We hand ``feedparser``
16+
the already-fetched bytes (never the URL) so our bounded, safe fetch below stays the
17+
only network path.
18+
"""
19+
20+
from __future__ import annotations
21+
22+
from pathlib import PurePosixPath
23+
from urllib.parse import urlsplit
24+
25+
from pydantic import BaseModel, Field
26+
27+
from aai_cli.core import youtube
28+
29+
# A feed lives at an extensionless URL (e.g. feeds.simplecast.com/<id>) or a feed
30+
# document (.xml/.rss/.atom). Every other path — .mp3, .txt, .pdf — is never a feed,
31+
# so it is left for the single-source path and never fetched here.
32+
_FEED_URL_SUFFIXES = frozenset({"", ".xml", ".rss", ".atom"})
33+
34+
# Bound the download so a hostile or huge URL can't exhaust memory; 10 MB of feed
35+
# already holds thousands of episodes, far past any realistic batch.
36+
_MAX_FEED_BYTES = 10 * 1024 * 1024 # pragma: no mutate -- tuning knob, not behavior
37+
_FETCH_TIMEOUT_SECONDS = 15.0 # pragma: no mutate -- tuning knob, not behavior
38+
39+
40+
class _Enclosure(BaseModel):
41+
"""One ``<enclosure>`` / Atom enclosure link; ``href`` is the media URL."""
42+
43+
href: str = ""
44+
45+
46+
class _Entry(BaseModel):
47+
# default_factory (not a shared `= []`) so each entry gets its own list, and the
48+
# typed factory keeps the field's element type known under pyright strict.
49+
enclosures: list[_Enclosure] = Field(default_factory=list[_Enclosure])
50+
51+
52+
class _ParsedFeed(BaseModel):
53+
"""The slice of feedparser's untyped result we use, validated into a real type
54+
(the project pattern for untyped third-party returns — cf. core/wer.py)."""
55+
56+
# feedparser sets ``version`` to a non-empty id ("rss20", "atom10", …) for a
57+
# recognized feed and to "" for anything it doesn't recognize as one.
58+
version: str = ""
59+
entries: list[_Entry] = Field(default_factory=list[_Entry])
60+
61+
62+
def feed_episode_urls(url: str) -> list[str] | None:
63+
"""The episode media URLs if `url` is a podcast feed, else ``None``.
64+
65+
Returns ``None`` (stay single-source) for a direct-media URL, a yt-dlp page,
66+
an unreachable URL, or any content that isn't a feed carrying enclosures.
67+
"""
68+
if not _looks_like_feed_url(url) or youtube.is_downloadable_url(url):
69+
return None
70+
body = _fetch(url)
71+
if body is None:
72+
return None
73+
return _episode_urls(body)
74+
75+
76+
def _looks_like_feed_url(url: str) -> bool:
77+
"""True when the URL path is feed-shaped: extensionless or a feed document."""
78+
suffix = PurePosixPath(urlsplit(url).path).suffix.lower()
79+
return suffix in _FEED_URL_SUFFIXES
80+
81+
82+
def _episode_urls(body: str) -> list[str] | None:
83+
"""The enclosure URLs in a feed body, deduped in document order; ``None`` when
84+
feedparser doesn't recognize it as a feed or it carries no enclosures."""
85+
import feedparser
86+
87+
# feedparser ships only partial inline types (its parse signature is Unknown),
88+
# so the result is validated through _ParsedFeed below; mirror remotefs.py's
89+
# fsspec shim in ignoring the unavoidable unknown-member report on the call.
90+
raw = feedparser.parse(body) # pyright: ignore[reportUnknownMemberType]
91+
parsed = _ParsedFeed.model_validate(raw)
92+
if not parsed.version:
93+
return None
94+
urls = [enc.href for entry in parsed.entries for enc in entry.enclosures if enc.href]
95+
deduped = list(dict.fromkeys(urls))
96+
return deduped or None
97+
98+
99+
def _fetch(url: str) -> str | None:
100+
"""Up to ``_MAX_FEED_BYTES`` of `url` decoded as text, or ``None`` on any failure
101+
or when the response is obviously binary media (audio/video/image)."""
102+
import httpx2 as httpx
103+
104+
chunks: list[bytes] = []
105+
try:
106+
with (
107+
httpx.Client(timeout=_FETCH_TIMEOUT_SECONDS, follow_redirects=True) as client,
108+
client.stream("GET", url) as response,
109+
):
110+
if not response.is_success:
111+
return None
112+
content_type = response.headers.get("content-type", "").lower()
113+
if content_type.startswith(("audio/", "video/", "image/")):
114+
return None
115+
total = 0
116+
for chunk in response.iter_bytes():
117+
chunks.append(chunk)
118+
total += len(chunk)
119+
if total >= _MAX_FEED_BYTES:
120+
break
121+
except (httpx.HTTPError, OSError):
122+
return None
123+
return b"".join(chunks).decode("utf-8", "replace")

‎aai_cli/app/transcribe/run.py‎

Lines changed: 6 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -356,7 +356,12 @@ def run_transcribe(opts: TranscribeOptions, state: AppState, *, json_mode: bool)
356356
transcribe_validate.validate_speakers_expected(merged)
357357

358358
sources = transcribe_sources.expand_sources(
359-
opts.source, from_stdin=opts.from_stdin, sample=opts.sample
359+
opts.source,
360+
from_stdin=opts.from_stdin,
361+
sample=opts.sample,
362+
# --show-code must never touch the network; skip the feed probe and treat a
363+
# URL as a single source for code generation.
364+
detect_feeds=not opts.show_code,
360365
)
361366
if sources is not None:
362367
transcribe_sources.reject_single_source_flags(

‎aai_cli/app/transcribe/sources.py‎

Lines changed: 22 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -49,24 +49,41 @@
4949
_GLOB_CHARS = frozenset("*?[")
5050

5151

52-
def expand_sources(source: str | None, *, from_stdin: bool, sample: bool) -> list[str] | None:
52+
def expand_sources(
53+
source: str | None, *, from_stdin: bool, sample: bool, detect_feeds: bool = True
54+
) -> list[str] | None:
5355
"""The batch source list, or ``None`` when this is a single-source invocation.
5456
5557
Batch mode triggers on ``--from-stdin``, a directory (scanned recursively for
56-
audio files), a glob pattern that names no existing file, or a bucket URL
57-
that is a glob or trailing-slash folder. A plain file, URL, ``-`` (audio
58-
piped on stdin), or ``--sample`` stays on the single-source path.
58+
audio files), a glob pattern that names no existing file, a bucket URL that is
59+
a glob or trailing-slash folder, or an http(s) URL that turns out to be a
60+
podcast RSS/Atom feed (each episode becomes one batch source). A plain file,
61+
direct media URL, ``-`` (audio piped on stdin), or ``--sample`` stays on the
62+
single-source path. ``detect_feeds=False`` skips the feed probe (and its
63+
network fetch) for paths that must not touch the network, e.g. ``--show-code``.
5964
"""
6065
if from_stdin:
6166
return _stdin_sources(source, sample=sample)
6267
# `not source` (rather than `is None`) also catches the empty string — e.g. an
6368
# unset shell variable in `assembly transcribe "$FILE"`. `Path("")` is `Path(".")`,
6469
# so it would otherwise fall into the directory branch and batch-transcribe the
6570
# whole working directory; instead it stays single-source and fails validation.
66-
if not source or sample or source == "-" or source.startswith(URL_PREFIXES):
71+
if not source or sample or source == "-":
6772
return None
73+
if source.startswith(URL_PREFIXES):
74+
# A podcast feed URL expands into its episode enclosure URLs (batch mode);
75+
# a direct media URL or ordinary page returns None and stays single-source.
76+
from aai_cli.app.transcribe import feed
77+
78+
return feed.feed_episode_urls(source) if detect_feeds else None
6879
if remotefs.is_remote_url(source):
6980
return _remote_sources(source)
81+
return _local_sources(source)
82+
83+
84+
def _local_sources(source: str) -> list[str] | None:
85+
"""Batch sources for a local path: a directory's audio files or a glob's matches,
86+
else ``None`` (a single file, which the single-source path handles)."""
7087
path = Path(source)
7188
if path.is_dir():
7289
return _directory_sources(path)

‎aai_cli/commands/agent/_exec.py‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -22,7 +22,7 @@
2222
from aai_cli.agent.voices import VOICE_NAMES
2323
from aai_cli.app.agent_shared import resolve_system_prompt as _resolve_system_prompt
2424
from aai_cli.app.context import AppState
25-
from aai_cli.core import choices, client
25+
from aai_cli.core import choices, client, signals
2626
from aai_cli.core.errors import UsageError
2727
from aai_cli.streaming.session import resolve_output_modes
2828
from aai_cli.streaming.sources import FileSource
@@ -130,7 +130,10 @@ def run_agent(opts: AgentOptions, state: AppState, *, json_mode: bool) -> None:
130130
exit_after_reply=from_file,
131131
)
132132
try:
133-
run_session(api_key, renderer=renderer, player=player, mic=audio, config=run_config)
133+
# SIGTERM stops the agent as cleanly as Ctrl-C, so an external supervisor
134+
# (Hammerspoon, a service manager, a wrapper's `kill`) can end the session.
135+
with signals.terminate_as_interrupt():
136+
run_session(api_key, renderer=renderer, player=player, mic=audio, config=run_config)
134137
except KeyboardInterrupt:
135138
renderer.stopped()
136139
except BrokenPipeError as exc:

‎aai_cli/commands/agent_cascade/_exec.py‎

Lines changed: 5 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -22,7 +22,7 @@
2222
from aai_cli.agent_cascade.config import DEFAULT_MAX_HISTORY, CascadeConfig
2323
from aai_cli.app.agent_shared import resolve_system_prompt as _resolve_system_prompt
2424
from aai_cli.app.context import AppState
25-
from aai_cli.core import choices, client, config_builder, llm
25+
from aai_cli.core import choices, client, config_builder, llm, signals
2626
from aai_cli.core.errors import UsageError
2727
from aai_cli.streaming import turn_presets
2828
from aai_cli.streaming.session import resolve_output_modes
@@ -213,7 +213,10 @@ def run_agent_cascade(opts: AgentCascadeOptions, state: AppState, *, json_mode:
213213
stt_params = _build_stt_params(opts, sample_rate)
214214
deps = engine.CascadeDeps.real(api_key, config, audio=audio, stt_params=stt_params)
215215
try:
216-
engine.run_cascade(renderer=renderer, player=player, config=config, deps=deps)
216+
# SIGTERM stops the cascade as cleanly as Ctrl-C, so an external supervisor
217+
# (Hammerspoon, a service manager, a wrapper's `kill`) can end the session.
218+
with signals.terminate_as_interrupt():
219+
engine.run_cascade(renderer=renderer, player=player, config=config, deps=deps)
217220
except KeyboardInterrupt:
218221
renderer.stopped()
219222
except BrokenPipeError as exc:

‎aai_cli/commands/dictate/__init__.py‎

Lines changed: 8 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -24,6 +24,10 @@
2424
[
2525
("Dictate: Enter starts a recording, Enter transcribes it", "assembly dictate"),
2626
("One utterance, then exit", "assembly dictate --once"),
27+
(
28+
"Pipe one utterance into another command",
29+
'assembly dictate | assembly llm "write a conventional commit"',
30+
),
2731
("Dictate in Spanish", "assembly dictate --language es"),
2832
(
2933
"Bias recognition toward tricky terms",
@@ -51,7 +55,7 @@ def dictate(
5155
None, "--word-boost", help="Bias recognition toward a term (repeatable)"
5256
),
5357
device: int | None = typer.Option(None, "--device", help="Microphone device index"),
54-
once: bool = typer.Option(False, "--once", help="Transcribe one utterance, then exit"),
58+
once: bool = typer.Option(False, "--once", help="Record one utterance immediately, then exit"),
5559
max_seconds: float = typer.Option(
5660
float(MAX_AUDIO_SECONDS),
5761
"--max-seconds",
@@ -72,7 +76,9 @@ def dictate(
7276
Press Enter (or Space) to start recording and press it again to stop; the
7377
utterance is sent to the AssemblyAI Sync API and the transcript prints
7478
immediately — no polling. Press q (or Esc/Ctrl-C) to finish. Each utterance
75-
can be up to 120 seconds long.
79+
can be up to 120 seconds long. With --once, or when stdout is piped,
80+
recording starts immediately and dictate exits after one utterance so the
81+
transcript flows to the next command.
7682
"""
7783
opts = dictate_exec.DictateOptions(
7884
language=language,

0 commit comments

Comments
 (0)