Skip to content

feat(mobile): add offline iPhone voice input - #233

Merged
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input
Sep 1, 2026
Merged

feat(mobile): add offline iPhone voice input#233
rynfar merged 5 commits into
pylonfrom
upstream/2026-09-01-voice-input

Conversation

@rynfar

@rynfar rynfar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

Fourth of the mobile batch. Adopted from upstream pingdotgg/t3code#8614 (352710d49). Follows #223, #224, #231.

Mobile gains on-device voice input: hold the mic in the composer, speak, and the transcript commits into the draft. Brings a voice-input controller to client-runtime, native transcription, the dictation UI, and two patched native deps.

The composer restructure, settled

This batch has cost a conflict in every cherry-pick because upstream's #8793 restructure arrives as context in each one. It turns out to be two independent changes, and treating them as one decision was the mistake:

  • ComposerSurface — animate borderRadius on a shared value, absolute glass layer, bounded collapsed radius. Touches nothing in the toolbar. Adopted (ceaae36aa); our file is now structurally identical to upstream there, so these hunks should stop conflicting.
  • The toolbar row — drop ComposerToolbarScroller for a fixed flex row. Declined, permanently. ComposerToolbarScroller is upstream's own component and they still ship it; they stopped using it because their toolbar holds four controls. Ours holds sixteen, because ControlPillMenu — which does not exist upstream at all — carries Refine, session goal, context window, agent count, the input queue, depth, resources, and reload.

Counts unchanged throughout: scroller 3, ControlPillMenu 13, QuickQuestionTrigger 2, ContextWindowIndicator 3.

Three defects I introduced, and fixed

The dictation controls had to be placed by hand rather than ported, since upstream places them inside the restructure we declined. Review found three ways that stranded the user. All 36 ported files were byte-identical to upstream — every defect was in the hand-written JSX, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen. NewTaskDraftScreen passes readOnly={voiceInput.freezesEditor}; the thread composer did not. The keyboard stayed live during recording, and one keystroke makes resolveTranscriptCommit see a changed draft and discard the whole transcript as stale — up to five minutes of speech, silently. It also made every native read-only guard this branch adds dead code on that surface.

Send vanished on any dictation error. Gated on isVoiceInputPresented rather than voicePresentation.showsSend. Those look interchangeable but diverge in exactly one phase: error has showsSend: true and a non-null statusLabel. So a denied permission, "no speech detected", or the stale-draft error above left no send control. Worse on entry, since dispose() no-ops in the error phase and a stale error survives navigating away and back.

Stop was unreachable for the whole dictation window. Placed inside ComposerToolbarScroller, which is the else branch of the dictation ternary — so across preparing → recording → transcribing → error there was no way to stop a running agent.

Also fixed: blocksSubmission guards on both submission entry points (canSend is derived above voiceInput and structurally cannot include it); a 4px spacer overflowing ComposerDictationToolbar's fixed 44px box and clipping the collapsed strip; the dropped pointerEvents="none" on the glass layer.

The showsSend divergence ships with a regression test — it is the one defect that is a pure predicate rather than JSX placement, and the one most likely to be reintroduced by someone reasoning that the two predicates look the same.

Pylon branding

Two leaks caught, both user-visible. app.config.ts carried upstream's "Allow T3 Code to use your microphone for voice input" — iOS renders that verbatim in the permission dialog, and the camera permission two lines below already said Pylon. Plus two T3 Code strings in docs/user/composer.md and docs/internals/voice-input.md.

No speech-recognition permission is needed, verified rather than assumed: @react-native-ai/apple uses SpeechAnalyzer/SpeechTranscriber, Apple's on-device framework, not SFSpeechRecognizer. That is what makes this offline.

Verification

Typecheck clean, lint clean, 1044 tests passing. Native build green with AppleLLM 0.12.0 and ExpoAudio.

This needs a native rebuild, not an OTA — two new native modules plus native Swift changes in t3-composer-editor/ios. runtimeVersion.policy is fingerprint, which covers new native deps and patches/, so a JS update cannot land on a stale binary.

Simulator pass verified the unavailable path only: with voice unavailable the composer renders correctly, full toolbar, no gap where the mic would be. The dictation UI itself has zero simulator coverageAppleTranscription.isAvailable() is a native check whose own error string reads "requires a supported device with iOS 26 or later", so the on-device model is absent on a simulator. Mic, status row, cancel, and the toolbar swap need a device pass before this merges.

Reviewed and integrated with Claude Opus 5 in Claude Code.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

NOT READY TO MERGE. Parked as a branch commit so the conflict resolution
is not lost.

Done: all five conflicts resolved, the two new patched native deps
(@react-native-ai/apple, expo-audio) installed, the voice-input feature
directories and client-runtime module landed intact, and ThreadComposer
compiles with the controller wired (composerOwnerKey, useVoiceInputController,
resolveVoiceComposerPresentation, showsCompactDictation).

Also resolved use-composer-command-menu.ts Pylon-first: upstream's file is far
smaller than Pylon's, so every Pylon export (buildComposerCommandItems,
resolveComposerProviderSlashCommands, and the ranking helpers) is kept and only
upstream's composerSelectionAtEnd helper and owner-key ref are added.

Not done, and the reason this is parked: the dictation UI is imported but not
rendered. ComposerDictationToolbar, ComposerDictationPrimaryAction,
ComposerDictationStatus and ComposerDictationCancelAction are all unused, so the
mic never appears. Wiring them means restructuring Pylon's composer rather than
patching it — upstream wraps its toolbar row directly, while Pylon's is
ComposerToolbarRow > ComposerToolbarScroller with thirteen controls. Pylon also
declares canSend roughly 750 lines above where voiceInput can exist, so even
`canSend && !voiceInput.blocksSubmission` needs the declaration order changed.

Shipping it compiling-but-inert would look done and do nothing, so it waits.
…lbar

#8793 changed two independent things in the mobile composer, and treating
them as one decision has cost a conflict in every cherry-pick from
upstream's 2026-08-30 batch since.

The first is ComposerSurface: animate borderRadius on a shared value, put
the glass on an absolute layer, render children in their own animated view,
and bound the collapsed pill radius so the morph interpolates instead of
travelling from 999. Nothing in it touches the toolbar row.

The second is the toolbar row: drop ComposerToolbarScroller for a fixed flex
row. That one genuinely conflicts. ComposerToolbarScroller is upstream's own
component and they still ship it; they stopped using it because their
toolbar holds four controls. Pylon's holds sixteen, because ControlPillMenu
- which does not exist upstream at all - carries Refine, session goal,
context window, agent count, the input queue, depth, resources, and reload.
Those need the scroller.

So take the first, decline the second. ComposerSurface is now structurally
identical to upstream (animatedBorderRadius, AnimatedGlassSurface,
layoutTransition, animatedShapeStyle, and the bounded radius all match), and
the toolbar is untouched: scroller, 13 ControlPillMenus, QuickQuestionTrigger
and ContextWindowIndicator all at their previous counts.

Pylon's shadow wrapper survives with its comment; upstream has no equivalent,
and the radius it carries is now animated alongside the surface.
Completes the #8614 port's UI half. Upstream places the dictation controls
inside its restructured collapsed row and fixed toolbar; Pylon declined that
restructure, so they are placed into Pylon's own structure instead.

The toolbar now shows whenever isToolbarVisible rather than only when
expanded, so dictation stays reachable from the collapsed pill, and it is
wrapped in ComposerDictationToolbar. The cancel action leads the row; while
dictating, ComposerDictationStatus replaces the toolbar scroller rather than
upstream's fixed left group, so Pylon's thirteen ControlPillMenu controls
keep their scroller when not dictating. The mic sits beside send in both the
collapsed row and the toolbar, and send is hidden while dictation owns the
row.

Scroller, ControlPillMenu, QuickQuestionTrigger and ContextWindowIndicator
are all at unchanged counts.
The #8614 port carried upstream's string verbatim: "Allow T3 Code to use
your microphone for voice input." iOS shows that text in the permission
dialog, so it is product copy, not a compatibility identifier. The camera
permission two lines below already reads "Allow Pylon to access your
camera", so this was purely adoption drift.

Verified while checking permissions that no speech-recognition key is
needed: @react-native-ai/apple uses SpeechAnalyzer and SpeechTranscriber,
Apple's on-device Speech framework, rather than SFSpeechRecognizer. Only
NSMicrophoneUsageDescription applies, and it is present.
Adversarial review of the hand-placed dictation UI found three ways the
composer strands the user. All three are in code written by hand rather
than ported, and none was caught by typecheck, lint, or 1042 tests.

The editor was never frozen. NewTaskDraftScreen passes
readOnly={voiceInput.freezesEditor}; the thread composer did not, so the
keyboard stayed live during recording. One keystroke makes
resolveTranscriptCommit see a changed draft and discard the entire
transcript as stale - up to five minutes of speech, silently. It also made
every native read-only guard this branch adds dead code on this surface.

Send was gated on isVoiceInputPresented rather than
voicePresentation.showsSend. Those look equivalent but diverge in exactly
one phase: error shows a status label AND keeps send. Any dictation failure
- denied permission, no speech detected, the stale-draft error above - left
the composer with no send control until the user found the dismiss button.
Worse on entry, since dispose() no-ops in the error phase, so a stale error
survives navigating away and back.

Stop sat inside ComposerToolbarScroller, which is the else branch of the
dictation ternary, so an agent was unstoppable for the whole recording and
transcription window. Moved to the always-rendered right cluster.

Also: guard both submission entry points on blocksSubmission, since canSend
is derived above voiceInput and cannot include it; move the 4px spacer
outside ComposerDictationToolbar's fixed 44px box, where it was overflowing
and clipping the collapsed dictation strip; restore pointerEvents="none" on
the glass layer with a comment matching the new sibling structure; and
rebrand two T3 Code strings in docs.

The showsSend divergence now has a regression test. It is the one defect
here that is a pure predicate rather than JSX placement, and it is the one
most likely to be reintroduced.
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:XXL labels Sep 1, 2026
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown

Thread transfer impact

✅ Thread transfer remains within every enforced ceiling.

Provider Metric Main baseline This PR Impact PR ceiling
Codex Total thread wire 13.7 KiB 13.7 KiB +8 B (+0.1%) 15.1 KiB
Codex Thread snapshot wire 6.9 KiB 6.9 KiB −3 B (−0.0%) 7.3 KiB
Codex Live turn WebSocket wire 6.8 KiB 6.8 KiB +11 B (+0.2%) 7.8 KiB
Codex Live turn WebSocket decoded 58.7 KiB 58.7 KiB 0 B (0.0%) 66.4 KiB
Codex Live turn messages 11 11 0 (0.0%) 21
Claude Total thread wire 13.5 KiB 13.8 KiB +230 B (+1.7%) 15.1 KiB
Claude Thread snapshot wire 6.9 KiB 6.9 KiB +2 B (+0.0%) 7.3 KiB
Claude Live turn WebSocket wire 6.6 KiB 6.8 KiB +228 B (+3.4%) 7.8 KiB
Claude Live turn WebSocket decoded 58.0 KiB 59.5 KiB +1.5 KiB (+2.6%) 66.4 KiB
Claude Live turn messages 9 11 +2 (+22.2%) 21

Baseline: c908bb2 · PR result: b21cb8f · Source CI: success

Scenario and decoded snapshot size

10 historical turns, 5 command tools per turn, 878.9 KiB retained MCP result per historical turn, and a 1.05 MiB retained result in the measured turn.

  • Codex decoded thread snapshot: 109.5 KiB
  • Claude decoded thread snapshot: 110.2 KiB

Updated in place by a trusted workflow. PR artifacts are strictly validated and never executed.

@rynfar
rynfar merged commit 74377fd into pylon Sep 1, 2026
19 checks passed
@rynfar
rynfar deleted the upstream/2026-09-01-voice-input branch September 1, 2026 23:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant