Adaptive planning for speech-driven text generation.
Compose emails, messages and other longer-form text by voice, hands-free and eyes-free, one question at a time.
This repository contains the research prototype for the UIST '25 paper StepWrite: Adaptive Planning for Speech-Driven Text Generation by Hamza El Alaoui, Atieh Taheri, Yi-Hao Peng and Jeffrey P. Bigham.
Paper (ACM DL) · arXiv · Project page · Citation
Dictation works well for short replies, but composing a longer, structured piece of text by voice while walking, commuting or otherwise on the move is hard: you have to plan the whole message in your head, hold it in memory while you speak, and you cannot look at a screen to revise it.
StepWrite is an LLM-driven, voice-based system that enables structured, hands-free and eyes-free composition of longer-form text. Instead of asking the user to dictate a finished message, it takes on the planning: it asks one short, context-aware question at a time, reads each question aloud, and listens for the spoken answer. Each new question is chosen based on everything said so far, so the conversation adapts to the task. When it has enough information, StepWrite writes the full text, checks it against what the user actually said, and reads it back. Users can go back and change an earlier answer at any point; StepWrite works out which later questions depend on that answer and discards only those, keeping the rest of the conversation.
- Choose a task. Start New writes a text from scratch. Reply composes a response to a
text you received (the prototype uses a sample email defined in
client/src/config.jsascore.reply_email). - Answer step-by-step questions. The server asks the model for the single best next question given the conversation so far, along with a decision on whether more information is needed. The client speaks the question with text-to-speech, detects speech in the browser with voice activity detection, filters noise and sends the audio for transcription.
- Background drafts. Once at least three questions are answered, a draft is generated in the background after each answer, and a Draft ready indicator appears. Saying "show draft" opens the latest draft in the editor.
- Final text with fact-checking. When no follow-up is needed (or the user says "finish writing"), the server classifies the intended tone, generates the text, and runs a fact-check loop that compares the output to the user's answers and corrects discrepancies (up to five attempts).
- Review and revise. The editor reads the text aloud when it opens and supports editing, undo/redo, font size, a one-sentence-per-line view and copy to clipboard. Saying "go back to questions" returns to the conversation, where any answer can be changed.
In hands-free mode everything can be done by voice. Phrases are matched fuzzily (configurable),
and each command has many variants in client/src/config.js
(handsFree.commands). Examples:
| Where | Command | Effect |
|---|---|---|
| Questions | "skip this question" | Skip the current question |
| Questions | "next question" / "previous question" | Move between answered questions |
| Questions | "change my answer" | Re-answer a question; only dependent later questions are discarded |
| Questions | "show draft" | Open the latest draft in the editor |
| Questions | "pause writing" / "continue writing" | Pause and resume listening |
| Questions | "finish writing" | Stop asking questions and generate the final text |
| Questions | "go to home page" | Return to the home screen |
| Editor | "go back to questions" | Return to the question flow |
| Editor | "read it again" / "stop reading" | Replay or stop the spoken text |
The interface is responsive and works in a phone browser (for example over an ngrok HTTPS tunnel).
Browser (React 18, Create React App, MUI)
├─ Voice activity detection: @ricky0123/vad-web + onnxruntime-web (loaded from jsDelivr in index.html)
├─ Pre-VAD noise filtering and silence checks (client/src/utils/audioUtils.js)
├─ Hands-free, text-only or text-and-voice input (REACT_APP_INPUT_MODE)
└─ HTTP → Node.js server on :3001
├─ POST /api/write next question, background draft, or final text
│ (Start New and Reply; Reply sends the received text as context)
├─ POST /api/reply, /api/edit standalone reply and edit endpoints
├─ POST /api/generate-initial-reply-question
├─ POST /api/classify/tone, /api/classify/text-type, /api/fact-check
├─ POST /api/draft/current, /api/control/followup
├─ POST /api/transcribe speech-to-text (OpenAI whisper-1)
├─ POST /api/tts/generate, GET /tts/* text-to-speech (OpenAI tts-1), cached on disk
└─ POST /api/save-experiment-data study logging (see below)
The server (server/server.js, server/openaiModule.js)
keeps no conversation state: the client sends the full conversation planning object (the list
of questions and answers plus a followup_needed flag) with every request.
All model behaviour is defined by the prompts in server/prompts/:
| Prompt | Purpose |
|---|---|
writePrompts.js |
Choose the next question for a new text, and write the final text |
replyPrompts.js |
Same for replies, grounded in the received message |
editPrompts.js |
Edit flow (currently disabled in the UI) |
toneClassificationPrompt.js |
Pick one of 14 tones (formal, urgent, empathetic, ...) to guide the final text |
factCheckingPrompt.js |
Verify the output against the user's answers and generate corrections |
dependencyAnalysisPrompt.js |
Decide which later questions are affected when an earlier answer changes |
textTypeClassificationPrompt.js |
Classify text as email, letter, message or other |
memoryPrompt.js |
Optional user profile context (disabled by default) |
- Node.js and npm (a current LTS release)
- An OpenAI API key
- A Chromium-based browser or Safari with microphone access. Browsers only allow microphone access
on
localhostor over HTTPS.
git clone https://github.com/helalaou/StepWrite.git
cd StepWrite
npm run install-all # installs the root, client and server dependenciescp server/.env.template server/.envThen set your key in server/.env (see Configuration).
npm startThis starts the server on http://localhost:3001 and the React app on
http://localhost:3000 together. Alternatively, start them separately with npm start in
server/ and in client/, or use ./run.sh.
Open http://localhost:3000, allow microphone access, and choose Start New or Reply.
server/.env (copied from server/.env.template) is loaded with dotenv
when the server starts from the server/ directory:
| Variable | Required | Description |
|---|---|---|
OPENAI_API_KEY |
Yes | Used for question generation, drafting, speech-to-text and text-to-speech |
The client reads these optional build-time variables (for example in client/.env):
| Variable | Default | Description |
|---|---|---|
REACT_APP_INPUT_MODE |
HANDS_FREE |
HANDS_FREE, TEXT_ONLY or TEXT_AND_VOICE |
REACT_APP_TTS_MODE |
ENABLED |
ENABLED or DISABLED spoken questions and output |
REACT_APP_TEXT_EDITOR_TTS_PREFIX |
The final output is: |
Phrase spoken before the final text is read aloud |
REACT_APP_COMMAND_MATCHING_MODE |
FUZZY |
Voice command matching: FUZZY, CONTAINS or EXACT_MATCH |
server/config.js holds the model, temperature and token limit for every
step (all default to chatgpt-4o-latest), the fact-checking loop, tone classification, the
continuous-draft policy, the TTS voice (nova) and the Whisper settings. Logs are written to
server/temp/app.log.
client/src/config.js holds the API URL and deployment mode, the reply
email used in the Reply flow, voice activity detection and noise thresholds, voice command
phrases, and the continuous-draft display settings. The continuous-draft settings exist in both
config files and should be kept in sync.
To use StepWrite from a phone, expose both ports over HTTPS. Set DEPLOY_MODE = 'ngrok' and the
NGROK.SERVER_URL / NGROK.CLIENT_URL values in client/src/config.js,
and start tunnels for ports 3000 and 3001 (see ngrok_config.yml for the
tunnel layout; use your own ngrok auth token and domain). The server accepts requests from ngrok
origins.
The prototype includes logging for running user studies. It is off by default.
- Set
experiment.enabled: trueinclient/src/config.js.experiment.enabledinserver/config.jsmust also betrue(the default). - With
participantIdRequired: true, each session starts by asking for a participant ID. - During the session the client tracks how often the participant skipped a question or modified an answer, and the time spent writing (answering questions) and revising (in the editor).
- In the editor, an end button (or saying "end writing" or "save and finish", detected with the browser's Web Speech API where available, e.g. in Chrome) saves the session.
Each session is saved on the server as server/temp/experiment_data/<participantId>_<write|reply>_experiment.txt
and contains the participant ID, skip and modify counts, writing and revision time, every
question and answer, the generated text, and the participant's revised text. The server/temp/
directory is git-ignored.
├── client/ React app (Create React App)
│ ├── public/ index.html (loads the VAD scripts), icons, logo, click sound
│ └── src/
│ ├── components/ LandingPage, HandsFreeInterface, ChatInterface, TextEditor,
│ │ ParticipantIdInput, ExperimentEndButton, ...
│ ├── hooks/ conversation, chat and editor state
│ ├── utils/ audio filtering and sound helpers
│ └── config.js client configuration
├── server/ Express API
│ ├── prompts/ all LLM prompts
│ ├── memory/ optional user profile context
│ ├── data/ user_memories.json (used only when memory is enabled)
│ ├── utils/ logging and temp-file cleanup
│ ├── openaiModule.js OpenAI calls, fact-checking and dependency analysis
│ ├── server.js routes
│ ├── config.js server configuration
│ └── .env.template
├── docs/assets/ screenshots
├── install.sh, run.sh, deploy.sh
└── ngrok_config.yml
If you use StepWrite in your research, please cite:
@inproceedings{elalaoui2025stepwrite,
author = {El Alaoui, Hamza and Taheri, Atieh and Peng, Yi-Hao and Bigham, Jeffrey P.},
title = {StepWrite: Adaptive Planning for Speech-Driven Text Generation},
booktitle = {Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology},
series = {UIST '25},
year = {2025},
publisher = {ACM},
doi = {10.1145/3746059.3747610},
url = {https://doi.org/10.1145/3746059.3747610}
}Hamza El Alaoui, Atieh Taheri, Yi-Hao Peng, and Jeffrey P. Bigham. 2025. StepWrite: Adaptive Planning for Speech-Driven Text Generation. In Proceedings of the 38th Annual ACM Symposium on User Interface Software and Technology (UIST '25). ACM. https://doi.org/10.1145/3746059.3747610
Citation metadata is also available in CITATION.cff.
Issues and pull requests are welcome. Read CONTRIBUTING.md and follow the Code of Conduct.
Hamza El Alaoui, Carnegie Mellon University · helalaou@cs.cmu.edu




