Roadmap¶
Planned work on localvoxtral, then smaller issues that are ready for a PR.
Planned¶
- Hotword boosting in the speech model itself, to bias transcription toward your repo's vocabulary, not only polishing (#316, waiting on the eval corpus in #315)
- Claude Code session joins on more terminals: WezTerm is next, then tmux support
- Repo context for remote sessions: collect git status, diffs and recently touched files on the enrolled host over the app's own ssh, so polishing for a joined remote session gets the same context as a local one
- Documentation website: a visual end-user guide beyond these docs
- Revisable transcripts for the Overlay Buffer, so a streaming model can correct text it has already emitted (#383; the second streaming ASR model it asked for, NVIDIA Nemotron 3.5 ASR Streaming 0.6B, shipped in #463)
Specified and open¶
These are smaller than the planned items. Each issue already states its scope, the constraints this repo adds, and the proof a PR must carry.
- Duck other audio while dictating, and fade it back
- A dropped WebSocket ends the dictation instead of reconnecting
- Let the overlay be moved, and remember where
- The shortcut recorder refuses function keys that would work fine
Several of these came from people who forked the repo and solved the problem for themselves. The from-fork label tracks them. A PR that takes one up credits the original author.