A full-screen Markdown editor with voice dictation. Speak, pause, and the recognized text is appended as a new paragraph. The editor is a CodeMirror Markdown editor (list continuation on Enter, syntax highlighting, follows the browser light/dark theme).
This repo is only the frontend (deployed to GitHub Pages). It is a plain
OpenAI-compatible client and does no ML itself — it talks to a backend that
exposes /v1/audio/transcriptions (e.g. rapid-mlx serving NVIDIA Parakeet).
- Microphone →
AudioWorkletcaptures PCM16 @ 16 kHz (public/pcm-worklet.js). - A browser VAD (energy-based) segments speech; on a pause the utterance is
encoded to WAV and
POSTed to/v1/audio/transcriptions. - The returned text is appended to the document at the end (the caret is preserved, so you can keep editing elsewhere while dictating).
- Ctrl/Cmd+Enter commits the current utterance immediately.
- Stopping the recording auto-copies the whole document to the clipboard.
Endpoints are same-origin /v1/* by default, so a local proxy serves this page
and routes /v1 to the backend. A hosted service can be set in Settings
(Backend URL + API key).
npm install
npm run dev # http://localhost:5173The dev server proxies /v1 to a local backend by sub-path
(see vite.config.ts): /v1/audio → :8002, /v1/chat → :8000. Point those at
your local OpenAI-compatible servers (e.g. the local-llm rapid-mlx stack
running Parakeet on :8002).
The mic needs a secure context: http://localhost works directly, or serve over
HTTPS (e.g. tailscale serve) when opening from another device.
cd front
npm run build # static output in front/dist/- Backend: any OpenAI-compatible server exposing
/v1/audio/transcriptions(Parakeet viarapid-mlx, a Whisper server, etc.). - Local proxy (
scriber-local, separate repo): serves this frontend and routes/v1to the local backend so everything is same-origin on localhost.