Muse Voice is an online speech-to-text workspace that takes a recording from intake to a reviewed handoff in the browser. Creators, journalists, students, and teams use Muse Voice to transcribe meetings, interviews, podcasts, lectures, and video files, then correct the draft before they publish captions or share notes.
Upload common audio formats such as MP3, WAV, M4A, and FLAC or video formats such as MP4 and MOV, up to 1 GB per file. Record live in the page when there is no file yet, or paste a hosted media URL so a large clip does not have to be downloaded and uploaded again. Muse Voice speech to text can auto-detect the spoken language or use a language you choose, covering about 100 languages. Quality still depends on audio, accents, and noise, so every automatic result is treated as a draft.
Key Features
- Three ways to start: file upload, in-browser recording, or a hosted media URL
- Optional speaker separation with rename-once labels that apply across matching turns
- Word-level timestamps you can click to replay the matching moment
- An editor for search, name fixes, punctuation, and speaker cleanup
- Six exports from the same approved transcript: TXT, DOCX, PDF, SRT, VTT, and JSON
- Account-tied private storage until you delete a recording
Use Cases Meeting owners turn calls into decision logs. Interviewers keep quotations with speaker names. Podcasters draft show notes. Students capture lectures. Video publishers generate SRT or VTT captions. Accessibility work uses the same reviewed text as captions or a reading copy.
Pricing New accounts can verify and receive 5 transcription minutes to test the full workflow. Paid plans sell transcription minutes as monthly or yearly allowances, and one-time credit packs cover occasional backlogs. The same balance also covers available AI tools on the site. Current prices and minute amounts are listed at https://musevoice.pro/pricing.
Muse Voice is a practical browser workflow around speech to text, not Meta's official live model demo. The landing page documents Meta Muse Voice Transcribe research separately from what the workspace can do today.

