One fixed Kokoro voice keeps the tool honest and predictable; there is no uploaded reference voice or identity cloning.
Text to Speech
Turn writing into a local voice. Generate an English AI voice in your browser and download the result as MP3 or WAV.
What does free text to speech generate?
Text to speech converts written words into a generated voice waveform. SoundTools divides English text into manageable sentence chunks, runs a pinned Kokoro q8 model inside your browser, joins the resulting audio locally, and lets you preview or download an MP3 or WAV without sending the text to a speech API.
- AI text to speech
- text to voice AI free
- AI TTS generator
- text to speech MP3 download
Know the voice, text limit, and first-run cost
This focused generator uses one fixed English voice named Heart. Enter up to 1,000 characters; SoundTools splits longer passages into sentence-aware chunks of about 200 characters, synthesizes them in order, and joins the audio with short natural gaps. It does not offer a voice library, speed control, language selection, or voice cloning.
The pinned Kokoro 82M q8 model and WebAssembly runtime download only after you choose Generate. The first run therefore takes longer; a compatible browser can reuse cached assets later. Your text is passed to the model inside the tab, and the finished mono speech can be saved as WAV or locally encoded 192 kbps MP3.
Local text to speech vs cloud TTS vs voice cloning
Choose by privacy, voice range, and identity requirements rather than treating every speech generator as the same product.
| Option | Best for | What it uses | Important limit |
|---|---|---|---|
| SoundTools local TTS | Short private English drafts and MP3 downloads | One pinned Kokoro q8 voice running through browser WebAssembly | One voice, 1,000 characters, fixed speed, and a slower first model load |
| Cloud multi-voice TTS | Language, speaker, style, or API variety | Text is sent to a hosted synthesis service with a larger voice catalog | Privacy, quotas, accounts, and licensing depend on the provider |
| Voice cloning | Authorized reproduction of a specific speaker | A reference recording conditions or trains a speaker identity model | Requires explicit consent and is not part of this text-to-speech tool |
How to use free text to speech and download MP3
- 01
Write a short English script
Enter up to 1,000 characters. Add normal punctuation so the sentence chunker has clear places to pause, and review names or abbreviations that may be pronounced unexpectedly.
- 02
Choose MP3 or WAV and generate
Select MP3 for a compact listening file or WAV for uncompressed editing. On the first run, keep the tab open while the pinned model and runtime download and initialize.
- 03
Preview the complete narration
Listen for pronunciation and pacing, revise the text if needed, generate again, and download the finished local file in the selected format.
What this tool actually does
Clear limits are part of a useful tool. These values describe the processor currently running in this page.
Sentence-aware chunks of about 200 characters are synthesized sequentially and joined locally.
Generated mono samples can remain uncompressed or pass through the browser MP3 encoder before download.
Where one private English TTS voice is useful
- Scratch narration
Create a temporary voice track for a storyboard, video timing pass, slide deck, or edit before recording final narration.
- Accessibility listening copy
Turn a short English notice, instruction, or personal draft into a downloadable audio version.
- Pronunciation and copy review
Listen to sentence flow and catch awkward wording before a human read, while checking names and abbreviations manually.
Questions about this tool
Answers based on the current browser processor—not promises about a future version.
01Can I download free text to speech as MP3?
Yes. Choose MP3 before generation. SoundTools synthesizes the speech as local audio samples and then encodes the result at 192 kbps in the browser.
02Which text-to-speech voice does this tool use?
It uses the fixed English Heart voice from a pinned Kokoro 82M q8 model. The current interface does not provide additional speakers or emotional styles.
03What is the text limit?
Each generation accepts up to 1,000 English characters. The text is divided into sentence-aware chunks of about 200 characters and rejoined as one audio file.
04Why is the first generation slower?
The browser must download and initialize the pinned model and WebAssembly runtime. Later generations can be faster when those assets remain in the browser cache.
05Can this tool generate other languages or change speaking speed?
No. The current release supports one English voice at a fixed speed. It does not present unsupported language, accent, or pacing controls.
06Can I upload a voice to clone it?
No. This generator does not accept a reference recording or imitate a person’s identity. SoundTools has a separate consent-focused voice-cloning experiment for authorized samples.
07Is my text sent to a speech API?
No. The model assets are downloaded to the browser, but the entered text and generated audio stay in the current tab during synthesis and export.