Skip to content
Models, voice and media

Spoken replies

Set up on-device synthesis or a compatible speech server and test playback.

In this topic

Spoken replies turn an assistant message into audio. They use a speech engine independently of the model that wrote the answer. Choose an on-device voice for local synthesis, or connect a compatible speech server when you need its language or voice catalog.

Recognition and playback have separate models and permissions. A working microphone test does not confirm that a reply can be spoken.

Get started

Open Settings → Voice → Text To Speech and enable text-to-speech. With the on-device engine selected, download the PocketTTS Model and wait for preparation to finish. Choose a voice, use Preview → Play to check audible output, and then use a reply's speaker control in chat.

PocketTTS supports English speech. The default voice is Alba. Its temperature control changes variation in delivery. Test ordinary text first, without code blocks or long lists, so you can distinguish voice quality from text formatting.

Choosing where the voice comes from

EngineWhat Mellow needsWhere the text goes
On-device PocketTTSDownloaded, loadable voice assetsSynthesized on this Mac
OpenAI-compatible serverReachable endpoint, accepted model and voice, credentials if requiredSent to the configured speech server

Select the engine in Advanced. A server on your own computer and a hosted speech provider can implement the same protocol; their data handling is different. Use the actual endpoint's policy when deciding which text to send.

Using a speech server

Enter the server's base endpoint, its speech model identifier and its voice identifier. Supply an API key only if required by that server. Mellow's separate speech-key store uses Keychain; a language-model provider key does not automatically configure speech.

Use Test Connection to request a short synthesis sample and check that Mellow receives decodable audio. This check does not play the sample. Then use Preview → Play to verify audible output. A successful connection check validates that endpoint, credentials, model, voice and decoding work together at that moment; it does not guarantee every voice or every later request will work.

FieldHow to choose it
EndpointBase address of the server implementing the speech API
ModelExact identifier accepted by that server
VoiceExact voice identifier supported by the selected model
API keyCredential for this speech service, when required
SpeedAdjust within the offered range, then replay a sample

The initial configuration contains http://localhost:5050, tts-1 and alloy as server defaults. These values are not an installed service. Replace them with values from the server you actually operate or subscribe to.

Playback and agent speech

The reply speaker control reads a message. Agents can also use the speak capability when it is available to them. A long response may be synthesized in chunks, so an error can occur after playback has begun.

Keep the output device audible and check system volume. If you switch between headphones and speakers, use Preview to test playback again. Stopping a speech preview does not cancel a separate agent task or delete the message being read.

Model preparation and recovery

A download must produce loadable model assets before PocketTTS can run. A model-load or compiled-model error is different from a missing microphone permission.

  1. Read the error in the PocketTTS Model card.
  2. Use the offered download or repair action and let preparation complete.
  3. Retry a short preview before returning to chat.
  4. If loading still fails, record the app version and exact error for diagnosis.

Do not remove the entire Mellow profile to repair a speech model. That would affect unrelated agents and configuration without identifying the failing asset.

Under the hood

The server engine uses the /v1/audio/speech contract. Mellow's client handles supported audio encodings and reports decoding failures rather than treating every response body as playable audio. A server that returns an authentication error as HTML or JSON has not returned a voice sample, even if the network connection succeeded.

For a compatibility investigation, check request acceptance and audio decoding separately. Confirm the server's advertised output format and sample properties match the client path. Changing a model name will not repair an unsupported encoding.

Speech configuration stores the chosen engine and voice parameters separately from the server secret. For configuration and support exports, keep the credential out of copied settings and logs.

Troubleshooting

SymptomUseful next check
Speaker control leads to setupEngine enabled and required voice assets installed
PocketTTS cannot load a modelComplete or repair its assets; capture the model-load error
Server rejects the sampleEndpoint, model, voice and speech API credential
Sample accepted but no audioOutput device, volume and decoding error
Unsupported audio-format errorServer encoding and format compatibility
Voice sounds rushedSpeed setting, then replay the same sample
Recognition works but playback failsDiagnose the speech engine independently of Parakeet

Continue with Voice input and dictation to configure the other direction of the conversation.

Continue exploring · Models, voice and mediaImages and video →Choose media capabilities, manage local model preparation and inspect hosted jobs.