Control Speaklone from your scripts, apps, and automations.
The API runs on localhost:7849 whenever Speaklone is open on your Mac. It supports both the native Speaklone API and OpenAI-compatible endpoints for drop-in integration with third-party tools.
What’s new in 1.3
audio/wav) for broader client compatibility — NDJSON and SSE remain available for event-metadata clients.opus output added to /v1/audio/speech, alongside wav and m4a.mp3, flac, aac) now return OpenAI-style invalid_request_error envelopes.Open Speaklone → Settings → scroll to Local API → click Copy next to the token.
a1b2c3d4e5f6a7b8...
Copy
Think of this as a password for the API. Keep it private.
Open Terminal and paste this (replace YOUR_TOKEN with the token you just copied):
curl -X POST http://localhost:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "Hello from the API!"}'
You should hear Speaklone speak on your Mac. The terminal will show:
{"status": "ok"}
How it works: The API triggers speech on your Mac's speakers. No audio data is returned over HTTP. Speaklone must be running for the API to respond.
OpenAI-compatible endpoints: Speaklone also exposes /v1/audio/speech (one-shot and streaming) and /v1/audio/transcriptions for drop-in use with tools like Open WebUI, OpenClaw, and the OpenAI SDK. See Client Presets below for setup guides.
Quick setup guides for popular tools that support OpenAI-compatible audio endpoints.
Use Speaklone as a TTS/STT provider in Open WebUI. Set both engines to OpenAI in Admin → Settings → Audio.
# TTS settings
TTS Engine: OpenAI
TTS Base URL: http://127.0.0.1:7849/v1
TTS API Key: YOUR_SPEAKLONE_TOKEN
TTS Model: speaklone
TTS Voice: aiden
# STT settings
STT Engine: OpenAI
STT Base URL: http://127.0.0.1:7849/v1
STT API Key: YOUR_SPEAKLONE_TOKEN
STT Model: speaklone-asr
Safari note: Safari blocks mixed-content requests (HTTPS → HTTP localhost). Use Chrome for local Open WebUI setups, or serve Open WebUI over HTTP.
/status
Health check. Returns Speaklone's version. The only public route — no authentication required.
// Response
{"status": "ok", "version": "1.3.0"}
/speak
Trigger text-to-speech. Audio plays on your Mac.
Requires Authorization: Bearer YOUR_TOKEN header.
Send as a JSON body with Content-Type: application/json.
| Parameter | Type | Required | Description |
|---|---|---|---|
text |
string | Yes | The text to speak. |
voice |
string | No | Voice name or ID (case-insensitive). Defaults to the last voice you selected in the app. See Voices. |
language |
string | No | Language code: en, zh, ja, ko, de, fr, ru, pt, es, it, or auto. Default: auto. |
direction |
string | No | Emotion or style instruction, e.g. "excited", "whispering". Only works with preset voices. |
seed |
number | No | Integer for deterministic output. Same seed + same text + same voice = same audio every time. |
wait |
boolean | No | If true, the response waits until playback finishes (up to 60s). Default: false. |
output |
string | No | speaker (default — plays on Mac, returns JSON), audio (returns audio bytes), or both (plays and returns bytes). |
format |
string | No | Audio container when output is audio or both. wav (default), m4a, or opus. |
stream |
boolean | No | Opt-in chunked streaming. Requires output: "audio" and format: "wav". Default: false. |
stream_format |
string | No | Stream transport when stream: true. Default "audio" (raw chunked WAV bytes). Also: "ndjson" or "sse" for event-metadata streams. |
Default (output: "speaker", wait: false) — returns immediately:
{"status": "ok"}
Blocking (wait: true) — returns after playback:
{
"status": "ok",
"waited": true,
"duration_ms": 3200,
"audio_samples": 76800,
"audio_sha256": "a1b2c3d4..."
}
output: "audio" — returns the audio bytes directly (Content-Type: audio/wav, audio/mp4, or audio/opus).
stream: true (requires output: "audio", format: "wav") — returns chunked WAV bytes. See streaming contract on /v1/audio/speech for the shared event format when stream_format is "ndjson" or "sse".
/speak preserves its legacy error shape for backwards compatibility:
// 400 - Missing or invalid parameters
{"error": "missing text parameter"}
{"error": "invalid seed parameter"}
{"error": "text too long"}
// 401 - Bad or missing token
{"error": "unauthorized"}
The OpenAI-compatible /v1/* routes use an { "error": { "message", "type", "param", "code" } } envelope instead. See below.
/voices
List all available voices — presets plus any you've created (cloned or designed).
Requires Authorization: Bearer YOUR_TOKEN header.
// Response (array)
[
{
"id": "aiden",
"name": "Aiden",
"type": "preset",
"gender": "male",
"description": "Sunny American male, clear midrange"
},
{
"id": "E3F1A2B4-...",
"name": "Spongebob",
"type": "cloned",
"gender": "male",
"description": "Custom voice from recording"
}
]
These endpoints follow the OpenAI API spec, so Speaklone works as a drop-in replacement in any tool that supports custom OpenAI-compatible audio endpoints. Errors return OpenAI-style error envelopes.
/v1/audio/speech
Text-to-speech. Returns audio data as a binary response (one-shot) or as a chunked stream (opt-in).
| Parameter | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Use speaklone (or any string — Speaklone ignores unknown model IDs). |
input |
string | Yes | The text to speak. |
voice |
string | No | Voice ID. Defaults to the app's current voice. |
response_format |
string | No | One-shot: wav (default), m4a, or opus. Stream mode: wav only. Unsupported values (mp3, flac, aac) return invalid_request_error. |
speed |
number | No | Playback rate multiplier. OpenAI-compatible. |
stream |
boolean | No | Opt-in chunked streaming. Default false. When true, response_format must be wav. |
stream_format |
string | No | Stream transport selector. Default "audio" → raw chunked WAV bytes. Other values: "ndjson", "sse". |
direction, language, seed |
string / string / number | No | Speaklone extensions. Same semantics as on /speak. |
stream: false)Returns the complete audio file in a single response. Supported response_format values: wav, m4a, opus. Any other value (including mp3, flac, aac) is rejected with an OpenAI-style invalid_request_error.
stream: true)stream_format |
Content-Type | Body |
|---|---|---|
"audio" (default) |
audio/wav |
Chunked WAV bytes: one progressive WAV header followed by PCM payload. The Content-Length is unknown until the stream ends — strict file validators should use one-shot mode instead. |
"ndjson" |
application/x-ndjson |
One JSON object per line. Event order: response.started → response.audio.delta* → response.completed. |
"sse" |
text/event-stream; charset=utf-8 |
Server-Sent Events carrying the same event sequence as NDJSON. |
NDJSON / SSE event payload shape:
{"type":"response.started","request_id":"...","sample_rate_hz":24000,"audio_encoding":"wav_base64"}
{"type":"response.audio.delta","index":0,"audio":"<base64 wav chunk>","sample_count":24000,"duration_ms":1000}
{"type":"response.completed","request_id":"...","chunks":8,"audio_samples":192000,"duration_ms":8000}
If streaming parameters are invalid (e.g. stream_format: "banana" or response_format: "m4a" with stream: true), Speaklone returns a standard HTTP error payload — not a stream.
{
"error": {
"message": "...",
"type": "invalid_request_error",
"param": "response_format",
"code": null
}
}
/v1/audio/transcriptions
Speech-to-text. Send an audio file and receive a transcription.
Multipart body only. 25 MB maximum.
| Parameter | Type | Required | Description |
|---|---|---|---|
file |
file | Yes | Audio file (multipart form upload). |
model |
string | Yes | Use speaklone-asr (or any string). |
language |
string | No | ISO language hint to bias recognition. |
response_format |
string | No | json, text, or verbose_json. Default: json. |
temperature |
number | No | Sampling temperature. OpenAI-compatible. |
Missing file, oversized uploads, and malformed multipart bodies return OpenAI-style invalid_request_error envelopes.
/v1/voices
Lists available voices in OpenAI-compatible format. Same data as /voices.
/v1/models
Lists available models.
// Response
{"data": [
{"id": "speaklone"},
{"id": "speaklone-asr"}
]}
| Code | Meaning |
|---|---|
200 | Success |
400 | Bad request — missing or invalid parameters |
401 | Unauthorized — missing or wrong Bearer token |
404 | Not found — endpoint doesn't exist |
500 | Server error — generation failed (blocking mode only) |
Auth scope: Speaklone's Local API is macOS-only. GET /status is the only public route — every other endpoint (legacy /speak, /voices, and all /v1/* routes) requires Authorization: Bearer <token>.
Network mode: When Local Network mode is enabled in Speaklone settings, the API is reachable from other devices on your LAN. The port can also be changed in settings.
Speaklone has three types of voices. Use any of them with the voice parameter.
9 built-in voices. These are the only voices that support the direction parameter for emotion control.
Voices you create from a recording. Use the voice name from Speaklone's sidebar. Does not support direction.
Voices you create by describing them in text (macOS only). Use the same way as cloned voices. Does not support direction.
Use the ID or Name as the voice parameter. Case-insensitive.
| ID | Name | Gender | Description |
|---|---|---|---|
aiden | Aiden | Male | Sunny American male, clear midrange |
ryan | Ryan | Male | Dynamic male, strong rhythmic drive |
vivian | Vivian | Female | Bright, slightly edgy young female |
serena | Serena | Female | Warm, gentle young female |
ono_anna | Ono Anna | Female | Playful female, light and nimble |
sohee | Sohee | Female | Warm female, rich emotion |
uncle_fu | Uncle Fu | Male | Seasoned male, low mellow timbre |
dylan | Dylan | Male | Youthful male, clear natural timbre |
eric | Eric | Male | Lively male, slightly husky |
Tip: Not sure what voices you have? Call GET /voices to see everything available on your Mac, including any cloned or designed voices.
# Say something with the default voice
curl -X POST http://localhost:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "Hello from the API!"}'
# Choose a voice and add emotion
curl -X POST http://localhost:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "This is incredible!", "voice": "vivian", "direction": "excited and happy"}'
# Use a cloned voice
curl -X POST http://localhost:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "Has anyone seen Patrick?", "voice": "Spongebob"}'
# Wait until playback finishes before continuing
curl -X POST http://localhost:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"text": "Waiting until I finish.", "wait": true}'
# List all your voices
curl http://localhost:7849/voices \
-H "Authorization: Bearer YOUR_TOKEN"
# Health check (no token needed)
curl http://localhost:7849/status
# One-shot TTS: get a WAV file
curl -X POST http://127.0.0.1:7849/v1/audio/speech \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "speaklone",
"voice": "aiden",
"input": "Hello from Speaklone.",
"response_format": "wav"
}' \
--output speech.wav
# Streaming TTS: default raw chunked WAV bytes (audio/wav)
curl -N -X POST http://127.0.0.1:7849/v1/audio/speech \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "speaklone",
"voice": "aiden",
"input": "Streaming hello from Speaklone.",
"response_format": "wav",
"stream": true
}' \
--output stream.wav
# Streaming TTS with event metadata (application/x-ndjson)
curl -N -X POST http://127.0.0.1:7849/v1/audio/speech \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "speaklone",
"input": "With metadata events.",
"response_format": "wav",
"stream": true,
"stream_format": "ndjson"
}'
# Speech-to-text (multipart upload)
curl -X POST http://127.0.0.1:7849/v1/audio/transcriptions \
-H "Authorization: Bearer YOUR_TOKEN" \
-F "file=@speech.wav" \
-F "model=speaklone-asr"
# /speak with the optional stream extension (raw chunked WAV)
curl -N -X POST http://127.0.0.1:7849/speak \
-H "Authorization: Bearer YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"text": "Legacy endpoint streaming bytes.",
"stream": true,
"output": "audio",
"format": "wav"
}' \
--output speak-stream.wav
/speak's default output: "speaker", audio plays on your Mac and the HTTP response is a JSON status — ideal for automations that just need voice output on your local machine. If you need the audio bytes, set output: "audio" (or "both") on /speak, or use the OpenAI-compatible /v1/audio/speech endpoint — which also supports opt-in chunked WAV streaming.
http://localhost:7849/speak and include your Bearer token as the Authorization header. The assistant sends text, and Speaklone speaks it on your Mac. Check your assistant's documentation for the specific configuration steps, or see the Client Presets section for ready-made configs.
wait: true and send requests one at a time.
direction is an emotion or style instruction for the voice. Examples: "excited", "whispering", "sad and slow", "speaking quickly with authority". It only works with the 9 preset voices — cloned and designed voices ignore it.
/speak returns immediately after queuing the text. With wait: true, the HTTP connection stays open until Speaklone finishes playing the audio (up to 60 seconds). The response includes timing details like duration_ms. This is useful for scripts that need to speak multiple lines in sequence.
127.0.0.1 (localhost). If you enable Local Network mode in Speaklone settings, the API becomes reachable from other devices on your LAN.
Authorization header with the correct Bearer token. The token is in Speaklone → Settings → Local API → Copy. It should look like Authorization: Bearer a1b2c3d4.... The /status endpoint doesn't need a token, so you can use that to check if the server is running.
GET /voices first.
/v1/audio/speech (TTS, one-shot and streaming), /v1/audio/transcriptions (STT), /v1/voices, and /v1/models, all following the OpenAI API spec with OpenAI-style error envelopes. This lets you use Speaklone as a drop-in replacement in any tool that supports custom OpenAI-compatible audio endpoints. See the Client Presets section for setup guides.
stream: false) supports wav, m4a, and opus. Streaming mode supports wav only. Requests for mp3, flac, or aac return a standard OpenAI-style invalid_request_error.
"stream": true with "response_format": "wav" to /v1/audio/speech. By default you get raw chunked WAV bytes (audio/wav) that most audio players and HTTP clients can consume on the fly. For event-driven clients, set stream_format to "ndjson" or "sse" to get response.started / response.audio.delta / response.completed events with base64-encoded WAV chunks. If a client disconnects mid-stream, Speaklone cancels generation server-side.