Speaklone Speaklone API Docs

Local API

Control Speaklone from your scripts, apps, and automations.

The API runs on localhost:7849 whenever Speaklone is open on your Mac. It supports both the native Speaklone API and OpenAI-compatible endpoints for drop-in integration with third-party tools.

macOS only · $39.99 · No subscription

What’s new in 1.3

  • OpenAI-compatible audio routes expanded and hardened.
  • Streaming now defaults to raw chunked WAV bytes (audio/wav) for broader client compatibility — NDJSON and SSE remain available for event-metadata clients.
  • One-shot opus output added to /v1/audio/speech, alongside wav and m4a.
  • Invalid formats (mp3, flac, aac) now return OpenAI-style invalid_request_error envelopes.

Quick Start

1

Find your API token

Open Speaklone → Settings → scroll to Local API → click Copy next to the token.

Token a1b2c3d4e5f6a7b8... Copy

Think of this as a password for the API. Keep it private.

2

Make your first request

Open Terminal and paste this (replace YOUR_TOKEN with the token you just copied):

curl -X POST http://localhost:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello from the API!"}'
3

Hear the result

You should hear Speaklone speak on your Mac. The terminal will show:

{"status": "ok"}

How it works: The API triggers speech on your Mac's speakers. No audio data is returned over HTTP. Speaklone must be running for the API to respond.

OpenAI-compatible endpoints: Speaklone also exposes /v1/audio/speech (one-shot and streaming) and /v1/audio/transcriptions for drop-in use with tools like Open WebUI, OpenClaw, and the OpenAI SDK. See Client Presets below for setup guides.

Client Presets

Quick setup guides for popular tools that support OpenAI-compatible audio endpoints.

Open WebUI

Use Speaklone as a TTS/STT provider in Open WebUI. Set both engines to OpenAI in Admin → Settings → Audio.

# TTS settings
TTS Engine:          OpenAI
TTS Base URL:        http://127.0.0.1:7849/v1
TTS API Key:         YOUR_SPEAKLONE_TOKEN
TTS Model:           speaklone
TTS Voice:           aiden

# STT settings
STT Engine:          OpenAI
STT Base URL:        http://127.0.0.1:7849/v1
STT API Key:         YOUR_SPEAKLONE_TOKEN
STT Model:           speaklone-asr

Safari note: Safari blocks mixed-content requests (HTTPS → HTTP localhost). Use Chrome for local Open WebUI setups, or serve Open WebUI over HTTP.

Endpoints

Native API

GET /status

Health check. Returns Speaklone's version. The only public route — no authentication required.

// Response
{"status": "ok", "version": "1.3.0"}
POST /speak

Trigger text-to-speech. Audio plays on your Mac.

Requires Authorization: Bearer YOUR_TOKEN header.

Parameters

Send as a JSON body with Content-Type: application/json.

Parameter Type Required Description
text string Yes The text to speak.
voice string No Voice name or ID (case-insensitive). Defaults to the last voice you selected in the app. See Voices.
language string No Language code: en, zh, ja, ko, de, fr, ru, pt, es, it, or auto. Default: auto.
direction string No Emotion or style instruction, e.g. "excited", "whispering". Only works with preset voices.
seed number No Integer for deterministic output. Same seed + same text + same voice = same audio every time.
wait boolean No If true, the response waits until playback finishes (up to 60s). Default: false.
output string No speaker (default — plays on Mac, returns JSON), audio (returns audio bytes), or both (plays and returns bytes).
format string No Audio container when output is audio or both. wav (default), m4a, or opus.
stream boolean No Opt-in chunked streaming. Requires output: "audio" and format: "wav". Default: false.
stream_format string No Stream transport when stream: true. Default "audio" (raw chunked WAV bytes). Also: "ndjson" or "sse" for event-metadata streams.

Response

Default (output: "speaker", wait: false) — returns immediately:

{"status": "ok"}

Blocking (wait: true) — returns after playback:

{
  "status": "ok",
  "waited": true,
  "duration_ms": 3200,
  "audio_samples": 76800,
  "audio_sha256": "a1b2c3d4..."
}

output: "audio" — returns the audio bytes directly (Content-Type: audio/wav, audio/mp4, or audio/opus).

stream: true (requires output: "audio", format: "wav") — returns chunked WAV bytes. See streaming contract on /v1/audio/speech for the shared event format when stream_format is "ndjson" or "sse".

Errors

/speak preserves its legacy error shape for backwards compatibility:

// 400 - Missing or invalid parameters
{"error": "missing text parameter"}
{"error": "invalid seed parameter"}
{"error": "text too long"}

// 401 - Bad or missing token
{"error": "unauthorized"}

The OpenAI-compatible /v1/* routes use an { "error": { "message", "type", "param", "code" } } envelope instead. See below.

GET /voices

List all available voices — presets plus any you've created (cloned or designed).

Requires Authorization: Bearer YOUR_TOKEN header.

// Response (array)
[
  {
    "id": "aiden",
    "name": "Aiden",
    "type": "preset",
    "gender": "male",
    "description": "Sunny American male, clear midrange"
  },
  {
    "id": "E3F1A2B4-...",
    "name": "Spongebob",
    "type": "cloned",
    "gender": "male",
    "description": "Custom voice from recording"
  }
]

OpenAI-Compatible

These endpoints follow the OpenAI API spec, so Speaklone works as a drop-in replacement in any tool that supports custom OpenAI-compatible audio endpoints. Errors return OpenAI-style error envelopes.

POST /v1/audio/speech

Text-to-speech. Returns audio data as a binary response (one-shot) or as a chunked stream (opt-in).

Parameter Type Required Description
model string Yes Use speaklone (or any string — Speaklone ignores unknown model IDs).
input string Yes The text to speak.
voice string No Voice ID. Defaults to the app's current voice.
response_format string No One-shot: wav (default), m4a, or opus. Stream mode: wav only. Unsupported values (mp3, flac, aac) return invalid_request_error.
speed number No Playback rate multiplier. OpenAI-compatible.
stream boolean No Opt-in chunked streaming. Default false. When true, response_format must be wav.
stream_format string No Stream transport selector. Default "audio" → raw chunked WAV bytes. Other values: "ndjson", "sse".
direction, language, seed string / string / number No Speaklone extensions. Same semantics as on /speak.

One-shot output (stream: false)

Returns the complete audio file in a single response. Supported response_format values: wav, m4a, opus. Any other value (including mp3, flac, aac) is rejected with an OpenAI-style invalid_request_error.

Streaming contract (stream: true)

stream_format Content-Type Body
"audio" (default) audio/wav Chunked WAV bytes: one progressive WAV header followed by PCM payload. The Content-Length is unknown until the stream ends — strict file validators should use one-shot mode instead.
"ndjson" application/x-ndjson One JSON object per line. Event order: response.started → response.audio.delta* → response.completed.
"sse" text/event-stream; charset=utf-8 Server-Sent Events carrying the same event sequence as NDJSON.

NDJSON / SSE event payload shape:

{"type":"response.started","request_id":"...","sample_rate_hz":24000,"audio_encoding":"wav_base64"}
{"type":"response.audio.delta","index":0,"audio":"<base64 wav chunk>","sample_count":24000,"duration_ms":1000}
{"type":"response.completed","request_id":"...","chunks":8,"audio_samples":192000,"duration_ms":8000}

If streaming parameters are invalid (e.g. stream_format: "banana" or response_format: "m4a" with stream: true), Speaklone returns a standard HTTP error payload — not a stream.

Error envelope

{
  "error": {
    "message": "...",
    "type": "invalid_request_error",
    "param": "response_format",
    "code": null
  }
}
POST /v1/audio/transcriptions

Speech-to-text. Send an audio file and receive a transcription.

Multipart body only. 25 MB maximum.

Parameter Type Required Description
file file Yes Audio file (multipart form upload).
model string Yes Use speaklone-asr (or any string).
language string No ISO language hint to bias recognition.
response_format string No json, text, or verbose_json. Default: json.
temperature number No Sampling temperature. OpenAI-compatible.

Missing file, oversized uploads, and malformed multipart bodies return OpenAI-style invalid_request_error envelopes.

GET /v1/voices

Lists available voices in OpenAI-compatible format. Same data as /voices.

GET /v1/models

Lists available models.

// Response
{"data": [
  {"id": "speaklone"},
  {"id": "speaklone-asr"}
]}

HTTP Status Codes

Code Meaning
200Success
400Bad request — missing or invalid parameters
401Unauthorized — missing or wrong Bearer token
404Not found — endpoint doesn't exist
500Server error — generation failed (blocking mode only)

Auth scope: Speaklone's Local API is macOS-only. GET /status is the only public route — every other endpoint (legacy /speak, /voices, and all /v1/* routes) requires Authorization: Bearer <token>.

Network mode: When Local Network mode is enabled in Speaklone settings, the API is reachable from other devices on your LAN. The port can also be changed in settings.

Voices

Speaklone has three types of voices. Use any of them with the voice parameter.

Preset

9 built-in voices. These are the only voices that support the direction parameter for emotion control.

Cloned

Voices you create from a recording. Use the voice name from Speaklone's sidebar. Does not support direction.

Designed

Voices you create by describing them in text (macOS only). Use the same way as cloned voices. Does not support direction.

Preset Voices

Use the ID or Name as the voice parameter. Case-insensitive.

ID Name Gender Description
aidenAidenMaleSunny American male, clear midrange
ryanRyanMaleDynamic male, strong rhythmic drive
vivianVivianFemaleBright, slightly edgy young female
serenaSerenaFemaleWarm, gentle young female
ono_annaOno AnnaFemalePlayful female, light and nimble
soheeSoheeFemaleWarm female, rich emotion
uncle_fuUncle FuMaleSeasoned male, low mellow timbre
dylanDylanMaleYouthful male, clear natural timbre
ericEricMaleLively male, slightly husky

Tip: Not sure what voices you have? Call GET /voices to see everything available on your Mac, including any cloned or designed voices.

Code Examples

# Say something with the default voice
curl -X POST http://localhost:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "Hello from the API!"}'
# Choose a voice and add emotion
curl -X POST http://localhost:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "This is incredible!", "voice": "vivian", "direction": "excited and happy"}'
# Use a cloned voice
curl -X POST http://localhost:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "Has anyone seen Patrick?", "voice": "Spongebob"}'
# Wait until playback finishes before continuing
curl -X POST http://localhost:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "Waiting until I finish.", "wait": true}'
# List all your voices
curl http://localhost:7849/voices \
  -H "Authorization: Bearer YOUR_TOKEN"
# Health check (no token needed)
curl http://localhost:7849/status

OpenAI-compatible

# One-shot TTS: get a WAV file
curl -X POST http://127.0.0.1:7849/v1/audio/speech \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "speaklone",
    "voice": "aiden",
    "input": "Hello from Speaklone.",
    "response_format": "wav"
  }' \
  --output speech.wav
# Streaming TTS: default raw chunked WAV bytes (audio/wav)
curl -N -X POST http://127.0.0.1:7849/v1/audio/speech \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "speaklone",
    "voice": "aiden",
    "input": "Streaming hello from Speaklone.",
    "response_format": "wav",
    "stream": true
  }' \
  --output stream.wav
# Streaming TTS with event metadata (application/x-ndjson)
curl -N -X POST http://127.0.0.1:7849/v1/audio/speech \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "speaklone",
    "input": "With metadata events.",
    "response_format": "wav",
    "stream": true,
    "stream_format": "ndjson"
  }'
# Speech-to-text (multipart upload)
curl -X POST http://127.0.0.1:7849/v1/audio/transcriptions \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -F "file=@speech.wav" \
  -F "model=speaklone-asr"
# /speak with the optional stream extension (raw chunked WAV)
curl -N -X POST http://127.0.0.1:7849/speak \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Legacy endpoint streaming bytes.",
    "stream": true,
    "output": "audio",
    "format": "wav"
  }' \
  --output speak-stream.wav

Tips & FAQ

Audio plays on my Mac, not in the response — is that right?
By default, yes. With /speak's default output: "speaker", audio plays on your Mac and the HTTP response is a JSON status — ideal for automations that just need voice output on your local machine. If you need the audio bytes, set output: "audio" (or "both") on /speak, or use the OpenAI-compatible /v1/audio/speech endpoint — which also supports opt-in chunked WAV streaming.
How do I use this with Open Claw or other AI assistants?
Most AI assistant platforms let you configure a custom tool or HTTP action. Point it at http://localhost:7849/speak and include your Bearer token as the Authorization header. The assistant sends text, and Speaklone speaks it on your Mac. Check your assistant's documentation for the specific configuration steps, or see the Client Presets section for ready-made configs.
What happens if I send a request while it's already speaking?
The new request interrupts the current speech. Speaklone stops what it's saying and starts speaking the new text immediately. If you need to queue lines, use wait: true and send requests one at a time.
How does the "direction" parameter work?
direction is an emotion or style instruction for the voice. Examples: "excited", "whispering", "sad and slow", "speaking quickly with authority". It only works with the 9 preset voices — cloned and designed voices ignore it.
What does "wait: true" do exactly?
By default, /speak returns immediately after queuing the text. With wait: true, the HTTP connection stays open until Speaklone finishes playing the audio (up to 60 seconds). The response includes timing details like duration_ms. This is useful for scripts that need to speak multiple lines in sequence.
Can I call the API from a browser tab?
Yes. CORS is enabled, so JavaScript running in any browser tab can call the API. This works because the API only listens on localhost — your Bearer token is the access control.
Can I reach the API from another computer?
By default, the API only listens on 127.0.0.1 (localhost). If you enable Local Network mode in Speaklone settings, the API becomes reachable from other devices on your LAN.
I'm getting a 401 error — what's wrong?
Make sure you're including the Authorization header with the correct Bearer token. The token is in Speaklone → Settings → Local API → Copy. It should look like Authorization: Bearer a1b2c3d4.... The /status endpoint doesn't need a token, so you can use that to check if the server is running.
What if I send a voice name that doesn't exist?
Speaklone will fall back to the last voice you selected in the app. No error is returned. If you want to verify voice names, call GET /voices first.
What are the OpenAI-compatible endpoints?
Speaklone 1.3+ exposes /v1/audio/speech (TTS, one-shot and streaming), /v1/audio/transcriptions (STT), /v1/voices, and /v1/models, all following the OpenAI API spec with OpenAI-style error envelopes. This lets you use Speaklone as a drop-in replacement in any tool that supports custom OpenAI-compatible audio endpoints. See the Client Presets section for setup guides.
Which audio formats are supported?
One-shot TTS (stream: false) supports wav, m4a, and opus. Streaming mode supports wav only. Requests for mp3, flac, or aac return a standard OpenAI-style invalid_request_error.
How does streaming work?
Add "stream": true with "response_format": "wav" to /v1/audio/speech. By default you get raw chunked WAV bytes (audio/wav) that most audio players and HTTP clients can consume on the fly. For event-driven clients, set stream_format to "ndjson" or "sse" to get response.started / response.audio.delta / response.completed events with base64-encoded WAV chunks. If a client disconnects mid-stream, Speaklone cancels generation server-side.