feat(voice-bridge): Kokoro-FR fallback (TTS chain)

Adds Kokoro-82M FR (ff_siwis voice, mlx-audio runtime) as the
second-tier TTS fallback after F5, before Piper, on the /voice/ws
path. Reason: Piper Tower is English-only as of 2026-05 and butchers
French utterances; Kokoro-FR is small (1.5 GB venv), fast on warm
(~1.4 s for 5 words), and FR-native.

New tools/macstudio/kokoro-fr/ houses the self-contained FastAPI
server (POST /synthesize, drop-in field shape with Piper). Persisted
on Studio :8002 via crontab @reboot. Verified roundtrip: Kokoro FR
synth → Kyutai STT → 'USON est sensible, vous savez, ...' (clean).

voice-bridge gains:
  - KOKORO_URL env (default http://localhost:8002, set to 0 to disable)
  - _kokoro_fallback() helper symmetric to _piper_fallback()
  - /voice/ws chain: cache -> F5 -> Kokoro -> Piper -> empty speak_end
This commit is contained in:
L'électron rare
2026-05-24 01:45:54 +02:00
parent 6d5edc7025
commit 95e4bf6c21
4 changed files with 304 additions and 5 deletions
+114
View File
@@ -0,0 +1,114 @@
# Kokoro-FR HTTP server — runbook
**Status** : LIVE on MacStudio :8002 since 2026-05-24. Drop-in
third-tier TTS fallback for `voice-bridge` after F5 and Piper.
Why : Piper Tower (`192.168.0.120:8001`) is English-only as of 2026-05
and produces wrong-language pronunciation on French Zacus utterances.
Kokoro's `ff_siwis` is the only open-source FR voice that fits the
"small + fast + acceptable quality" niche today (B- grade on the
authors' scale, <11 h training data).
Validated 2026-05-24: a Kokoro-FR synthesis fed back into Kyutai STT
roundtrips cleanly (`USON est sensible, vous savez, il réagit à
certaines fréquences précises, cherchez bien autour de vous.`).
## Files on Studio
```
/Users/clems/kokoro-fr/
server.py # FastAPI app, /health + /synthesize
requirements.txt
.venv/ # uv-managed, ~1.5 GB (mlx-audio + misaki + spacy)
logs/server.log
```
The HF model (`prince-canuma/Kokoro-82M`) is auto-downloaded into
`~/.cache/huggingface/` on first request (~150 MB).
## API
```
GET /health → {status, model, default_voice, lang, sample_rate}
POST /synthesize {"text": "...", "speaker_id": "ff_siwis"?}
→ audio/wav PCM16 mono 24 kHz
X-Kokoro-Latency-Ms + X-Kokoro-Voice headers
```
Field name `speaker_id` is intentional Piper-compat — lets
`voice-bridge._piper_fallback` shape be reused verbatim by
`_kokoro_fallback`.
## Run
Persistent via crontab `@reboot` :
```cron
@reboot cd /Users/clems/kokoro-fr && \
/Users/clems/kokoro-fr/.venv/bin/python -m uvicorn server:app \
--host 0.0.0.0 --port 8002 --app-dir /Users/clems/kokoro-fr \
>> /Users/clems/kokoro-fr/logs/server.log 2>&1
```
Manual restart :
```bash
ssh electron-server "ssh [email protected] 'pkill -f \"uvicorn server:app\" ; \
sleep 2 ; cd /Users/clems/kokoro-fr ; \
nohup .venv/bin/python -m uvicorn server:app --host 0.0.0.0 --port 8002 \
--app-dir /Users/clems/kokoro-fr </dev/null >> logs/server.log 2>&1 & disown'"
```
## Smoke test
```bash
curl -sS http://100.116.92.12:8002/health
curl -sS -X POST http://100.116.92.12:8002/synthesize \
-H 'Content-Type: application/json' \
-d '{"text":"Bonjour, ceci est un test."}' \
-o /tmp/test.wav
file /tmp/test.wav # → WAVE audio, 16 bit, mono 24000 Hz
afplay /tmp/test.wav # macOS playback
```
## Latency
| Scenario | Wall time |
|----------|-----------|
| Cold (first call after boot, model load) | ~9 s |
| Warm, short utterance (~5 words) | ~1.4 s |
| Warm, medium utterance (~20 words) | ~3 s |
Far from F5 quality (no voice cloning) but **faster on warm path**
and FR-native. Wired as the 2nd fallback in `voice-bridge` `/voice/ws`
(after F5, before Piper) since the WS path is FR-dominant.
## Voice catalog
Only `ff_siwis` is FR. Other useful voices on this model :
| Voice | Lang | Notes |
|-------|------|-------|
| `ff_siwis` | FR | Female, SIWIS dataset, only FR voice |
| `af_heart` | EN | Female US, default English voice |
| `am_michael` | EN | Male US |
To change the default: `KOKORO_VOICE=<id>` env var before launching.
## Known limitations
- ~1.5 GB venv (mlx-audio + spacy + transformers + misaki). Future
Kokoro releases may slim this down.
- First call after boot is slow (~9 s) — model + tokenizer load. The
`@reboot` crontab does NOT pre-warm; consider a curl probe in
a post-boot script if cold-start ever matters for a demo.
- Single voice for FR means we can't differentiate NPCs the way F5
voice cloning would. Kokoro is for "anonymous system voice" /
fallback, not for persona work.
## Roadmap
- Streaming first-chunk delivery (`mlx_audio` exposes `stream=True`
but our HTTP wrapper still does end-to-end synth + return). Worth
doing the day we want sub-500 ms first-byte from the fallback path.
- Tune `cfg_scale` / `temperature` against playtest recordings.
@@ -0,0 +1,10 @@
fastapi>=0.115
uvicorn[standard]>=0.30
pydantic>=2
numpy>=1.26
soundfile>=0.12
# Kokoro-82M MLX inference engine. Pulls in misaki (G2P), spacy, etc.
# Total install ~1.5 GB once warm.
mlx-audio>=0.4
kokoro>=0.9
# transformers version follows mlx-audio's own pin — let pip resolve.
+129
View File
@@ -0,0 +1,129 @@
"""Kokoro-82M FR HTTP TTS server — minimal /synthesize endpoint.
Designed as the second-tier TTS fallback in the voice-bridge chain:
cache → F5-TTS (primary, voice-cloned Zacus persona)
→ Piper (Tower :8001, EN-only today)
→ Kokoro-FR (this server, MacStudio :8002, fast FR neutral voice)
Surface mirrors the Tower Piper server so the voice-bridge
``_piper_fallback`` helper can be reused with just a URL change:
POST /synthesize body: {"text": "...", "speaker_id": "ff_siwis"}
returns: audio/wav PCM16 mono 24 kHz
Voice ``ff_siwis`` is Kokoro's only French voice as of 2026-05;
B- grade on the Kokoro authors' scale (cf. VOICES.md). The
underlying model is ``prince-canuma/Kokoro-82M`` which ships the
MLX weights ; the alternative ``hexgrad/Kokoro-82M`` works through
the regular kokoro PyPI package but is non-MLX.
Configuration:
KOKORO_MODEL HuggingFace repo (default prince-canuma/Kokoro-82M)
KOKORO_VOICE voice id (default ff_siwis)
KOKORO_LANG language code (default 'f' for French)
KOKORO_PORT listen port (default 8002)
"""
from __future__ import annotations
import io
import logging
import os
import time
import uuid
from pathlib import Path
import numpy as np
import soundfile as sf
from fastapi import FastAPI, HTTPException
from fastapi.responses import Response
from mlx_audio.tts.generate import generate_audio
from pydantic import BaseModel, Field
KOKORO_MODEL = os.getenv("KOKORO_MODEL", "prince-canuma/Kokoro-82M")
KOKORO_VOICE_DEFAULT = os.getenv("KOKORO_VOICE", "ff_siwis")
KOKORO_LANG = os.getenv("KOKORO_LANG", "f")
KOKORO_SR = 24_000
KOKORO_TMP = Path(os.getenv("KOKORO_TMP", "/tmp/kokoro-tts"))
KOKORO_TMP.mkdir(parents=True, exist_ok=True)
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
LOG = logging.getLogger("kokoro-fr")
app = FastAPI(title="kokoro-fr", version="0.1.0")
class SynthesizeRequest(BaseModel):
text: str = Field(min_length=1, max_length=2000)
# speaker_id is the field name Piper Tower uses — kept for drop-in
# compatibility with the voice-bridge _piper_fallback helper.
speaker_id: str | None = None
@app.get("/health")
def health() -> dict:
return {
"status": "ok",
"model": KOKORO_MODEL,
"default_voice": KOKORO_VOICE_DEFAULT,
"lang": KOKORO_LANG,
"sample_rate": KOKORO_SR,
}
@app.post("/synthesize")
def synthesize(req: SynthesizeRequest) -> Response:
voice = req.speaker_id or KOKORO_VOICE_DEFAULT
# generate_audio writes to disk and we read it back — not pretty,
# but mlx-audio 0.4.3 has no in-memory API yet. Use a unique
# prefix so concurrent requests don't race.
rid = uuid.uuid4().hex[:12]
prefix = KOKORO_TMP / f"kk_{rid}"
t0 = time.monotonic()
try:
generate_audio(
text=req.text,
model=KOKORO_MODEL,
voice=voice,
lang_code=KOKORO_LANG,
file_prefix=str(prefix),
audio_format="wav",
verbose=False,
)
except Exception as exc: # noqa: BLE001
LOG.exception("generate_audio failed")
raise HTTPException(status_code=500, detail=f"kokoro: {exc}") from exc
latency_ms = int((time.monotonic() - t0) * 1000)
# generate_audio appends "_000.wav" to the prefix; for very long
# inputs it may emit several files. We concatenate them in order.
parts = sorted(KOKORO_TMP.glob(f"kk_{rid}_*.wav"))
if not parts:
raise HTTPException(status_code=500, detail="kokoro produced no audio")
try:
chunks = []
sr_seen = None
for p in parts:
data, sr = sf.read(p, dtype="int16", always_2d=False)
if sr_seen is None:
sr_seen = sr
chunks.append(data)
merged = np.concatenate(chunks).astype(np.int16)
buf = io.BytesIO()
sf.write(buf, merged, sr_seen or KOKORO_SR, format="WAV", subtype="PCM_16")
wav_bytes = buf.getvalue()
finally:
for p in parts:
try:
p.unlink()
except OSError:
pass
LOG.info("synth ok text=%d voice=%s latency_ms=%d bytes=%d",
len(req.text), voice, latency_ms, len(wav_bytes))
return Response(
content=wav_bytes,
media_type="audio/wav",
headers={"X-Kokoro-Latency-Ms": str(latency_ms), "X-Kokoro-Voice": voice},
)
+51 -5
View File
@@ -119,6 +119,12 @@ LITELLM_URL = os.getenv("LITELLM_URL", "http://localhost:4000")
LITELLM_DEFAULT_KEY = "sk-zacus-local-dev-do-not-share" # placeholder, log warns at boot
LITELLM_KEY = os.environ.get("LITELLM_MASTER_KEY", LITELLM_DEFAULT_KEY)
PIPER_URL = os.getenv("PIPER_URL", "http://192.168.0.120:8001")
# Kokoro-FR fallback (MacStudio :8002) — third-tier TTS, used when F5
# fails AND Piper fails / produces wrong-language output. The /tts
# HTTP handler still gives Piper first crack for backwards-compat, but
# the /voice/ws path prefers Kokoro since Piper Tower is EN-only.
# Set KOKORO_URL=0 to disable.
KOKORO_URL = os.getenv("KOKORO_URL", "http://localhost:8002")
F5_TIMEOUT_S = float(os.getenv("F5_TIMEOUT_S", "8.0"))
F5_MODEL = os.getenv("F5_MODEL", "lucasnewman/f5-tts-mlx")
F5_DEFAULT_STEPS = int(os.getenv("F5_DEFAULT_STEPS", "4"))
@@ -831,6 +837,31 @@ async def _piper_fallback(text: str) -> Optional[bytes]:
return None
async def _kokoro_fallback(text: str) -> Optional[bytes]:
"""POST text to Kokoro :8002/synthesize. Returns WAV bytes or None.
Third-tier TTS, used by /voice/ws when F5 and Piper both fail or
when Kokoro is configured as the FR-preferred fallback (since
Piper Tower is EN-only as of 2026-05). Returns None on any error
so the caller can fall through to the next tier (or empty
speak_end on the WS path).
"""
if KOKORO_URL in {"", "0", "off", "disabled"}:
return None
try:
async with httpx.AsyncClient(timeout=15.0) as client:
resp = await client.post(
f"{KOKORO_URL}/synthesize",
json={"text": text},
)
if resp.status_code == 200:
return resp.content
_jlog("kokoro_fallback_nonok", status=resp.status_code)
except (httpx.ConnectError, httpx.TimeoutException, httpx.HTTPError) as exc:
_jlog("kokoro_fallback_unreachable", err=type(exc).__name__)
return None
def _service_down_response(request_id: str, latency_ms: int,
err_kind: Optional[str]) -> Optional[Response]:
"""Return service_down.wav when both F5 and Piper are unavailable."""
@@ -1839,11 +1870,26 @@ async def voice_ws(ws: WebSocket) -> None:
_jlog("ws_tts_f5_not_loaded", request_id=request_id,
load_err=_F5_LOAD_ERR)
# 5c. Piper fallback on F5 failure — mirrors the /tts handler.
# Same wav-to-pcm conversion; we resample on output only if the
# Piper sample-rate diverges from the WS contract (24 kHz). The
# fallback is intentionally best-effort: any error here just
# leaves pcm=None and the WS sends an empty speak_end below.
# 5c. Fallback chain on F5 failure — mirrors the /tts handler
# for the wav→pcm conversion, but ordered Kokoro-first then
# Piper since Piper Tower is EN-only today and the WS path
# is overwhelmingly FR. Each tier is best-effort; on any
# error we fall through to the next one, and a final
# pcm=None means the WS just emits an empty speak_end.
if pcm is None and speak_text:
wav_fallback = await _kokoro_fallback(speak_text)
if wav_fallback is not None:
try:
pcm = _wav_to_pcm16(wav_fallback)
tts_backend_used = "kokoro_fallback"
tts_err = None
_jlog("ws_tts_kokoro_fallback",
request_id=request_id, bytes=len(pcm))
except Exception as exc: # noqa: BLE001
_jlog("ws_tts_kokoro_decode_err",
request_id=request_id,
err=type(exc).__name__, msg=str(exc))
if pcm is None and speak_text:
wav_fallback = await _piper_fallback(speak_text)
if wav_fallback is not None: