feat(voice-bridge): Kokoro-FR fallback (TTS chain)
Adds Kokoro-82M FR (ff_siwis voice, mlx-audio runtime) as the second-tier TTS fallback after F5, before Piper, on the /voice/ws path. Reason: Piper Tower is English-only as of 2026-05 and butchers French utterances; Kokoro-FR is small (1.5 GB venv), fast on warm (~1.4 s for 5 words), and FR-native. New tools/macstudio/kokoro-fr/ houses the self-contained FastAPI server (POST /synthesize, drop-in field shape with Piper). Persisted on Studio :8002 via crontab @reboot. Verified roundtrip: Kokoro FR synth → Kyutai STT → 'USON est sensible, vous savez, ...' (clean). voice-bridge gains: - KOKORO_URL env (default http://localhost:8002, set to 0 to disable) - _kokoro_fallback() helper symmetric to _piper_fallback() - /voice/ws chain: cache -> F5 -> Kokoro -> Piper -> empty speak_end
This commit is contained in:
@@ -0,0 +1,114 @@
|
||||
# Kokoro-FR HTTP server — runbook
|
||||
|
||||
**Status** : LIVE on MacStudio :8002 since 2026-05-24. Drop-in
|
||||
third-tier TTS fallback for `voice-bridge` after F5 and Piper.
|
||||
|
||||
Why : Piper Tower (`192.168.0.120:8001`) is English-only as of 2026-05
|
||||
and produces wrong-language pronunciation on French Zacus utterances.
|
||||
Kokoro's `ff_siwis` is the only open-source FR voice that fits the
|
||||
"small + fast + acceptable quality" niche today (B- grade on the
|
||||
authors' scale, <11 h training data).
|
||||
|
||||
Validated 2026-05-24: a Kokoro-FR synthesis fed back into Kyutai STT
|
||||
roundtrips cleanly (`USON est sensible, vous savez, il réagit à
|
||||
certaines fréquences précises, cherchez bien autour de vous.`).
|
||||
|
||||
## Files on Studio
|
||||
|
||||
```
|
||||
/Users/clems/kokoro-fr/
|
||||
server.py # FastAPI app, /health + /synthesize
|
||||
requirements.txt
|
||||
.venv/ # uv-managed, ~1.5 GB (mlx-audio + misaki + spacy)
|
||||
logs/server.log
|
||||
```
|
||||
|
||||
The HF model (`prince-canuma/Kokoro-82M`) is auto-downloaded into
|
||||
`~/.cache/huggingface/` on first request (~150 MB).
|
||||
|
||||
## API
|
||||
|
||||
```
|
||||
GET /health → {status, model, default_voice, lang, sample_rate}
|
||||
POST /synthesize {"text": "...", "speaker_id": "ff_siwis"?}
|
||||
→ audio/wav PCM16 mono 24 kHz
|
||||
X-Kokoro-Latency-Ms + X-Kokoro-Voice headers
|
||||
```
|
||||
|
||||
Field name `speaker_id` is intentional Piper-compat — lets
|
||||
`voice-bridge._piper_fallback` shape be reused verbatim by
|
||||
`_kokoro_fallback`.
|
||||
|
||||
## Run
|
||||
|
||||
Persistent via crontab `@reboot` :
|
||||
|
||||
```cron
|
||||
@reboot cd /Users/clems/kokoro-fr && \
|
||||
/Users/clems/kokoro-fr/.venv/bin/python -m uvicorn server:app \
|
||||
--host 0.0.0.0 --port 8002 --app-dir /Users/clems/kokoro-fr \
|
||||
>> /Users/clems/kokoro-fr/logs/server.log 2>&1
|
||||
```
|
||||
|
||||
Manual restart :
|
||||
|
||||
```bash
|
||||
ssh electron-server "ssh [email protected] 'pkill -f \"uvicorn server:app\" ; \
|
||||
sleep 2 ; cd /Users/clems/kokoro-fr ; \
|
||||
nohup .venv/bin/python -m uvicorn server:app --host 0.0.0.0 --port 8002 \
|
||||
--app-dir /Users/clems/kokoro-fr </dev/null >> logs/server.log 2>&1 & disown'"
|
||||
```
|
||||
|
||||
## Smoke test
|
||||
|
||||
```bash
|
||||
curl -sS http://100.116.92.12:8002/health
|
||||
curl -sS -X POST http://100.116.92.12:8002/synthesize \
|
||||
-H 'Content-Type: application/json' \
|
||||
-d '{"text":"Bonjour, ceci est un test."}' \
|
||||
-o /tmp/test.wav
|
||||
file /tmp/test.wav # → WAVE audio, 16 bit, mono 24000 Hz
|
||||
afplay /tmp/test.wav # macOS playback
|
||||
```
|
||||
|
||||
## Latency
|
||||
|
||||
| Scenario | Wall time |
|
||||
|----------|-----------|
|
||||
| Cold (first call after boot, model load) | ~9 s |
|
||||
| Warm, short utterance (~5 words) | ~1.4 s |
|
||||
| Warm, medium utterance (~20 words) | ~3 s |
|
||||
|
||||
Far from F5 quality (no voice cloning) but **faster on warm path**
|
||||
and FR-native. Wired as the 2nd fallback in `voice-bridge` `/voice/ws`
|
||||
(after F5, before Piper) since the WS path is FR-dominant.
|
||||
|
||||
## Voice catalog
|
||||
|
||||
Only `ff_siwis` is FR. Other useful voices on this model :
|
||||
|
||||
| Voice | Lang | Notes |
|
||||
|-------|------|-------|
|
||||
| `ff_siwis` | FR | Female, SIWIS dataset, only FR voice |
|
||||
| `af_heart` | EN | Female US, default English voice |
|
||||
| `am_michael` | EN | Male US |
|
||||
|
||||
To change the default: `KOKORO_VOICE=<id>` env var before launching.
|
||||
|
||||
## Known limitations
|
||||
|
||||
- ~1.5 GB venv (mlx-audio + spacy + transformers + misaki). Future
|
||||
Kokoro releases may slim this down.
|
||||
- First call after boot is slow (~9 s) — model + tokenizer load. The
|
||||
`@reboot` crontab does NOT pre-warm; consider a curl probe in
|
||||
a post-boot script if cold-start ever matters for a demo.
|
||||
- Single voice for FR means we can't differentiate NPCs the way F5
|
||||
voice cloning would. Kokoro is for "anonymous system voice" /
|
||||
fallback, not for persona work.
|
||||
|
||||
## Roadmap
|
||||
|
||||
- Streaming first-chunk delivery (`mlx_audio` exposes `stream=True`
|
||||
but our HTTP wrapper still does end-to-end synth + return). Worth
|
||||
doing the day we want sub-500 ms first-byte from the fallback path.
|
||||
- Tune `cfg_scale` / `temperature` against playtest recordings.
|
||||
@@ -0,0 +1,10 @@
|
||||
fastapi>=0.115
|
||||
uvicorn[standard]>=0.30
|
||||
pydantic>=2
|
||||
numpy>=1.26
|
||||
soundfile>=0.12
|
||||
# Kokoro-82M MLX inference engine. Pulls in misaki (G2P), spacy, etc.
|
||||
# Total install ~1.5 GB once warm.
|
||||
mlx-audio>=0.4
|
||||
kokoro>=0.9
|
||||
# transformers version follows mlx-audio's own pin — let pip resolve.
|
||||
@@ -0,0 +1,129 @@
|
||||
"""Kokoro-82M FR HTTP TTS server — minimal /synthesize endpoint.
|
||||
|
||||
Designed as the second-tier TTS fallback in the voice-bridge chain:
|
||||
|
||||
cache → F5-TTS (primary, voice-cloned Zacus persona)
|
||||
→ Piper (Tower :8001, EN-only today)
|
||||
→ Kokoro-FR (this server, MacStudio :8002, fast FR neutral voice)
|
||||
|
||||
Surface mirrors the Tower Piper server so the voice-bridge
|
||||
``_piper_fallback`` helper can be reused with just a URL change:
|
||||
|
||||
POST /synthesize body: {"text": "...", "speaker_id": "ff_siwis"}
|
||||
returns: audio/wav PCM16 mono 24 kHz
|
||||
|
||||
Voice ``ff_siwis`` is Kokoro's only French voice as of 2026-05;
|
||||
B- grade on the Kokoro authors' scale (cf. VOICES.md). The
|
||||
underlying model is ``prince-canuma/Kokoro-82M`` which ships the
|
||||
MLX weights ; the alternative ``hexgrad/Kokoro-82M`` works through
|
||||
the regular kokoro PyPI package but is non-MLX.
|
||||
|
||||
Configuration:
|
||||
KOKORO_MODEL HuggingFace repo (default prince-canuma/Kokoro-82M)
|
||||
KOKORO_VOICE voice id (default ff_siwis)
|
||||
KOKORO_LANG language code (default 'f' for French)
|
||||
KOKORO_PORT listen port (default 8002)
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import io
|
||||
import logging
|
||||
import os
|
||||
import time
|
||||
import uuid
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import soundfile as sf
|
||||
from fastapi import FastAPI, HTTPException
|
||||
from fastapi.responses import Response
|
||||
from mlx_audio.tts.generate import generate_audio
|
||||
from pydantic import BaseModel, Field
|
||||
|
||||
|
||||
KOKORO_MODEL = os.getenv("KOKORO_MODEL", "prince-canuma/Kokoro-82M")
|
||||
KOKORO_VOICE_DEFAULT = os.getenv("KOKORO_VOICE", "ff_siwis")
|
||||
KOKORO_LANG = os.getenv("KOKORO_LANG", "f")
|
||||
KOKORO_SR = 24_000
|
||||
KOKORO_TMP = Path(os.getenv("KOKORO_TMP", "/tmp/kokoro-tts"))
|
||||
KOKORO_TMP.mkdir(parents=True, exist_ok=True)
|
||||
|
||||
logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(message)s")
|
||||
LOG = logging.getLogger("kokoro-fr")
|
||||
|
||||
app = FastAPI(title="kokoro-fr", version="0.1.0")
|
||||
|
||||
|
||||
class SynthesizeRequest(BaseModel):
|
||||
text: str = Field(min_length=1, max_length=2000)
|
||||
# speaker_id is the field name Piper Tower uses — kept for drop-in
|
||||
# compatibility with the voice-bridge _piper_fallback helper.
|
||||
speaker_id: str | None = None
|
||||
|
||||
|
||||
@app.get("/health")
|
||||
def health() -> dict:
|
||||
return {
|
||||
"status": "ok",
|
||||
"model": KOKORO_MODEL,
|
||||
"default_voice": KOKORO_VOICE_DEFAULT,
|
||||
"lang": KOKORO_LANG,
|
||||
"sample_rate": KOKORO_SR,
|
||||
}
|
||||
|
||||
|
||||
@app.post("/synthesize")
|
||||
def synthesize(req: SynthesizeRequest) -> Response:
|
||||
voice = req.speaker_id or KOKORO_VOICE_DEFAULT
|
||||
# generate_audio writes to disk and we read it back — not pretty,
|
||||
# but mlx-audio 0.4.3 has no in-memory API yet. Use a unique
|
||||
# prefix so concurrent requests don't race.
|
||||
rid = uuid.uuid4().hex[:12]
|
||||
prefix = KOKORO_TMP / f"kk_{rid}"
|
||||
t0 = time.monotonic()
|
||||
try:
|
||||
generate_audio(
|
||||
text=req.text,
|
||||
model=KOKORO_MODEL,
|
||||
voice=voice,
|
||||
lang_code=KOKORO_LANG,
|
||||
file_prefix=str(prefix),
|
||||
audio_format="wav",
|
||||
verbose=False,
|
||||
)
|
||||
except Exception as exc: # noqa: BLE001
|
||||
LOG.exception("generate_audio failed")
|
||||
raise HTTPException(status_code=500, detail=f"kokoro: {exc}") from exc
|
||||
latency_ms = int((time.monotonic() - t0) * 1000)
|
||||
|
||||
# generate_audio appends "_000.wav" to the prefix; for very long
|
||||
# inputs it may emit several files. We concatenate them in order.
|
||||
parts = sorted(KOKORO_TMP.glob(f"kk_{rid}_*.wav"))
|
||||
if not parts:
|
||||
raise HTTPException(status_code=500, detail="kokoro produced no audio")
|
||||
try:
|
||||
chunks = []
|
||||
sr_seen = None
|
||||
for p in parts:
|
||||
data, sr = sf.read(p, dtype="int16", always_2d=False)
|
||||
if sr_seen is None:
|
||||
sr_seen = sr
|
||||
chunks.append(data)
|
||||
merged = np.concatenate(chunks).astype(np.int16)
|
||||
buf = io.BytesIO()
|
||||
sf.write(buf, merged, sr_seen or KOKORO_SR, format="WAV", subtype="PCM_16")
|
||||
wav_bytes = buf.getvalue()
|
||||
finally:
|
||||
for p in parts:
|
||||
try:
|
||||
p.unlink()
|
||||
except OSError:
|
||||
pass
|
||||
|
||||
LOG.info("synth ok text=%d voice=%s latency_ms=%d bytes=%d",
|
||||
len(req.text), voice, latency_ms, len(wav_bytes))
|
||||
return Response(
|
||||
content=wav_bytes,
|
||||
media_type="audio/wav",
|
||||
headers={"X-Kokoro-Latency-Ms": str(latency_ms), "X-Kokoro-Voice": voice},
|
||||
)
|
||||
@@ -119,6 +119,12 @@ LITELLM_URL = os.getenv("LITELLM_URL", "http://localhost:4000")
|
||||
LITELLM_DEFAULT_KEY = "sk-zacus-local-dev-do-not-share" # placeholder, log warns at boot
|
||||
LITELLM_KEY = os.environ.get("LITELLM_MASTER_KEY", LITELLM_DEFAULT_KEY)
|
||||
PIPER_URL = os.getenv("PIPER_URL", "http://192.168.0.120:8001")
|
||||
# Kokoro-FR fallback (MacStudio :8002) — third-tier TTS, used when F5
|
||||
# fails AND Piper fails / produces wrong-language output. The /tts
|
||||
# HTTP handler still gives Piper first crack for backwards-compat, but
|
||||
# the /voice/ws path prefers Kokoro since Piper Tower is EN-only.
|
||||
# Set KOKORO_URL=0 to disable.
|
||||
KOKORO_URL = os.getenv("KOKORO_URL", "http://localhost:8002")
|
||||
F5_TIMEOUT_S = float(os.getenv("F5_TIMEOUT_S", "8.0"))
|
||||
F5_MODEL = os.getenv("F5_MODEL", "lucasnewman/f5-tts-mlx")
|
||||
F5_DEFAULT_STEPS = int(os.getenv("F5_DEFAULT_STEPS", "4"))
|
||||
@@ -831,6 +837,31 @@ async def _piper_fallback(text: str) -> Optional[bytes]:
|
||||
return None
|
||||
|
||||
|
||||
async def _kokoro_fallback(text: str) -> Optional[bytes]:
|
||||
"""POST text to Kokoro :8002/synthesize. Returns WAV bytes or None.
|
||||
|
||||
Third-tier TTS, used by /voice/ws when F5 and Piper both fail or
|
||||
when Kokoro is configured as the FR-preferred fallback (since
|
||||
Piper Tower is EN-only as of 2026-05). Returns None on any error
|
||||
so the caller can fall through to the next tier (or empty
|
||||
speak_end on the WS path).
|
||||
"""
|
||||
if KOKORO_URL in {"", "0", "off", "disabled"}:
|
||||
return None
|
||||
try:
|
||||
async with httpx.AsyncClient(timeout=15.0) as client:
|
||||
resp = await client.post(
|
||||
f"{KOKORO_URL}/synthesize",
|
||||
json={"text": text},
|
||||
)
|
||||
if resp.status_code == 200:
|
||||
return resp.content
|
||||
_jlog("kokoro_fallback_nonok", status=resp.status_code)
|
||||
except (httpx.ConnectError, httpx.TimeoutException, httpx.HTTPError) as exc:
|
||||
_jlog("kokoro_fallback_unreachable", err=type(exc).__name__)
|
||||
return None
|
||||
|
||||
|
||||
def _service_down_response(request_id: str, latency_ms: int,
|
||||
err_kind: Optional[str]) -> Optional[Response]:
|
||||
"""Return service_down.wav when both F5 and Piper are unavailable."""
|
||||
@@ -1839,11 +1870,26 @@ async def voice_ws(ws: WebSocket) -> None:
|
||||
_jlog("ws_tts_f5_not_loaded", request_id=request_id,
|
||||
load_err=_F5_LOAD_ERR)
|
||||
|
||||
# 5c. Piper fallback on F5 failure — mirrors the /tts handler.
|
||||
# Same wav-to-pcm conversion; we resample on output only if the
|
||||
# Piper sample-rate diverges from the WS contract (24 kHz). The
|
||||
# fallback is intentionally best-effort: any error here just
|
||||
# leaves pcm=None and the WS sends an empty speak_end below.
|
||||
# 5c. Fallback chain on F5 failure — mirrors the /tts handler
|
||||
# for the wav→pcm conversion, but ordered Kokoro-first then
|
||||
# Piper since Piper Tower is EN-only today and the WS path
|
||||
# is overwhelmingly FR. Each tier is best-effort; on any
|
||||
# error we fall through to the next one, and a final
|
||||
# pcm=None means the WS just emits an empty speak_end.
|
||||
if pcm is None and speak_text:
|
||||
wav_fallback = await _kokoro_fallback(speak_text)
|
||||
if wav_fallback is not None:
|
||||
try:
|
||||
pcm = _wav_to_pcm16(wav_fallback)
|
||||
tts_backend_used = "kokoro_fallback"
|
||||
tts_err = None
|
||||
_jlog("ws_tts_kokoro_fallback",
|
||||
request_id=request_id, bytes=len(pcm))
|
||||
except Exception as exc: # noqa: BLE001
|
||||
_jlog("ws_tts_kokoro_decode_err",
|
||||
request_id=request_id,
|
||||
err=type(exc).__name__, msg=str(exc))
|
||||
|
||||
if pcm is None and speak_text:
|
||||
wav_fallback = await _piper_fallback(speak_text)
|
||||
if wav_fallback is not None:
|
||||
|
||||
Reference in New Issue
Block a user