14 KiB
Migration du module voix : CosyVoice (.cvps) → OmniVoice (.ovsp)
Date : 23/06/2026 · Auteur : Richard & Claude (côté app/engine) · Pour : dev Kazeia-central
Objet : adapter kazeia_central/voice/ pour préparer les voix du nouveau TTS OmniVoice
(l'app patiente a retiré CosyVoice le 22/06/2026, vc17+). Les voix sont désormais des
clone-prompts .ovsp, plus des .cvps.
L'architecture du module ne change pas. Bridge worker en venv encodeur, orchestrator (pull → archive chiffré → transcription → enrôlement → push → suppression WAV device), verrou d'exclusivité (
locked_profile_id), autosync : tout reste valable. Seuls changent l'encodeur (enroll.py), la transcription (transcribe.py), les chemins/suffixes (transfer.py), le venv (bridge.py) et la glue (orchestrator.py).
0. Faits vérifiés sur device/PC (ne pas re-découvrir)
- Encodeur OmniVoice =
OmniVoice.create_voice_clone_prompt(ref_audio=<wav>, ref_text=None|str)→ renvoie unvcpavecvcp.ref_audio_tokens(tensor[8, T_ref]int) etvcp.ref_text(str). Siref_text=None→ ASR automatique (Whisper internetransformers),vcp.ref_textrempli. - venv encodeur =
ov_venv(/opt/Kazeia/ov_venv). Contient : torch, librosa, soundfile, transformers, omnivoice. PAS openai-whisper → la transcription (relecture opérateur) doit passer par l'ASR interne d'OmniVoice, pasimport whisper. - Cap de la réf à ~16 s (≠ CosyVoice qui voulait 22 s). OmniVoice paie la longueur de réf
au runtime :
T_ref ∝ durée, etc_len = T_ref + ciblechoisit le bucket de synthèse. Réf ~16 s →T_ref ~390→ bucket 640 (rapide, niveau damien validé 3,5 s). Réf 22 s →T_ref ~550→ bucket 768/1024 (plus lent) + risque de saturer le plus grand bucket. prepare_wavpartagé entre transcription et enrôlement (déterministe) → le texte relu correspond exactement à l'audio enrôlé. (sinon mismatch texte/audio = mauvais clone-prompt.)- Validé :
make_ovspclip 15 s →T_ref=364,.ovsp11,7 Ko ; auto-ASR sur clip 16 s tronqué OK. - Référence d'implémentation :
/opt/Kazeia-engine/dist/omnivoice/scripts/make_ovsp.py.
1. Format .ovsp (binaire, lu par le moteur tablette nativeLoadVoice)
int32 magic = 0x4F565350 ('OVSP', little-endian)
int32 T_ref (frames audio de réf, par codebook)
int32 text_len (octets UTF-8 du ref_text)
int32 ref_audio_tokens[8 * T_ref] (row-major [8, T_ref])
char ref_text[text_len] (UTF-8)
⚠ Le moteur ne tolère aucune dérive de ce format. Plus de conteneur GGUF (c'était CosyVoice).
2. Déploiement sur device
- Dossier cible :
<MODELS_DIR>/omnivoice/voices/<voice_id>.ovsp(était<MODELS_DIR>/cosyvoice/<id>.cvps). MODELS_DIRreste dérivé duwav_pathdu provider/voices(même logique qu'avant ; marqueur_WAV_MARKER = "/../voix/voix/"inchangé).- Le provider
/voicesde l'app patiente a déjà été migré : il inventorieomnivoice/voices/*.ovspet la suppression de voix retire le.ovsp. La colonne d'état « ready » =.ovspprésent.
3. Changements fichier par fichier
3.1 voice/enroll.py — réécriture complète (encodeur OmniVoice)
Remplacer intégralement par :
"""Enrôlement vocal OmniVoice : WAV de référence → clone-prompt `.ovsp`.
⚠️ Sous l'env encodeur `ov_venv` (torch + omnivoice + transformers + librosa/soundfile),
invoqué par le worker subprocess (pas en-process API py3.14).
Réimplémente /opt/Kazeia-engine/dist/omnivoice/scripts/make_ovsp.py.
Format .ovsp : int32 'OVSP' | int32 T_ref | int32 text_len | int32 tokens[8*T_ref] | char ref_text[].
"""
from __future__ import annotations
import logging, os, struct, sys
import numpy as np
for _noisy in ("numba", "matplotlib", "transformers", "omnivoice", "huggingface_hub"):
logging.getLogger(_noisy).setLevel(logging.WARNING)
OV_MODEL = os.environ.get("OV_MODEL", "k2-fsa/OmniVoice")
OVSP_MAGIC = 0x4F565350
# Cap ~16 s : T_ref ~390 -> bucket 640 (rapide). Au-delà = bucket 768/1024 (lent) +
# risque de saturer le plus grand bucket. prepare_wav partagé avec la transcription.
OV_TARGET_S = float(os.environ.get("OV_TARGET_S", "16.0"))
_PREP_SR = 24000
_HARD_MAX_S = 20.0
_MIN_WARN_S = 6.0
_model_singleton = None
class EnrollmentError(RuntimeError):
pass
def prepare_wav(in_path: str, out_path: str, target_s: float = OV_TARGET_S) -> dict:
import librosa, soundfile as sf
y, _ = librosa.load(in_path, sr=_PREP_SR, mono=True)
orig_s = len(y) / _PREP_SR
yt, _ = librosa.effects.trim(y, top_db=30)
if yt.size == 0:
yt = y
yt = yt[: int(min(target_s, _HARD_MAX_S) * _PREP_SR)]
final_s = len(yt) / _PREP_SR
sf.write(out_path, yt, _PREP_SR, subtype="PCM_16")
return {"orig_s": round(orig_s, 1), "final_s": round(final_s, 1), "too_short": final_s < _MIN_WARN_S}
def load_model():
global _model_singleton
if _model_singleton is not None:
return _model_singleton
try:
import torch
from omnivoice.models.omnivoice import OmniVoice
except Exception as e:
raise EnrollmentError(f"import OmniVoice échoué ({e}). Exécuter sous ov_venv.") from e
_model_singleton = OmniVoice.from_pretrained(OV_MODEL, torch_dtype=torch.float32).to("cpu").eval()
return _model_singleton
def _clone_prompt(prep_wav: str, ref_text):
import torch
m = load_model()
with torch.no_grad():
return m.create_voice_clone_prompt(ref_audio=prep_wav, ref_text=ref_text or None)
def transcribe_ref(wav_path: str) -> dict:
"""ASR du segment préparé (16 s) via OmniVoice, pour relecture opérateur."""
import tempfile
if not os.path.isfile(wav_path):
raise EnrollmentError(f"WAV introuvable: {wav_path}")
with tempfile.TemporaryDirectory() as td:
prep = os.path.join(td, "prep.wav")
prep_info = prepare_wav(wav_path, prep)
vcp = _clone_prompt(prep, None)
return {"text": (getattr(vcp, "ref_text", "") or "").strip(),
"language": None, "model": "omnivoice-asr", "prep": prep_info}
def enroll(wav_path: str, text: str, out_path: str) -> dict:
import tempfile
if not os.path.isfile(wav_path):
raise EnrollmentError(f"WAV introuvable: {wav_path}")
text = (text or "").strip()
with tempfile.TemporaryDirectory() as td:
prep = os.path.join(td, "prep.wav")
prep_info = prepare_wav(wav_path, prep)
vcp = _clone_prompt(prep, text if text else None)
toks = vcp.ref_audio_tokens.cpu().numpy().astype(np.int32) # [8, T_ref]
C, T = toks.shape
if C != 8:
raise EnrollmentError(f"attendu 8 codebooks, eu {C}")
ref_text = (getattr(vcp, "ref_text", "") or text or "").strip()
txt = ref_text.encode("utf-8")
os.makedirs(os.path.dirname(os.path.abspath(out_path)), exist_ok=True)
with open(out_path, "wb") as f:
f.write(struct.pack("<i", OVSP_MAGIC))
f.write(struct.pack("<i", T))
f.write(struct.pack("<i", len(txt)))
f.write(toks.tobytes())
f.write(txt)
return {"out_path": out_path, "bytes": os.path.getsize(out_path), "t_ref": int(T),
"ref_seconds": round(T / 25.0, 1), "text_bytes": len(txt),
"ref_text": ref_text, "prep": prep_info}
if __name__ == "__main__":
import json
ref, ref_txt, out = sys.argv[1], sys.argv[2], sys.argv[3]
text = "" if ref_txt == "-" else (open(ref_txt).read() if ref_txt.endswith(".txt") and os.path.isfile(ref_txt) else ref_txt)
print(json.dumps(enroll(ref, text, out), ensure_ascii=False))
3.2 voice/transcribe.py — ASR via OmniVoice (plus d'openai-whisper)
ov_venv n'a pas openai-whisper. Remplacer le corps par une délégation à
enroll.transcribe_ref (qui transcrit le segment préparé 16 s via l'ASR OmniVoice) :
"""Transcription du segment de référence (16 s) via l'ASR interne d'OmniVoice — ov_venv.
Le texte du .ovsp doit correspondre à l'audio réellement enrôlé ; prepare_wav étant
partagé/déterministe, le segment transcrit == le segment enrôlé."""
from __future__ import annotations
from .enroll import transcribe_ref
def transcribe(wav_path: str, model: str | None = None) -> dict:
return transcribe_ref(wav_path) # {text, language, model, prep}
3.3 voice/transfer.py — chemins/suffixe
- Renommer
cosyvoice_dir_for→omnivoice_dir_for, retournerposixpath.join(models, "omnivoice", "voices")(et le fallback…/kazeia/models/omnivoice/voices). - Renommer
deployed_cvps→deployed_ovsp, filtrer.ovsp(line[:-5]reste valable :.ovsp= 5 car). - Renommer
push_cvps→push_ovsp,remote = …/<voice_id>.ovsp. pull_wav/delete_device_wavinchangés.
3.4 voice/bridge.py — venv encodeur
CV_PYTHON→OV_PYTHON = os.environ.get("OV_PYTHON", "/opt/Kazeia/ov_venv/bin/python").- Adapter le message d'erreur (« provisionner ov_venv »). Le worker (
worker.py) est générique (dispatchping/transcribe/enroll/shutdown) → inchangé.
3.5 voice/orchestrator.py — glue
transfer.cosyvoice_dir_for→transfer.omnivoice_dir_for(var localecosy/ov).transfer.deployed_cvps→transfer.deployed_ovsp;transfer.push_cvps→transfer.push_ovsp.- Remplacer les littéraux
.cvps→.ovsp(chemins temporaires) et la clé de retourcvps_bytes→ovsp_bytes(et l'auditdetail={"ovsp_bytes": …}). - Le verrou d'exclusivité (
locked_profile_id/verrou_profil_absent) et l'archivage WAV chiffré : INCHANGÉS — la sémantique « ne pas déployer une voix exclusive sur une tablette sans le profil propriétaire » reste exactement la même.
3.6 voice/__init__.py, docstrings, docs/
__init__.py: docstring « enrôlement OmniVoice WAV → .ovsp ».- Mettre à jour
docs/VOICE_ENROLLMENT_SPEC.mdetdocs/VOICE_DEPLOYMENT_SPEC.md:.cvps→.ovsp, dossieromnivoice/voices, cap 16 s, ASR OmniVoice, format binaire OVSP. tests/test_voice.py: adapter les attentes (magicOVSP, suffixe.ovsp,omnivoice_dir_for,t_ref/ovsp_bytesau lieu des shapes GGUF/cvps_bytes).
4. Provisioning attendu (env)
ov_venvopérationnel (/opt/Kazeia/ov_venv) avec le modèlek2-fsa/OmniVoiceen cache HF (déjà le cas sur le PC de dev). Variables :OV_PYTHON,OV_MODEL,OV_TARGET_S(défaut 16).- Plus besoin de
CV_REPO/CV_TEACHER/cosyvoice-repo/teacher 0.5B.
5. Validation suggérée
ov_venv/bin/python -m kazeia_central.voice.enroll <wav> - /tmp/x.ovsp(ASR auto) → vérifier magicOVSP,t_ref~300–400 pour ~15 s,ref_textcohérent.- Flux complet sur une tablette :
prepare(transcription relue) →enroll_deploy→.ovspprésent dans…/omnivoice/voices/, voix « ready » dans l'admin, lecture OK. - Exclusivité : une voix
locked_profile_idnon présent sur la tablette →verrou_profil_absent.
6. Notes
- Le consentement légal complet (~22 s, texte intégral) reste dans le manifeste de
l'admin ; le
.ovspne porte que les 16 s + leur transcription (clone-prompt). C'est voulu. - Côté app, le défaut de débit voix est réglable (curseur admin) — sans impact sur l'enrôlement.
- Binaires/catalogue OmniVoice déjà publiés sur box (
catalog v13, composanttts_omnivoice).
7. ⚠ FIX OBLIGATOIRE — mot parasite en début de chaque phrase (bleed de fin de réf)
Symptôme observé en prod (voix « Richard » enrôlée via la lecture de consentement, 2026-06-23) : un mot étrange au début de CHAQUE phrase synthétisée.
Cause (diagnostiquée, preuve à l'appui) : le cap dur à 16 s coupe la réf en plein
mot. Ex. la phrase « …le chant joyeux des oiseaux, près d'un feu… » est coupée à 16 s ;
l'auto-ASR transcrit le bout partiel en « présenter » → ref_text se termine par ce
mot tronqué. OmniVoice recrache ce dernier mot de la réf au début de chaque génération.
Vérifié : synthèse de « Le chat dort sur le canapé. » → ASR = « Présenter le chat dort… ».
Correctif (validé device — le mot parasite disparaît) : ne JAMAIS terminer la réf sur un mot/segment tronqué.
- Couper l'audio sur une frontière de silence (jamais à
target_sbrut). Garder les segments non-silencieux jusqu'à ~target_s, finir à la fin du dernier segment complet :import librosa, soundfile as sf y,_ = librosa.load(in_path, sr=24000, mono=True) y,_ = librosa.effects.trim(y, top_db=30) end = 0 for s,e in librosa.effects.split(y, top_db=30): if e/24000 <= target_s: end = e else: break if end == 0: end = min(len(y), int(target_s*24000)) sf.write(out_path, y[:end], 24000, subtype="PCM_16") # finit en silence, mots entiers - Tronquer
ref_textà la dernière phrase complète (anti-bleed, ceinture+bretelles) :
(L'audio finit en silence → le léger surplus audio vs texte = silence, sans mot recraché.)import re ref = (vcp.ref_text or "").strip() last = list(re.finditer(r'[.!?]', ref)) if last: ref = ref[:last[-1].end()].strip() # finit sur . ! ?
→ À intégrer dans voice/enroll.py (prepare_wav pour le point 1, enroll pour le point 2).
Validé : ref_text « …le chant joyeux des oiseaux. » (au lieu de « …des oiseaux présenter. »)
→ plus de mot parasite. Toutes les voix déjà enrôlées via l'ancien cap doivent être
ré-enrôlées (la voix Richard de test a été re-patchée à la main en attendant).