# Optimisations CPU TTS — bilan session (29/05) Suite à l'élimination du HTP comme accélérateur (cf `PERF_INFRA_TESTS.md`), session dédiée à pousser le CPU au maximum. **Résultat principal : RTF 2.93 → 2.45 (-16% pipeline) via NEON dot-product sur les heads CP**, bit-exact. ## Mesure baseline (post-cleanup) Pad3, KZTTS_THREADS=6, GGML_NUM_THREADS=8, KZTTS_CP_CACHE=1, seed=42 : ``` Phrase "Bonjour Kazeia" (33 frames, 2.75 s audio) : pipeline total : 8.05 s per-frame : talker 33 ms / cp 107 ms / decoder 100 ms RTF : 2.93 ``` ## T1.bench — profile fin CP CPU pur `KZTTS_CP_PROFILE=1` ajouté dans `cp_forward_cached_step`. Breakdown moyen sur 30 sub-steps : | Étape | ms/step | Total/frame | % | |---|---:|---:|---:| | `ggml_init` | 0.00 | 0 | 0% | | `graph build + expand` | 0.22 | 3.3 | 3% | | **`ggml_graph_compute_with_ctx`** | **4.49** | **67.4** | **63%** | | `memcpy + free` | 0.00 | 0 | 0% | | **Delta** (cp_predict_cached host code) | — | **~24** | **22%** | | **Total mesuré** | | **107** | 100% | **Découvertes invalidant des hypothèses précédentes** : - `ggml_init / ggml_free` coûte 0 ms → l'angle « ctx persistant » est mort. - Les copies sont gratuites → l'angle « pré-alloc buffers » est marginal. - Le delta ~24 ms est entièrement le `sample_head` CPU host. ## T1.1 — thread sweep (sans affinity) `KZTTS_THREADS={2,4,6,8} × GGML_NUM_THREADS={4,6,8}` : | KZTTS_THREADS | GGML_NUM_THREADS | Total | RTF | |---:|---:|---:|---:| | 2 | 8 | 9.43 s | 3.43 | | **4** | **8** | **8.21 s** | **2.99** | | 6 | 8 | 8.09 s | 2.94 | | 8 | * | 70-114 s | **40+** ⚠️ | - `T=8` catastrophe : 8 threads ggml + main thread sur 8 cores → contention. - `T=4-6` équivalents (CP BW-bound, pas compute-bound). - `G=8` toujours mieux (decoder bénéficie de plus de threads). - **Best : T=6 G=8** = RTF 2.94. ## T1.2 — affinity (`taskset`) Topologie SM8750 : 6× Cortex-A720 @ 3.53 GHz (cpus 0-5) + 2× Cortex-X4 @ 4.32 GHz (cpus 6-7). Tout perf, pas d'efficiency. - `taskset 0xFF` (all) : RTF 2.88 — légèrement mieux que sans, marge bruit. - Masks restrictifs (`0xFC`, `0xF8`) : process froze sur Android. Taskset Android tatillon ; mécanisme non-fiable. - **Skip**, garder par défaut tout-cores. ## T2.1 — NEON 16-way sur heads ✨ GROS GAIN `cp_predict_cached.sample_head` faisait 2048 dot products de 1024 floats en boucle SCALAIRE pour chacune des 15 heads/frame. Total 31.5 M ops/frame de host-side code. Le compilateur n'auto-vectorise pas (probable à cause des indices avec head_idx). Implémentation explicite `dot_neon_1024` aarch64 : ```cpp float32x4_t s0,s1,s2,s3 = vdupq_n_f32(0); for (int i = 0; i < 1024; i += 16) { s0 = vfmaq_f32(s0, vld1q_f32(a+i+0), vld1q_f32(b+i+0)); // ... 4 accumulators in parallel (instruction-level parallelism) } return vaddvq_f32(vaddq_f32(vaddq_f32(s0,s1), vaddq_f32(s2,s3))); ``` | Config | CP/frame | Pipeline | RTF | WAV md5 | |---|---:|---:|---:|---| | Scalar (KZTTS_CP_NO_NEON=1) | 106.8 ms | 8.05 s | 2.93 | c3dd71... | | **NEON 16-way** | **72.1 ms** | **6.75 s** | **2.45** | **c3dd71... ✓** | **-33% sur CP, -16% sur le pipeline, BIT-EXACT.** Commit `9a0c63f`. Confirmation sur 3 phrases (33, 41, 60 frames) : RTF stable à 2.44-2.45. ## T2.2 — threadpool ggml persistent (skipped après analyse) `ggml_graph_compute_with_ctx` crée un threadpool éphémère à chaque appel. `ggml_threadpool_new` + `ggml_graph_plan` + `ggml_graph_compute` permet de réutiliser un threadpool global. **Estimation : ~0.5 ms / appel × 15 sub-forwards + 5 stages decoder = ~10 ms / phrase.** Sur 6.75 s = **0.15% de gain**. Pas la peine de refactor 20 sites d'appel. ## État final post-optims | Composant | Per-frame | % total | Plafond effort raisonnable | |---|---:|---:|---| | Talker | 32 ms | 15% | llama.cpp interne, peu de marge sans HTP | | **CP** | **71 ms** | **35%** | NEON heads sorti, reste BW-bound ggml-cpu | | **Decoder** | **99 ms** | **49%** | BigVGAN ggml-cpu NEON déjà, BW-bound | | Prefill | 3 ms one-shot | 1% | text_projection NEON potentiel marginal | | **Total** | **~203 ms** | 100% | **RTF 2.45** | ## Leviers résiduels et leur ROI 1. **Streaming pipeline** (1 session, neutre RTF) : TTFB perçu 6.7 s → ~500 ms. **Le seul gain UX disponible.** 2. **Decoder : ggml_conv_1d → ggml_im2col + ggml_mul_mat explicite** (½ session) pour permettre de batcher les im2col (réduire BW conv1d). Gain incertain, probable 5-10% sur BigVGAN. 3. **text_projection NEON** (½ session) : ~50 ms one-shot au prefill → gain marginal pipeline -0.5%. 4. **next_embed NEON + pre-alloc** : ~1 ms / frame → -0.5%. Aucun ne dépasse 5% additionnel sans toucher au kernel hexagon ou au backend ggml interne. **L'angle "optim CPU pure" a livré son maximum.** ## Pour RTF<1 sans HTP Demande : - Soit refactor decoder pour kernel HVX/HMX bit-exact (frontier R&D, semaines, hors scope) - Soit Vulkan backend Adreno pour le decoder (mal supporté ggml-vulkan + conv1d, semaines) - Soit accepter audio dégradé pour gagner BW (quant Q4/Q8 CP/decoder, test précis nécessaire) Pour shipper en l'état : **streaming pipeline** rend RTF 2.45 acceptable côté UX (audio commence à jouer ~500 ms après le clic). ## Commits associés - `101d1cd` : tests infra HTP/HMX — verdict HTP non-rentable - `9a0c63f` : T2.1 NEON heads CP — gain principal -33% CP -16% pipe - Cette doc.