diff --git a/dist/PERF_CPU_OPTIMS.md b/dist/PERF_CPU_OPTIMS.md new file mode 100644 index 0000000..aa16580 --- /dev/null +++ b/dist/PERF_CPU_OPTIMS.md @@ -0,0 +1,142 @@ +# Optimisations CPU TTS — bilan session (29/05) + +Suite à l'élimination du HTP comme accélérateur (cf `PERF_INFRA_TESTS.md`), +session dédiée à pousser le CPU au maximum. **Résultat principal : RTF 2.93 +→ 2.45 (-16% pipeline) via NEON dot-product sur les heads CP**, bit-exact. + +## Mesure baseline (post-cleanup) + +Pad3, KZTTS_THREADS=6, GGML_NUM_THREADS=8, KZTTS_CP_CACHE=1, seed=42 : + +``` +Phrase "Bonjour Kazeia" (33 frames, 2.75 s audio) : + pipeline total : 8.05 s + per-frame : talker 33 ms / cp 107 ms / decoder 100 ms + RTF : 2.93 +``` + +## T1.bench — profile fin CP CPU pur + +`KZTTS_CP_PROFILE=1` ajouté dans `cp_forward_cached_step`. Breakdown moyen +sur 30 sub-steps : + +| Étape | ms/step | Total/frame | % | +|---|---:|---:|---:| +| `ggml_init` | 0.00 | 0 | 0% | +| `graph build + expand` | 0.22 | 3.3 | 3% | +| **`ggml_graph_compute_with_ctx`** | **4.49** | **67.4** | **63%** | +| `memcpy + free` | 0.00 | 0 | 0% | +| **Delta** (cp_predict_cached host code) | — | **~24** | **22%** | +| **Total mesuré** | | **107** | 100% | + +**Découvertes invalidant des hypothèses précédentes** : +- `ggml_init / ggml_free` coûte 0 ms → l'angle « ctx persistant » est mort. +- Les copies sont gratuites → l'angle « pré-alloc buffers » est marginal. +- Le delta ~24 ms est entièrement le `sample_head` CPU host. + +## T1.1 — thread sweep (sans affinity) + +`KZTTS_THREADS={2,4,6,8} × GGML_NUM_THREADS={4,6,8}` : + +| KZTTS_THREADS | GGML_NUM_THREADS | Total | RTF | +|---:|---:|---:|---:| +| 2 | 8 | 9.43 s | 3.43 | +| **4** | **8** | **8.21 s** | **2.99** | +| 6 | 8 | 8.09 s | 2.94 | +| 8 | * | 70-114 s | **40+** ⚠️ | + +- `T=8` catastrophe : 8 threads ggml + main thread sur 8 cores → contention. +- `T=4-6` équivalents (CP BW-bound, pas compute-bound). +- `G=8` toujours mieux (decoder bénéficie de plus de threads). +- **Best : T=6 G=8** = RTF 2.94. + +## T1.2 — affinity (`taskset`) + +Topologie SM8750 : 6× Cortex-A720 @ 3.53 GHz (cpus 0-5) + 2× Cortex-X4 @ +4.32 GHz (cpus 6-7). Tout perf, pas d'efficiency. + +- `taskset 0xFF` (all) : RTF 2.88 — légèrement mieux que sans, marge bruit. +- Masks restrictifs (`0xFC`, `0xF8`) : process froze sur Android. Taskset + Android tatillon ; mécanisme non-fiable. +- **Skip**, garder par défaut tout-cores. + +## T2.1 — NEON 16-way sur heads ✨ GROS GAIN + +`cp_predict_cached.sample_head` faisait 2048 dot products de 1024 floats en +boucle SCALAIRE pour chacune des 15 heads/frame. Total 31.5 M ops/frame de +host-side code. Le compilateur n'auto-vectorise pas (probable à cause des +indices avec head_idx). + +Implémentation explicite `dot_neon_1024` aarch64 : +```cpp +float32x4_t s0,s1,s2,s3 = vdupq_n_f32(0); +for (int i = 0; i < 1024; i += 16) { + s0 = vfmaq_f32(s0, vld1q_f32(a+i+0), vld1q_f32(b+i+0)); + // ... 4 accumulators in parallel (instruction-level parallelism) +} +return vaddvq_f32(vaddq_f32(vaddq_f32(s0,s1), vaddq_f32(s2,s3))); +``` + +| Config | CP/frame | Pipeline | RTF | WAV md5 | +|---|---:|---:|---:|---| +| Scalar (KZTTS_CP_NO_NEON=1) | 106.8 ms | 8.05 s | 2.93 | c3dd71... | +| **NEON 16-way** | **72.1 ms** | **6.75 s** | **2.45** | **c3dd71... ✓** | + +**-33% sur CP, -16% sur le pipeline, BIT-EXACT.** Commit `9a0c63f`. + +Confirmation sur 3 phrases (33, 41, 60 frames) : RTF stable à 2.44-2.45. + +## T2.2 — threadpool ggml persistent (skipped après analyse) + +`ggml_graph_compute_with_ctx` crée un threadpool éphémère à chaque appel. +`ggml_threadpool_new` + `ggml_graph_plan` + `ggml_graph_compute` permet de +réutiliser un threadpool global. + +**Estimation : ~0.5 ms / appel × 15 sub-forwards + 5 stages decoder = ~10 ms / phrase.** +Sur 6.75 s = **0.15% de gain**. Pas la peine de refactor 20 sites d'appel. + +## État final post-optims + +| Composant | Per-frame | % total | Plafond effort raisonnable | +|---|---:|---:|---| +| Talker | 32 ms | 15% | llama.cpp interne, peu de marge sans HTP | +| **CP** | **71 ms** | **35%** | NEON heads sorti, reste BW-bound ggml-cpu | +| **Decoder** | **99 ms** | **49%** | BigVGAN ggml-cpu NEON déjà, BW-bound | +| Prefill | 3 ms one-shot | 1% | text_projection NEON potentiel marginal | +| **Total** | **~203 ms** | 100% | **RTF 2.45** | + +## Leviers résiduels et leur ROI + +1. **Streaming pipeline** (1 session, neutre RTF) : TTFB perçu 6.7 s → ~500 ms. + **Le seul gain UX disponible.** + +2. **Decoder : ggml_conv_1d → ggml_im2col + ggml_mul_mat explicite** (½ session) + pour permettre de batcher les im2col (réduire BW conv1d). Gain incertain, + probable 5-10% sur BigVGAN. + +3. **text_projection NEON** (½ session) : ~50 ms one-shot au prefill → gain + marginal pipeline -0.5%. + +4. **next_embed NEON + pre-alloc** : ~1 ms / frame → -0.5%. + +Aucun ne dépasse 5% additionnel sans toucher au kernel hexagon ou au +backend ggml interne. **L'angle "optim CPU pure" a livré son maximum.** + +## Pour RTF<1 sans HTP + +Demande : +- Soit refactor decoder pour kernel HVX/HMX bit-exact (frontier R&D, + semaines, hors scope) +- Soit Vulkan backend Adreno pour le decoder (mal supporté ggml-vulkan + + conv1d, semaines) +- Soit accepter audio dégradé pour gagner BW (quant Q4/Q8 CP/decoder, + test précis nécessaire) + +Pour shipper en l'état : **streaming pipeline** rend RTF 2.45 acceptable +côté UX (audio commence à jouer ~500 ms après le clic). + +## Commits associés + +- `101d1cd` : tests infra HTP/HMX — verdict HTP non-rentable +- `9a0c63f` : T2.1 NEON heads CP — gain principal -33% CP -16% pipe +- Cette doc.