Kazeia-engine/dist/PERF_CPU_OPTIMS.md

143 lines
5.3 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Optimisations CPU TTS — bilan session (29/05)
Suite à l'élimination du HTP comme accélérateur (cf `PERF_INFRA_TESTS.md`),
session dédiée à pousser le CPU au maximum. **Résultat principal : RTF 2.93
→ 2.45 (-16% pipeline) via NEON dot-product sur les heads CP**, bit-exact.
## Mesure baseline (post-cleanup)
Pad3, KZTTS_THREADS=6, GGML_NUM_THREADS=8, KZTTS_CP_CACHE=1, seed=42 :
```
Phrase "Bonjour Kazeia" (33 frames, 2.75 s audio) :
pipeline total : 8.05 s
per-frame : talker 33 ms / cp 107 ms / decoder 100 ms
RTF : 2.93
```
## T1.bench — profile fin CP CPU pur
`KZTTS_CP_PROFILE=1` ajouté dans `cp_forward_cached_step`. Breakdown moyen
sur 30 sub-steps :
| Étape | ms/step | Total/frame | % |
|---|---:|---:|---:|
| `ggml_init` | 0.00 | 0 | 0% |
| `graph build + expand` | 0.22 | 3.3 | 3% |
| **`ggml_graph_compute_with_ctx`** | **4.49** | **67.4** | **63%** |
| `memcpy + free` | 0.00 | 0 | 0% |
| **Delta** (cp_predict_cached host code) | — | **~24** | **22%** |
| **Total mesuré** | | **107** | 100% |
**Découvertes invalidant des hypothèses précédentes** :
- `ggml_init / ggml_free` coûte 0 ms → l'angle « ctx persistant » est mort.
- Les copies sont gratuites → l'angle « pré-alloc buffers » est marginal.
- Le delta ~24 ms est entièrement le `sample_head` CPU host.
## T1.1 — thread sweep (sans affinity)
`KZTTS_THREADS={2,4,6,8} × GGML_NUM_THREADS={4,6,8}` :
| KZTTS_THREADS | GGML_NUM_THREADS | Total | RTF |
|---:|---:|---:|---:|
| 2 | 8 | 9.43 s | 3.43 |
| **4** | **8** | **8.21 s** | **2.99** |
| 6 | 8 | 8.09 s | 2.94 |
| 8 | * | 70-114 s | **40+** ⚠️ |
- `T=8` catastrophe : 8 threads ggml + main thread sur 8 cores → contention.
- `T=4-6` équivalents (CP BW-bound, pas compute-bound).
- `G=8` toujours mieux (decoder bénéficie de plus de threads).
- **Best : T=6 G=8** = RTF 2.94.
## T1.2 — affinity (`taskset`)
Topologie SM8750 : 6× Cortex-A720 @ 3.53 GHz (cpus 0-5) + 2× Cortex-X4 @
4.32 GHz (cpus 6-7). Tout perf, pas d'efficiency.
- `taskset 0xFF` (all) : RTF 2.88 — légèrement mieux que sans, marge bruit.
- Masks restrictifs (`0xFC`, `0xF8`) : process froze sur Android. Taskset
Android tatillon ; mécanisme non-fiable.
- **Skip**, garder par défaut tout-cores.
## T2.1 — NEON 16-way sur heads ✨ GROS GAIN
`cp_predict_cached.sample_head` faisait 2048 dot products de 1024 floats en
boucle SCALAIRE pour chacune des 15 heads/frame. Total 31.5 M ops/frame de
host-side code. Le compilateur n'auto-vectorise pas (probable à cause des
indices avec head_idx).
Implémentation explicite `dot_neon_1024` aarch64 :
```cpp
float32x4_t s0,s1,s2,s3 = vdupq_n_f32(0);
for (int i = 0; i < 1024; i += 16) {
s0 = vfmaq_f32(s0, vld1q_f32(a+i+0), vld1q_f32(b+i+0));
// ... 4 accumulators in parallel (instruction-level parallelism)
}
return vaddvq_f32(vaddq_f32(vaddq_f32(s0,s1), vaddq_f32(s2,s3)));
```
| Config | CP/frame | Pipeline | RTF | WAV md5 |
|---|---:|---:|---:|---|
| Scalar (KZTTS_CP_NO_NEON=1) | 106.8 ms | 8.05 s | 2.93 | c3dd71... |
| **NEON 16-way** | **72.1 ms** | **6.75 s** | **2.45** | **c3dd71... ✓** |
**-33% sur CP, -16% sur le pipeline, BIT-EXACT.** Commit `9a0c63f`.
Confirmation sur 3 phrases (33, 41, 60 frames) : RTF stable à 2.44-2.45.
## T2.2 — threadpool ggml persistent (skipped après analyse)
`ggml_graph_compute_with_ctx` crée un threadpool éphémère à chaque appel.
`ggml_threadpool_new` + `ggml_graph_plan` + `ggml_graph_compute` permet de
réutiliser un threadpool global.
**Estimation : ~0.5 ms / appel × 15 sub-forwards + 5 stages decoder = ~10 ms / phrase.**
Sur 6.75 s = **0.15% de gain**. Pas la peine de refactor 20 sites d'appel.
## État final post-optims
| Composant | Per-frame | % total | Plafond effort raisonnable |
|---|---:|---:|---|
| Talker | 32 ms | 15% | llama.cpp interne, peu de marge sans HTP |
| **CP** | **71 ms** | **35%** | NEON heads sorti, reste BW-bound ggml-cpu |
| **Decoder** | **99 ms** | **49%** | BigVGAN ggml-cpu NEON déjà, BW-bound |
| Prefill | 3 ms one-shot | 1% | text_projection NEON potentiel marginal |
| **Total** | **~203 ms** | 100% | **RTF 2.45** |
## Leviers résiduels et leur ROI
1. **Streaming pipeline** (1 session, neutre RTF) : TTFB perçu 6.7 s → ~500 ms.
**Le seul gain UX disponible.**
2. **Decoder : ggml_conv_1d → ggml_im2col + ggml_mul_mat explicite** (½ session)
pour permettre de batcher les im2col (réduire BW conv1d). Gain incertain,
probable 5-10% sur BigVGAN.
3. **text_projection NEON** (½ session) : ~50 ms one-shot au prefill → gain
marginal pipeline -0.5%.
4. **next_embed NEON + pre-alloc** : ~1 ms / frame → -0.5%.
Aucun ne dépasse 5% additionnel sans toucher au kernel hexagon ou au
backend ggml interne. **L'angle "optim CPU pure" a livré son maximum.**
## Pour RTF<1 sans HTP
Demande :
- Soit refactor decoder pour kernel HVX/HMX bit-exact (frontier R&D,
semaines, hors scope)
- Soit Vulkan backend Adreno pour le decoder (mal supporté ggml-vulkan
+ conv1d, semaines)
- Soit accepter audio dégradé pour gagner BW (quant Q4/Q8 CP/decoder,
test précis nécessaire)
Pour shipper en l'état : **streaming pipeline** rend RTF 2.45 acceptable
côté UX (audio commence à jouer ~500 ms après le clic).
## Commits associés
- `101d1cd` : tests infra HTP/HMX — verdict HTP non-rentable
- `9a0c63f` : T2.1 NEON heads CP — gain principal -33% CP -16% pipe
- Cette doc.