5.3 KiB
Optimisations CPU TTS — bilan session (29/05)
Suite à l'élimination du HTP comme accélérateur (cf PERF_INFRA_TESTS.md),
session dédiée à pousser le CPU au maximum. Résultat principal : RTF 2.93
→ 2.45 (-16% pipeline) via NEON dot-product sur les heads CP, bit-exact.
Mesure baseline (post-cleanup)
Pad3, KZTTS_THREADS=6, GGML_NUM_THREADS=8, KZTTS_CP_CACHE=1, seed=42 :
Phrase "Bonjour Kazeia" (33 frames, 2.75 s audio) :
pipeline total : 8.05 s
per-frame : talker 33 ms / cp 107 ms / decoder 100 ms
RTF : 2.93
T1.bench — profile fin CP CPU pur
KZTTS_CP_PROFILE=1 ajouté dans cp_forward_cached_step. Breakdown moyen
sur 30 sub-steps :
| Étape | ms/step | Total/frame | % |
|---|---|---|---|
ggml_init |
0.00 | 0 | 0% |
graph build + expand |
0.22 | 3.3 | 3% |
ggml_graph_compute_with_ctx |
4.49 | 67.4 | 63% |
memcpy + free |
0.00 | 0 | 0% |
| Delta (cp_predict_cached host code) | — | ~24 | 22% |
| Total mesuré | 107 | 100% |
Découvertes invalidant des hypothèses précédentes :
ggml_init / ggml_freecoûte 0 ms → l'angle « ctx persistant » est mort.- Les copies sont gratuites → l'angle « pré-alloc buffers » est marginal.
- Le delta ~24 ms est entièrement le
sample_headCPU host.
T1.1 — thread sweep (sans affinity)
KZTTS_THREADS={2,4,6,8} × GGML_NUM_THREADS={4,6,8} :
| KZTTS_THREADS | GGML_NUM_THREADS | Total | RTF |
|---|---|---|---|
| 2 | 8 | 9.43 s | 3.43 |
| 4 | 8 | 8.21 s | 2.99 |
| 6 | 8 | 8.09 s | 2.94 |
| 8 | * | 70-114 s | 40+ ⚠️ |
T=8catastrophe : 8 threads ggml + main thread sur 8 cores → contention.T=4-6équivalents (CP BW-bound, pas compute-bound).G=8toujours mieux (decoder bénéficie de plus de threads).- Best : T=6 G=8 = RTF 2.94.
T1.2 — affinity (taskset)
Topologie SM8750 : 6× Cortex-A720 @ 3.53 GHz (cpus 0-5) + 2× Cortex-X4 @ 4.32 GHz (cpus 6-7). Tout perf, pas d'efficiency.
taskset 0xFF(all) : RTF 2.88 — légèrement mieux que sans, marge bruit.- Masks restrictifs (
0xFC,0xF8) : process froze sur Android. Taskset Android tatillon ; mécanisme non-fiable. - Skip, garder par défaut tout-cores.
T2.1 — NEON 16-way sur heads ✨ GROS GAIN
cp_predict_cached.sample_head faisait 2048 dot products de 1024 floats en
boucle SCALAIRE pour chacune des 15 heads/frame. Total 31.5 M ops/frame de
host-side code. Le compilateur n'auto-vectorise pas (probable à cause des
indices avec head_idx).
Implémentation explicite dot_neon_1024 aarch64 :
float32x4_t s0,s1,s2,s3 = vdupq_n_f32(0);
for (int i = 0; i < 1024; i += 16) {
s0 = vfmaq_f32(s0, vld1q_f32(a+i+0), vld1q_f32(b+i+0));
// ... 4 accumulators in parallel (instruction-level parallelism)
}
return vaddvq_f32(vaddq_f32(vaddq_f32(s0,s1), vaddq_f32(s2,s3)));
| Config | CP/frame | Pipeline | RTF | WAV md5 |
|---|---|---|---|---|
| Scalar (KZTTS_CP_NO_NEON=1) | 106.8 ms | 8.05 s | 2.93 | c3dd71... |
| NEON 16-way | 72.1 ms | 6.75 s | 2.45 | c3dd71... ✓ |
-33% sur CP, -16% sur le pipeline, BIT-EXACT. Commit 9a0c63f.
Confirmation sur 3 phrases (33, 41, 60 frames) : RTF stable à 2.44-2.45.
T2.2 — threadpool ggml persistent (skipped après analyse)
ggml_graph_compute_with_ctx crée un threadpool éphémère à chaque appel.
ggml_threadpool_new + ggml_graph_plan + ggml_graph_compute permet de
réutiliser un threadpool global.
Estimation : ~0.5 ms / appel × 15 sub-forwards + 5 stages decoder = ~10 ms / phrase. Sur 6.75 s = 0.15% de gain. Pas la peine de refactor 20 sites d'appel.
État final post-optims
| Composant | Per-frame | % total | Plafond effort raisonnable |
|---|---|---|---|
| Talker | 32 ms | 15% | llama.cpp interne, peu de marge sans HTP |
| CP | 71 ms | 35% | NEON heads sorti, reste BW-bound ggml-cpu |
| Decoder | 99 ms | 49% | BigVGAN ggml-cpu NEON déjà, BW-bound |
| Prefill | 3 ms one-shot | 1% | text_projection NEON potentiel marginal |
| Total | ~203 ms | 100% | RTF 2.45 |
Leviers résiduels et leur ROI
-
Streaming pipeline (1 session, neutre RTF) : TTFB perçu 6.7 s → ~500 ms. Le seul gain UX disponible.
-
Decoder : ggml_conv_1d → ggml_im2col + ggml_mul_mat explicite (½ session) pour permettre de batcher les im2col (réduire BW conv1d). Gain incertain, probable 5-10% sur BigVGAN.
-
text_projection NEON (½ session) : ~50 ms one-shot au prefill → gain marginal pipeline -0.5%.
-
next_embed NEON + pre-alloc : ~1 ms / frame → -0.5%.
Aucun ne dépasse 5% additionnel sans toucher au kernel hexagon ou au backend ggml interne. L'angle "optim CPU pure" a livré son maximum.
Pour RTF<1 sans HTP
Demande :
- Soit refactor decoder pour kernel HVX/HMX bit-exact (frontier R&D, semaines, hors scope)
- Soit Vulkan backend Adreno pour le decoder (mal supporté ggml-vulkan
- conv1d, semaines)
- Soit accepter audio dégradé pour gagner BW (quant Q4/Q8 CP/decoder, test précis nécessaire)
Pour shipper en l'état : streaming pipeline rend RTF 2.45 acceptable côté UX (audio commence à jouer ~500 ms après le clic).
Commits associés
101d1cd: tests infra HTP/HMX — verdict HTP non-rentable9a0c63f: T2.1 NEON heads CP — gain principal -33% CP -16% pipe- Cette doc.