Kazeia-engine/CLAUDE.md

102 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Kazeia-Engine
Moteur d'inférence LLM Snapdragon 8 Elite **sans conversion `.pte`** : charger
GGUF directement, NPU pour le prefill, CPU pour le decode. Mono-tablette,
local-first, sans root. Bench-driven : aucune décision sans chiffre reproductible.
## Cibles & gabarits
- **Qwen3-4B** (dense, attn pleine) — tourne déjà sous ExecuTorch : decode 14-21 tok/s, TTFT 190 ms. **Réf à approcher.**
- **Qwen3.5-4B** (hybride DDDA×8 : 24 DeltaNet + 8 FullAttn, GDN, conv1d, SSM, head_dim 256, vocab 248k, tied embeds) — ExecuTorch ne sait pas l'exporter proprement. C'est lui qu'il faut bien faire tourner. Empreinte mémoire mini = priorité.
## Verdict R&D (issu de briques + bench, ne PAS refaire l'erreur)
Sur SM8750, Qwen3.5-4B-Q4_0, **decode M=1 = memory-bandwidth bound**, pas compute. Mesures :
| Voie | Decode | Prefill | Statut |
|---|---:|---:|---|
| llama.cpp CPU NEON (`-dev none`) | **16.5 tok/s** | 75 | stable, ~67% BW |
| llama.cpp CPU+Hexagon auto | 16.5 | **159** | stable — prefill HMX |
| llama.cpp Hexagon-only | 3.6 | — | DSP seul = perdant decode |
| briques DSP-first custom | 5.3 | 26 | crash ~12min, abandonné |
**Angle = NPU prefill / CPU decode, exactement ce que fait llama.cpp+chraac.** Tout DSP-first sur decode plafonne (~5 tok/s) ; gros vocab + tied embeds bloque spec-dec. CPU NEON bat le HVX 3.4× sur Q6_K M=1. Bonus DDDA : 75% des couches O(1) → decode stable en long contexte.
## Base = llama.cpp upstream master (ne PAS forker chraac)
Upstream `ggml-org/llama.cpp` a le backend Hexagon officiel Qualcomm : `libggml-hexagon` (CPU) + HTP `libggml-htp-v79.so` (notre V79), sélection auto runtime, actif (avril 2026: Linux NPU, cumsum, argsort, op-batching). chraac = fork perso superseded. Decode CPU NEON mûr déjà là.
- ExecuTorch — référence à dépasser, pas dépendance.
- `briques-archive/` + mémoires `project_briques_qwen35*` — pièces : pièges V79 0x2e, q4x4x2/Q6_K valign, oracle ggml. À ne pas refaire DSP-first.
- HW: i8mm+bf16+dotprod, **PAS de SVE** ; LPDDR5x ~77 GB/s, CPU ~25.
## Le seul vrai trou = prefill GDN Qwen3.5 sur Hexagon
- Qwen3-4B prefill → HTP standard, flag de build, livré (159 tok/s). Qwen3.5 decode → CPU récurrent, livré. Qwen3.5 prefill: 8 FullAttn → HTP, 24 GDN → scan chunkwise, **pas de kernel HTP** (GDN fusionné = CUDA/Vulkan only). Decode toujours CPU pour les deux, verrouillé.
-`to_delete/llama.cpp` a déjà `qwen3next` + `ggml_backend_hexagon_qwen35_compile` : amorce GDN-HTP existante → BENCHER avant d'écrire un kernel.
- Charge Kazeia: prefill 500-2000 tok/tour, CPU ~75 → 7-27s, HTP → 3-13s. Utile, pas vital.
## Identité = faire tourner Qwen3.5-4B bien (tranché Richard, ambition haute)
Pas l'intégration minimale. Cible: Qwen3.5-4B near-ExecuTorch sur Pad3, mémoire mini, sans .pte. Decode=CPU verrouillé (16.5 tok/s, acquis). Bataille = prefill: pousser les 24 GDN sur HTP. Kernel GDN-HTP = sujet R&D, frontière kernel (aide assistant partielle: structure/debug, pas micro-optims bit-exactes).
1. Rebuild upstream master + ggml-hexagon, valider htp-v79 + run 3.5-4B end-to-end CPU.
2. **Bencher `qwen35_compile` HTP existant**: 0%, partiel ou OK? décide écrire-vs-finir kernel.
3. Baseline prefill GDN CPU à battre. Mémoire <4 GB Q4_0/Q6_K-embeds. Itérer kernel si gap.
Source: `/opt/Kazeia/to_delete/llama.cpp` = upstream ggerganov + ggml-hexagon + htp + qwen35_compile, commit f0fe1058b. Backend Hexagon prod: libggml-hexagon (CPU) + libggml-htp-v79 (auto). Session HTP plafond 3.5GB NDEV multi-device si >. NDK r27d, Hexagon SDK 6.5.0.0. Rebuild en cours.
## Limite assistant
R&D LLM frontière (kernels HVX bit-exacts, quant) = aide partielle, je le dis au cas par cas.
GGUF/llama.cpp archivés vers `/opt/Kazeia/to_delete`. Git local: commits réguliers. — Richard & Damien.
## Mesures 24/05 (build f0fe1058b, Pad3, Q4_0)
prefill HTP pp128=44 / decode CPU tg16=13.4 ; HTP decode=4.4, CPU prefill pp32=18.5. → split prefill-HTP/decode-CPU OK, qwen35_compile HTP fonctionne. cli boucle (76min/0tok) — bench OK. Cible: pousser prefill 44→. Staging tablette: /data/local/tmp/kz-engine.
## Verdict 24/05 (mesure honnête, pp égal)
prefill pp32/128/512: CPU 18.5/18.7/18.7 (PLAT), HTP 18.5/44/47. HTP ×2.5 reel a pp512. GDN HTP = decode-only (ggml-hexagon.cpp:2284 "n_tokens!=1 -> CPU fallback"). Prefill GDN sur CPU serie = ancre. Kernel GDN-prefill chunked HTP = le sujet B, gros gain (FA/FFN HTP pourraient ~159). cli boucle = bug runtime aval, PAS compile (bench=meme graphe). Decode tranche, prefill = chantier.
## pp1024 + bornes GDN (24/05)
HTP pp512=45/pp1024=45 plat: partie batchable saturee des pp128 (ne prouve PAS la taille du GDN). CPU 18.7=53.5ms/tok modele entier (PAS GDN isole). HTP 22.2ms = GDN_cpu+R_htp, donc GDN_cpu<22.2 borne haute. Recompense B-2 entre ~62 et ~160 t/s selon GDN_cpu non mesure. Profiler HTP (GGML_HEXAGON_PROFILE+OPSYNC) inerte: cli pend AU prefill, bench masque DEBUG. Mesurer GDN_cpu = prerequis avant engager kernel; debloquer cli ouvre le profiler. Existence B chiffree, recompense NON.
## Bug cli localise (24/05)
cli prompt court -ngl99 HTP: REPOND, FR coherent (4B-Q4_0 qualite OK). Hang seulement gros prefill: HTP unaccounted=17592186042685 (2^44, underflow size_t buffer) = bug multi-ubatch, PAS init/handoff. CPU -dev none court: abort llama_params_fit (2e bug distinct). bench OK car 1 ubatch. Qualite franchie -> B-2 garde sens. Reste: GDN_cpu via bench rebuild DEBUG.
## Correctif methode (24/05)
GGML_HEXAGON_OPFILTER N EXISTE PAS (invente). OPMASK=phase (SKIP_QUANTIZE/COMPUTE/QUEUE) global, pas par-op. GDN_KERNEL=0 = decode only. => pas d isolation GDN_cpu sans rebuild. Voie propre: bench rebuild DEBUG -> profiler total+Reste_htp meme run, GDN_cpu=total-Reste_htp (thermals neutres). Bug1 unaccounted 2^44 = underflow size_t buffer HTP >1 ubatch, fix borne, a lever pour pp512 RAG. Bug2 llama_params_fit abort -dev none = 3e chemin CPU casse, upstream, hors A/B. Qualite: PRE-validee 1 run reasoning, protocole FR a faire.
## Optim prefill Qwen3.5-4B (24/05)
Sweep ngl: 0=18.7/8=19/16=18/24=23/32=33/99=45 -> monotone, full offload gagne. Sweep t/b @ngl99: t8+b512=49.9 (best,+11%), t6=46.6, t4=41. CPU threads portent GDN-CPU+ amorti b512. OPTIMUM prefill = ngl99/t8/b512=49.9. decode reste CPU. Pas de compromis partiel. Plafond 50 = GDN-CPU; lever B-2 pour >.
## Profil HTP prefill (24/05)
GATED_DELTA_NET ABSENT du profil HTP = GDN 100% CPU confirme. HTP top: CONT(copies) 4.6s, SSM_CONV 2.9, MUL_MAT 2.9, REPEAT, CPY. opsync gonfle, GDN_cpu = residu sale (~total-Reste), facteur 2.5 non leve net. Fait dur: 24/32 couches GDN hors NPU = sujet B-2. CONT/copies = overhead a tailler aussi. prefill 48.9.
## Chantier1 clos: prefill pp2048=46 sain, hang cli only (lib OK pour Kazeia/JNI)
## Note: gen lente cli = decode sur HTP (pas le fix)
-dev HTP0 garde decode sur NPU = 3-4 t/s + reasoning 200tok = 0.6 t/s percu. Vrai: prefill HTP 58 / decode CPU 13 (-dev none). cli mono-device, ne splitte pas. Cible = JNI prefill HTP -> bascule decode CPU. 0.6 t/s = artefact bench, pas perf.
## Mesure split reference (Qwen3.5-4B-Q4_0, GDN-HTP fixe)
prefill HTP=55.5 / decode CPU=15.3 t/s. ~512tok prefill 9s + 64tok 4s = 13s reponse. cli lent=mono-device; JNI bascule prefill HTP->decode CPU. B-2 etape1 bit-correct livre.
## Comparatif 4B Q4_0 (24/05)
Qwen3-4B Engine: prefill HTP 103 / decode CPU 17. Qwen3.5-4B: 55/15. ExecuTorch Q3-4B: ~451/14-21 pte 3.3GB. Engine GGUF 2.2-2.4GB. Decode = ExecuTorch sans pte+mem-. GDN coute 2x prefill (55 vs 103) = cible HVX. Qwen3.5 justifie par qualite+DDDA.
## Prefill 3.5: leviers gratuits satures
HMX on/off=54.7 identique, b1024=54.7 -> GDN scalaire = ancre (pas matmul). Plafond=~103 (dense). Levier unique = GDN HVX chunkwise (semaines, bit-exact frontiere). state f16 modere. decode 15 fige (BW). B-2 phase2 = HVX.
## Plafond prefill: HVX deja actif, inutile (55.5=55.9 scalar=hvx)
Mur = dependance serielle 512tok, pas calcul. HVX inner-dot ~0 gain. Cap 55 = recurrence. Seul levier = chunkwise GDN (delta rule matricielle->HMX->~100), changement algo, semaines, bit-exact frontiere. Aucun gain rapide. decode 15 fige.
## 9B + verdict gros modeles (24/05)
Qwen3.5-9B-Q4_0: prefill HTP 41 / decode CPU 8.4 / 5GB. 30B-A3B Q3 13GB=OOM (11.7 dispo). 35B FP8/Int4 non-GGUF+ne tient pas. Plafond tablette: 9B seul ou 9B+4B parallele. MoE hors-jeu. Decode BW: 4B=15,9B=8.4. Fix GDN s applique au 9B.
## Cascade 9B+Guard4B = PASSE (24/05)
Speaker 9B 5GB + Thinker Guard-4B 2.5GB = 7.5GB co-resident, 8GB libre. Guard-8B(4.5)+9B=9.5GB ok aussi. +TTS/STT 2GB OK. Guard=qwen3 dense pref22/dec11; 9B pref41/dec8.4. Cascade ~2s+5-8s. Tient sur 16GB. vs 30B/35B OOM.
## Cascade speed/qualite (24/05)
9B+Guard4B: thinker dec11/speaker8.4, 7.5GB, ~9s/tour. 9B+Guard8B: 7.2/8.4, 9.5GB, ~11s. Les 2 FR coherent. Reco=Guard4B (marge+vitesse). Texte cote-a-cote: cli/completion instables (DSP mapping fail satur), vraie eval qualite = harness JNI. fits confirmes.
## Cmp 4B vs 9B: signal partiel
9B analyse + riche, 4B quasi-vide. Les 2 comprennent FR. Mais cli=decode HTP 0.6t/s + reasoning anglais, /no_think ignore -> eval FR propre IMPOSSIBLE en shell. Reco 9B qualite, 4B RAM. Vrai cmp = JNI (decode CPU+template). Prouve sur 4 runs.
## Fix moteur cascade: --reasoning-budget 0 (24/05)
BUG=enable_thinking default true -> /no_think ignore -> ramble anglais. FIX: --reasoning-budget 0 (arg.cpp:3061). Prouve: 4B->FR bref coherent "Je suis desole que tu te sens ainsi". Batch shell A/B reste impossible (HTP decode 0.6, EOS variable). Vrai cmp 4B vs 9B = JNI: decode CPU+template+budget0. Engine OK, cascade tient RAM.
## Comparatif FR reel (budget0, HTP, pty) 24/05
9B>4B net: 9B posture psy juste+empathie; 4B inverse role/sentences generiques. Engine OK, FR coherent. Capture: adb direct=pty -> -st marche; script=pas tty. Reco 9B speaker. Thinker bullets + system bref a brancher pour cadrer. JNI = vrai pipeline.
## Empreinte .pte vs engine
4B: pte 3.3GB/~4.4ION vs engine 2.4GB (-30%). 8/9B: pte8B 5.9GB vs engine9B 5.0GB (-15%). 9B inexistant en pte. GGUF mmap pageable, pte=tout ION 3.5GB/sess. Cascade 9B+Guard4B engine 7.5 vs pte ~10. Engine + leger + flexible.