1.2 KiB
1.2 KiB
Statut & limites (MAJ 26/05) — à lire avant intégration
VALIDÉ device : load 4B, generate() bout-en-bout (test_native), decode CPU exact, FR cohérent,
thinking-off déterministe. Modèle tranché = q35-lmq4 (éval qualité ../eval/VERDICT.md).
Perf : prefill clos côté kernel GDN (pas le goulot), decode au plafond CPU.
LIMITES / OUVERT :
- Prefill HTP CORRIGÉ (27/05) : charabia = 1 op (SSM_CONV conv1d) cassée sur HTP en multi-token. Fix
intégré au backend (
libggml-hexagon.so: supported_ssm_conv n_t>1→CPU). Plus d'OPFILTER. Prefill HTP 110-180 t/s cohérent. JNI = option C (prefill HTP→transfert KV→decode CPU), validé multi-tour. Rebuild libggml-hexagon.so + libkazeia_engine.so avant ship. Options A/B/C dans HANDOFF.md déc.2. - KV = f16 (résolu, JNI corrigé) : f16=10.9 vs q8_0=6.5 au decode. Rebuild .so requis (livré=q8_0).
generate()applique déjà ChatML + thinking-off. App fournitsys/usr(voir system_fr.txt).- TTS Talker GGUF chargeable mais audio e2e non validé (= intégration app).
- STT inchangé (ORT-QAIRT). 9B/30B/35B écartés (lents/OOM).
- Le tg/pp de ce soir étaient throttlés (batterie ~11%) — reconfirmer les chiffres sur batterie pleine.