Kazeia-engine/dist/STATUS.md

1.1 KiB

Statut & limites (MAJ 26/05) — à lire avant intégration

VALIDÉ device : load 4B, generate() bout-en-bout (test_native), decode CPU exact, FR cohérent, thinking-off déterministe. Modèle tranché = q35-lmq4 (éval qualité ../eval/VERDICT.md). Perf : prefill clos côté kernel GDN (pas le goulot), decode au plafond CPU.

LIMITES / OUVERT :

  1. Prefill HTP = CASSÉ (sortie charabia), prefill OBLIGATOIREMENT CPU (~14 t/s). Vérifié 27/05 (cli -dev HTP0 + harness). Le débit HTP est réel mais l'état produit est corrompu. Dense sur HTP crashe (0x2e). Bridge CPU-only (ngl0) = la seule config correcte. cf PERF.md / HANDOFF.md décision 2.
  2. KV = f16 (résolu, JNI corrigé) : f16=10.9 vs q8_0=6.5 au decode. Rebuild .so requis (livré=q8_0).
  3. generate() applique déjà ChatML + thinking-off. App fournit sys/usr (voir system_fr.txt).
  4. TTS Talker GGUF chargeable mais audio e2e non validé (= intégration app).
  5. STT inchangé (ORT-QAIRT). 9B/30B/35B écartés (lents/OOM).
  6. Le tg/pp de ce soir étaient throttlés (batterie ~11%) — reconfirmer les chiffres sur batterie pleine.