Integration #277: hang=thinking non bride, fix budget0+cap; 4B>9B CPU

This commit is contained in:
Richard Loyer 2026-05-25 09:38:25 +02:00
parent 76f21a21cb
commit fe6826a3be
1 changed files with 3 additions and 0 deletions

View File

@ -105,3 +105,6 @@ Talker=qwen3 GGUF chargeable engine (poids OK). Decoder RTF0.96 chraac valide hi
## GDN chunkwise verdict
Chunkwise batchable, 55->100 credible MAIS HMX accumulateur FP16 -> bit-exact impossible. Voie=HVX qf32 3-5sem frontiere. Prefill GDN-HTP encore not-bit-correct (2287 default OFF), decode CPU bit-exact. Reco: rester 55/15, gain pas worth (4s masques). Optim sure=release Thinker+KV f16.
## #277 integration: cause hang = thinking non bride
Dev report: 4B/9B 6-13min/tour, jamais done. CAUSE: EngineLlmEngine generate() sans reasoning_budget0 -> thinking infini (pas decode lent: 4B CPU=15tok/s ok). FIX bridge: enable_thinking=false + maxTok 64. 9B trop lent CPU (8.4)->4B speaker. .pte garde en attendant. #277=2 fixes Kotlin/JNI pas kernel. RAM 4.5 vs 3 mmap accepte.