Optim prefill: ngl99/t8/b512=49.9 t/s (+11%), full offload, decode CPU

This commit is contained in:
Richard Loyer 2026-05-24 16:35:52 +02:00
parent 627f9a80e2
commit 3a5f735ab9
1 changed files with 3 additions and 0 deletions

View File

@ -55,3 +55,6 @@ cli prompt court -ngl99 HTP: REPOND, FR coherent (4B-Q4_0 qualite OK). Hang seul
## Correctif methode (24/05)
GGML_HEXAGON_OPFILTER N EXISTE PAS (invente). OPMASK=phase (SKIP_QUANTIZE/COMPUTE/QUEUE) global, pas par-op. GDN_KERNEL=0 = decode only. => pas d isolation GDN_cpu sans rebuild. Voie propre: bench rebuild DEBUG -> profiler total+Reste_htp meme run, GDN_cpu=total-Reste_htp (thermals neutres). Bug1 unaccounted 2^44 = underflow size_t buffer HTP >1 ubatch, fix borne, a lever pour pp512 RAG. Bug2 llama_params_fit abort -dev none = 3e chemin CPU casse, upstream, hors A/B. Qualite: PRE-validee 1 run reasoning, protocole FR a faire.
## Optim prefill Qwen3.5-4B (24/05)
Sweep ngl: 0=18.7/8=19/16=18/24=23/32=33/99=45 -> monotone, full offload gagne. Sweep t/b @ngl99: t8+b512=49.9 (best,+11%), t6=46.6, t4=41. CPU threads portent GDN-CPU+ amorti b512. OPTIMUM prefill = ngl99/t8/b512=49.9. decode reste CPU. Pas de compromis partiel. Plafond 50 = GDN-CPU; lever B-2 pour >.