diff --git a/CLAUDE.md b/CLAUDE.md index 3b36e6f..93f1007 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -117,3 +117,6 @@ CPU ngl0: t1=1.6, t8=14tok/s. in-app 0.32<1.6 = 1 thread + HTP init residuel. FI ## Comparatif final 4B in-app (dev) pte NPU: 15tok/s, 1.7s, 2.2GB. GGUF bridge: 0.32tok/s (1thread+HTP res), 4.5GB. GGUF standalone t8: 13. pte~standalone match, 2x moins RAM. #277 merge-able SI bridge t8 (->13~pte) mais RAM reste 2x. SEUL gain GGUF = 9B/modeles sans pte. 4B: pte gagne. Bloque sur: n_threads=8 + couper HTP init in-app. + +## 9B = vraie cible (pas de pte 9B) +9B in-app: CPU 11%=1 coeur (n_threads=8 perdu), 0.16tok/s, 6.8GB. standalone t8=8.4=~8s/tour utilisable. 8B pte 3GB existe vs 9B GGUF 6.8GB. merge 9B tient a t8 effectif. cap64 borne hang. Bench 8B pte vs 9B gguf qualite=vrai arbitrage. dev: forcer affinity 8 grands coeurs.