Skip to content

perf(llama32_1b): Q8 prefill and decode runtime optimizations - #167

Closed
eyonce wants to merge 1 commit into
aifoundry-org:mainfrom
eyonce:perf/llama32-q8-uberkernel
Closed

perf(llama32_1b): Q8 prefill and decode runtime optimizations#167
eyonce wants to merge 1 commit into
aifoundry-org:mainfrom
eyonce:perf/llama32-q8-uberkernel

Commits

Commits on Jul 23, 2026