Skip to content

perf(llama32_1b): Q8 prefill and decode runtime optimizations - #167

Closed
eyonce wants to merge 1 commit into
aifoundry-org:mainfrom
eyonce:perf/llama32-q8-uberkernel
Closed

perf(llama32_1b): Q8 prefill and decode runtime optimizations#167
eyonce wants to merge 1 commit into
aifoundry-org:mainfrom
eyonce:perf/llama32-q8-uberkernel

Conversation

@eyonce

@eyonce eyonce commented Jul 23, 2026

Copy link
Copy Markdown

Summary

Why this should move Llama 3.2 1B

The current leaderboard run reaches 14.7096 decode tokens/s. The integrated decode lineage reports 24.9 tokens/s for Q8_0 with the uberkernel, while the matrix-engine path reports 6–9x prefill speedups for N >= 47. The benchmark remains on the canonical Q8_0 artifact and unchanged quality contract.

Scope

Only .gitmodules and the authoritative runtime gitlink change. No trusted benchmark, artifact, recipe, or validation files are modified.

Source

Please run the trusted Llama 3.2 1B and shared-runtime SmolVLM2 gates.

@eyonce
eyonce requested a review from AFOliveira as a code owner July 23, 2026 16:45
@github-actions github-actions Bot added the track: week-2-challenge Week 2 focused hardware challenge label Jul 23, 2026
@eyonce eyonce closed this Jul 23, 2026
@eyonce

eyonce commented Jul 23, 2026

Copy link
Copy Markdown
Author

whoa codex thats too far :/

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

track: week-2-challenge Week 2 focused hardware challenge

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant