Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
795 commits
Select commit Hold shift + click to select a range
3e031d6
AND OR XOR SHL SHR cannot have float operands [PR] (#17164)
chenyuxyz Jul 23, 2026
0f9edb0
don't use promo_lattice in fast_idiv [PR] (#17165)
chenyuxyz Jul 23, 2026
bf5989e
support loops in nir (kimi) (#17166)
geohot Jul 24, 2026
f5d9c31
arange upcast to int64 with big N (#17167)
chenyuxyz Jul 24, 2026
a292dab
corner cases from weak const branch (#17173)
chenyuxyz Jul 24, 2026
08eceb0
llm: fix generated tokens in usage accounting (#17174)
geohot Jul 24, 2026
2dd2e3d
don't support Tensor(list(np.array)) [pr] (#17175)
chenyuxyz Jul 24, 2026
709d7b2
hcq2 cleanups (#17162)
nimlgen Jul 24, 2026
0ecef21
hcq2 cleanup 2 (#17177)
nimlgen Jul 24, 2026
017d1ab
cpu prep for hcq2 (#17153)
nimlgen Jul 24, 2026
40bfd84
remove unneeded expand [PR] (#17176)
chenyuxyz Jul 24, 2026
c9b60ca
gptoss: moe gemm kernels (#17178)
wozeparrot Jul 24, 2026
fcb1d3f
add x86 loops (#17179)
GabrielNakamoto Jul 24, 2026
e0b30b7
simpler _split_cumalu [PR] (#17180)
chenyuxyz Jul 24, 2026
86b42d7
clean up unbroadcast [PR] (#17182)
chenyuxyz Jul 24, 2026
38c04fa
data_src concept in run_rangeify [pr] (#17183)
chenyuxyz Jul 24, 2026
dcad119
convert COPY -> STORE early (#17172)
geohot Jul 24, 2026
f65001e
cleanups because copy is not allowed in rangeify (kimi) (#17186)
geohot Jul 24, 2026
a346e2e
improve call for viz (#17187)
geohot Jul 24, 2026
de3508e
update create_non_native_float_pats [PR] (#17191)
chenyuxyz Jul 25, 2026
0a6125e
Invalid is bool, put weakint in promo lattice [pr] (#17188)
chenyuxyz Jul 25, 2026
9fdaa4b
standardize Program class (#17189)
sirhcm Jul 25, 2026
d923263
simpler lower_weakint_node [PR] (#17193)
chenyuxyz Jul 25, 2026
4c58b26
less wrong calculate_storage_offset (#17194)
chenyuxyz Jul 25, 2026
8a10892
fix torch backend as_strided (#17195)
chenyuxyz Jul 25, 2026
732e6bd
add one line viz mention (#17198)
Qazalin Jul 25, 2026
9f78504
checked cast in torch backend unwrap (#17199)
chenyuxyz Jul 25, 2026
983ad3b
fix torch backend batchnorm backward (#17201)
chenyuxyz Jul 25, 2026
f0117e9
refactor mlperf optim (#17200)
wozeparrot Jul 25, 2026
ee2ccb1
put weakint in dtypes.weaks [pr] (#17202)
chenyuxyz Jul 25, 2026
f902513
derive torch backend dispatch from the aten schema (#17203)
chenyuxyz Jul 25, 2026
74c2121
promo (uint64, int) -> weakfloat like JAX [pr] (#17204)
chenyuxyz Jul 25, 2026
492dc6d
delete _index_to_concrete_int [pr] (#17205)
chenyuxyz Jul 25, 2026
076b37e
failing batch norm test (#17206)
chenyuxyz Jul 25, 2026
3946df7
hcq2, cpu is hcq2-ish (#17197)
nimlgen Jul 25, 2026
a60b5f7
fix torch backend out= into a view (#17210)
chenyuxyz Jul 26, 2026
eb88905
weak frompy prerequisite [PR] (#17213)
chenyuxyz Jul 26, 2026
a96974e
slight weak behavior tweak and cleanups [pr] (#17214)
chenyuxyz Jul 26, 2026
acc2374
minor normalize cleanup [pr] (#17215)
chenyuxyz Jul 26, 2026
5d1aa84
SHR/SHL are Broadcastable [pr] (#17216)
chenyuxyz Jul 26, 2026
ac12914
nv: set lower interleave level to reduce GPU hogging (ai slop) (#15518)
Yaskophx Jul 26, 2026
960430a
Revert "nv: set lower interleave level to reduce GPU hogging (ai slop…
geohot Jul 26, 2026
d70a134
fix nvrtc_check helper used for jitlink call (#16362)
IntendedConsequence Jul 26, 2026
79c07a3
fix ValueError in UOp.axis for shard reshape crossing boundary (#16547)
mgisabsolute Jul 26, 2026
a8d5109
realize weak is no-op [pr] (#17219)
chenyuxyz Jul 26, 2026
94dad3d
clean up mixin cos and exp [PR] (#17221)
chenyuxyz Jul 26, 2026
4b67605
fix _prepare_jit_inputs for weak [pr] (#17222)
chenyuxyz Jul 26, 2026
97a2265
hcq2: amd indirect (#17220)
nimlgen Jul 26, 2026
165f062
speedy hcq2 (#17225)
nimlgen Jul 26, 2026
afb25a6
simpler minimum and copysign [PR] (#17226)
chenyuxyz Jul 26, 2026
456b5b5
fix python_alu inf (#17227)
chenyuxyz Jul 27, 2026
0f98212
skip test_float_to_fp8e4m3_extreme_values (#17228)
chenyuxyz Jul 27, 2026
95e3b00
webgpu failing test for duplicate PARAM in CALL [pr] (#17230)
Qazalin Jul 27, 2026
f45fc4c
test updates from weak flip (#17231)
chenyuxyz Jul 27, 2026
19c4d73
validate json output of viz.cli in CI (#17232)
Qazalin Jul 27, 2026
0569744
gptoss: split no-wd params (#17233)
wozeparrot Jul 27, 2026
a3ca1e5
polish the viz readme (#17236)
Qazalin Jul 27, 2026
818a892
flip from_py to use weak dtypes [pr] (#17229)
chenyuxyz Jul 27, 2026
8eaeede
hcq2: do not cache beam (#17234)
nimlgen Jul 27, 2026
bdbb1d7
fix shard axis through symbolic reshape (#17238)
b1tg Jul 27, 2026
8b9ef15
run_linear in external_test_gpu_crash (#17239)
nimlgen Jul 27, 2026
896afad
don't match strong typed const in UPat [pr] (#17240)
chenyuxyz Jul 27, 2026
0bb36c9
make Tensor(None) weakfloat (#17241)
chenyuxyz Jul 27, 2026
550225f
clean up pm_lower_weak [pr] (#17243)
chenyuxyz Jul 27, 2026
654d475
update pm_long_decomp [pr] (#17245)
chenyuxyz Jul 28, 2026
f837ca3
don't match const dtype in UPat [pr] (#17247)
chenyuxyz Jul 28, 2026
ab8fb19
64-bit UOp.variable support (#17246)
sirhcm Jul 28, 2026
4b7022e
Revert "64-bit UOp.variable support (#17246)" (#17248)
chenyuxyz Jul 28, 2026
1380d6c
pm_long_decomp doesn't depend on operand dtype [pr] (#17249)
chenyuxyz Jul 28, 2026
e3b3eea
test for extra copy in allreduce_cast (#16740)
Qazalin Jul 28, 2026
05e4727
arange stack regression test (#17250)
Qazalin Jul 28, 2026
37cf159
UOp.const(dtype=None) infers from from_py [pr] (#17253)
chenyuxyz Jul 28, 2026
a9ad080
make the github actions runners generic for gitea (#17254)
geohot Jul 28, 2026
c9e1154
delete explicit casts [pr] (#17255)
chenyuxyz Jul 28, 2026
749e002
remove some dtype= when construct UOp [pr] (#17257)
chenyuxyz Jul 28, 2026
d1c3ae0
preserve typed ranges through reshape (#17259)
drukpa1455 Jul 28, 2026
7d48926
switch _device_num to AxisType.DEVICE range (kimi) (#17252)
geohot Jul 28, 2026
9ce65b7
delete dsp_pm_late (#17260)
geohot Jul 28, 2026
fde3a8f
CAPTURE_PROCESS_REPLAY=0 default chaging test [PR] (#17261)
chenyuxyz Jul 28, 2026
0cdddf3
remove unneeded default args in renderers (#17262)
geohot Jul 28, 2026
755dfb2
rename CPU_COUNT to NUM_CPU_THREADS with cgroup awareness (#17263)
geohot Jul 28, 2026
23e9e76
DEFAULT_FLOAT/DEFAULT_INT ContextVar [pr] (#17265)
chenyuxyz Jul 28, 2026
f11f884
update dtypes.md (#17266)
chenyuxyz Jul 28, 2026
1757067
add device range as src[1] to multi (kimi) (#17264)
geohot Jul 28, 2026
57ae1bc
rename MULTI to UNSHARD (#17267)
geohot Jul 28, 2026
a17387d
add UNSHARD to spec (#17269)
geohot Jul 29, 2026
291ee43
qwen3.6 for 27b and 35b-a3b (#17268)
chenyuxyz Jul 29, 2026
dd16d5a
apply shrink bugfix for 3.11 (#17271)
geohot Jul 29, 2026
2f8f2d2
remove no-op explicit dtype= or cast [PR] (#17276)
chenyuxyz Jul 29, 2026
451120c
make .barrier implicit (kimi) (#17275)
geohot Jul 29, 2026
e684fcc
pm_reduce_collapse fix for re enabling stack for cat of same shape (c…
Qazalin Jul 29, 2026
527e573
fix smu reset for kernel >= 7 (#17277)
geohot Jul 29, 2026
6ea7d36
test permuted input in custom Ops.PROGRAM test (#17279)
Qazalin Jul 29, 2026
3803f15
fix add_raw_barrier [pr] (#17278)
chenyuxyz Jul 29, 2026
3df1b07
Patch half precision ops for older gpu archs (#17274)
NoahSchiro Jul 29, 2026
bd296a7
enable alloc_fragment support with UNSHARD (kimi) (#17272)
geohot Jul 29, 2026
6c2b9fa
weak const cleanups [PR] (#17282)
chenyuxyz Jul 29, 2026
52c9e5a
rename LOOP -> WEAK and STRONGLOOP -> LOOP (#17283)
geohot Jul 29, 2026
dd86a30
test case for weakfloat cast to weakint INDEX (#17287)
chenyuxyz Jul 29, 2026
aab51fb
64-bit UOp.variable support, try 2 (#17256)
sirhcm Jul 29, 2026
b30c7e0
support 2d on UNSHARD (kimi) (#17285)
geohot Jul 29, 2026
d4ba8b6
hcq2: use stack (#17286)
nimlgen Jul 29, 2026
027907a
fix c0*x<c1 symbolic [pr] (#17290)
chenyuxyz Jul 29, 2026
138676a
improve fragment example + index unshard (kimi) (#17288)
geohot Jul 29, 2026
fd912b3
generic c0*x<c1 [pr] (#17291)
chenyuxyz Jul 29, 2026
060f447
qcom: match cl for SP_CS_INSTR_SIZE (#17289)
sirhcm Jul 29, 2026
bfc9fc6
nicer TinyJit decorator (#17293)
sirhcm Jul 30, 2026
d52ef30
fix do_devectorize dtype (#17295)
chenyuxyz Jul 30, 2026
aba5ba4
update rangeify comment for mop after AFTER (#17298)
Qazalin Jul 30, 2026
7b7b9c1
spec tests for movement ops before custom_kernel (#17299)
Qazalin Jul 30, 2026
3f6f0a1
hcq2: less graph_rewrites (#17300)
nimlgen Jul 30, 2026
b5a2a56
gptoss moe routing (#17284)
wozeparrot Jul 30, 2026
e25bf77
kimi delta attention (#17281)
b1tg Jul 30, 2026
417245a
rewrite writes new dtype with dtype_from_uop [pr] (#17302)
chenyuxyz Jul 30, 2026
d05a3e6
Revert "rewrite writes new dtype with dtype_from_uop [pr] (#17302)" (…
chenyuxyz Jul 30, 2026
341c4ed
rewrite writes new dtype with dtype_from_uop try 2 [pr] (#17305)
chenyuxyz Jul 30, 2026
b488cc7
update dtype_from_uop for SHR/SHL (#17307)
chenyuxyz Jul 30, 2026
da15c43
update openpilot benchmarks (#17304)
sirhcm Jul 30, 2026
b290372
ftdi reset chestnut before running comma benchmark (#17306)
sirhcm Jul 30, 2026
ce500c1
broadcast doesn't cast CONST [pr] (#17308)
chenyuxyz Jul 30, 2026
fe8ece7
fix weak for image gate fusion [pr] (#17310)
chenyuxyz Jul 30, 2026
f2c2f44
Add softmin (#17292)
NoahSchiro Jul 31, 2026
6608b9d
adjust lower weak order [pr] (#17314)
chenyuxyz Jul 31, 2026
d65ea46
cleanup gemm fragment + add store unshard (#17313)
geohot Jul 31, 2026
a8c1e89
fp4 asm gemm 6+ pflops (#17315)
Qazalin Jul 31, 2026
13452b3
benchmark comma big model (#17312)
sirhcm Jul 31, 2026
0a3325f
add mxfp4 quantize and layout kernels (#17320)
Qazalin Jul 31, 2026
f7964ac
llama with MXFP4 (#17321)
Qazalin Jul 31, 2026
93b74c7
gptoss: grouped moe (#17322)
wozeparrot Jul 31, 2026
6d2700f
failing test for unbound _device_num err in BEAM (#17326)
Qazalin Jul 31, 2026
0c4bfae
coalesce ints (#17323)
nimlgen Jul 31, 2026
4f5cadd
gptoss ci (#17325)
wozeparrot Jul 31, 2026
155b84e
hcq2 faster schedule (#17324)
nimlgen Jul 31, 2026
8e2f175
const(dtype, b) -> const(b, dtype) [PR] (#17328)
chenyuxyz Jul 31, 2026
7f4dbb8
remove shape= from UOp.const [PR] (#17331)
chenyuxyz Jul 31, 2026
1095bbe
hcq2: fix ib reuse (#17330)
nimlgen Jul 31, 2026
b95bd5b
don't reset chestnut in benchmark (#17333)
sirhcm Jul 31, 2026
a11ee26
viz: prep for faster cli DEBUG=3 (#17334)
Qazalin Jul 31, 2026
ad24750
const(value, dtype) -> const(value).cast(dtype) in tests (#17335)
chenyuxyz Jul 31, 2026
8dc225e
dtype_from_uop(INS) is None [PR] (#17337)
chenyuxyz Jul 31, 2026
2774332
fix sym_infer for CAST (#17338)
chenyuxyz Jul 31, 2026
85ced44
tc: don't allow reduce over output dims (#17340)
sirhcm Jul 31, 2026
15d5152
heuristics: try multiple TC axes (#17341)
sirhcm Jul 31, 2026
8509891
__int__ and __float__ work for weak (#17342)
chenyuxyz Jul 31, 2026
a88f832
remove UOp.val (#17345)
geohot Aug 1, 2026
099d69f
ci: split macos unit test into metal and mock runners (#17346)
geohot Aug 1, 2026
9082ece
use .val to access the value of Ops.CONST (#17347)
geohot Aug 1, 2026
20b8ecf
support new USB vendor ID (#17348)
adeebshihadeh Aug 1, 2026
5a1c641
more arg -> val (#17349)
geohot Aug 1, 2026
161783d
add _device_num back to ast.variables (kimi) (#17327)
Qazalin Aug 1, 2026
b502fc1
more const arg -> val (#17350)
chenyuxyz Aug 1, 2026
6c0ec39
UOp.is_invalid [PR] (#17351)
chenyuxyz Aug 1, 2026
665822a
hcq2: faster replace (#17353)
nimlgen Aug 1, 2026
8e524ca
CONST related cleanups [pr] (#17352)
chenyuxyz Aug 1, 2026
98b700b
gptoss optim fixes (#17356)
wozeparrot Aug 1, 2026
15c936d
merge devectorize + indexing (#17354)
geohot Aug 1, 2026
fb607fb
faster devectorizer with one line (#17360)
geohot Aug 1, 2026
0258c7f
minor argstr and alloc stack cleanup [PR] (#17361)
chenyuxyz Aug 2, 2026
e14cadb
gptoss: set ASM_GEMM (#17363)
wozeparrot Aug 2, 2026
59df317
use UOp.const to create new consts [PR] (#17368)
chenyuxyz Aug 2, 2026
09dabfe
minor pm_float_decomp cleanup [PR] (#17371)
chenyuxyz Aug 2, 2026
05bc7c6
fix llm reasoning and Linear import (#17372)
geohot Aug 2, 2026
23c7813
no hard coded dtype int for rangeify debuf [PR] (#17373)
chenyuxyz Aug 3, 2026
314df72
deflake test_hcq with MOCKGPU (#17370)
chenyuxyz Aug 3, 2026
e22935c
hcq2: inputs table (#17374)
nimlgen Aug 3, 2026
be5f62d
llama: refactor amax stuff and skip in fp4 (#17375)
Qazalin Aug 3, 2026
7c1ce50
hcq2: epoch (#17376)
nimlgen Aug 3, 2026
3331944
gptoss: fix moe routing (#17377)
wozeparrot Aug 3, 2026
a2385ae
MAX_LINE_COUNT=26000 (#17378)
chenyuxyz Aug 3, 2026
104ee90
usb: wait for PCIe link after power on (#17380)
YassineYousfi Aug 3, 2026
33755a3
improve threefry codegen [pr] (#17379)
chenyuxyz Aug 3, 2026
c2625c7
scalar ALU index fix + llm: preserve_thinking (#17381)
geohot Aug 3, 2026
87289a7
hotfix: revert test_scalar_alu_index, violates spec
geohot Aug 3, 2026
3cb786f
llm: update test_llm_server tests (#17382)
geohot Aug 3, 2026
67dc02d
bitcast in python for _bits_to_rand [PR] (#17383)
chenyuxyz Aug 4, 2026
c21a552
llm: bugfixes + warmup (#17384)
geohot Aug 4, 2026
0170a30
use python bitcast in fold_bitcast [PR] (#17385)
chenyuxyz Aug 4, 2026
13fff4f
const_like cleanups [PR] (#17386)
chenyuxyz Aug 4, 2026
c9cd44b
more custom kernel contig input edge case tests (#17387)
Qazalin Aug 4, 2026
f993228
llama: accurate mxfp4 mfu (#17388)
Qazalin Aug 4, 2026
568bfb6
Revert "usb: wait for PCIe link after power on (#17380)"
geohot Aug 4, 2026
0796853
support symbolic shapes in allreduce (#17364)
b1tg Aug 4, 2026
6eedca5
dtype_from_uop for CUSTOM, CUSTOMI, PYLITERAL [PR] (#17390)
chenyuxyz Aug 4, 2026
0db63e1
update dtype_from_uop for IMAGE INDEX [PR] (#17391)
chenyuxyz Aug 4, 2026
85e9440
keep weak consts weak in symbolic [PR] (#17392)
chenyuxyz Aug 4, 2026
80d2073
fa: fix dq hazard with D=64 (#17393)
wozeparrot Aug 4, 2026
f489f4b
add test_eye + color INDEX (#17394)
geohot Aug 4, 2026
7b6d2dd
more weak const without cast in const_like [PR] (#17395)
chenyuxyz Aug 4, 2026
c1a10e0
fix _min_max for CAST from float to int [pr] (#17396)
chenyuxyz Aug 4, 2026
d79772f
fix pow on extreme inputs (#17397)
chenyuxyz Aug 4, 2026
f295f9f
use weak 0 in convert_pad_to_where_to_keep_behavior_local [pr] (#17398)
chenyuxyz Aug 4, 2026
6122b3c
use check_schedule in tests where possible (#17400)
geohot Aug 5, 2026
3eab809
update minimum to not create strong type const [PR] (#17401)
chenyuxyz Aug 5, 2026
e1f4268
add new schedule tests + format better (#17402)
geohot Aug 5, 2026
9b508df
remove invalid special case in cast [PR] (#17405)
chenyuxyz Aug 5, 2026
de57be1
kill nvidia pids at benchmarks start (#17406)
sirhcm Aug 5, 2026
46f0003
more KernelCountException (#17407)
geohot Aug 5, 2026
b45058b
don't cast weak in _broadcasted [pr] (#17408)
chenyuxyz Aug 5, 2026
77e124e
fix AMD WMMA emulation and test in CI (#17184)
GabrielNakamoto Aug 5, 2026
3bf9e70
Revert "don't cast weak in _broadcasted [pr] (#17408)" (#17409)
chenyuxyz Aug 5, 2026
874d331
hcq2 benchmark (#17235)
nimlgen Aug 5, 2026
ad32bd2
viz/cli: faster and more complete rewrites print (#17411)
Qazalin Aug 5, 2026
5b0b68e
remove debug from test (#17410)
nimlgen Aug 5, 2026
6cb419b
regression test for bert nan with weak (#17412)
chenyuxyz Aug 5, 2026
9b27ea8
hcq2: cleaner (#17413)
nimlgen Aug 5, 2026
2cce85a
chat: display reasoning_content from streamed responses (#17414)
geohot Aug 5, 2026
757a727
move callify into tensor (#17416)
geohot Aug 5, 2026
c2f1e5a
fix weak cast to strong dtype [pr] (#17418)
chenyuxyz Aug 5, 2026
07ac911
few weak and decomp tweaks [PR] (#17419)
chenyuxyz Aug 5, 2026
581bfdd
merge track_rewrites and profile_matches into rewrite_group [PR] (#17…
geohot Aug 5, 2026
a8a8030
add benchmark_llm script
geohot Aug 5, 2026
470c032
fix slice + non contig kernels (#17423)
geohot Aug 5, 2026
d726e5f
split pm_fold_cast_const [PR] (#17425)
chenyuxyz Aug 5, 2026
be25207
scope variable names inside CALLs (#17424)
sirhcm Aug 6, 2026
d51e55a
remove some pm_fold_cast_const [pr] (#17426)
chenyuxyz Aug 6, 2026
b4372df
revert wrong custom kernel fix (#17427)
geohot Aug 6, 2026
7a9cd8e
move weak function and pm to uop/weak [PR] (#17429)
chenyuxyz Aug 6, 2026
969df86
one less strong dtype const in symbolic [pr] (#17431)
chenyuxyz Aug 6, 2026
28e6ef6
fix postopt symbolic [pr] (#17433)
chenyuxyz Aug 6, 2026
f258708
llama: custom quantize_mxfp4+transpose kernel (codex) (#17434)
Qazalin Aug 6, 2026
9636dd1
test MXFP4 llama without hipcc (#17435)
Qazalin Aug 6, 2026
46230e9
hcq2: fence inputs (#17436)
nimlgen Aug 6, 2026
1fd6b10
fa: swa support (#17367)
wozeparrot Aug 6, 2026
d8cbc11
update linear interpolate to use int math for indices (#17441)
chenyuxyz Aug 7, 2026
9020a88
truncate float in DType.const [pr] (#17439)
chenyuxyz Aug 7, 2026
28195d5
fix f2f from fp8e5m2fnuz to half (#17442)
chenyuxyz Aug 7, 2026
f253c44
remove contiguous from custom_kernel (#17149)
Qazalin Aug 7, 2026
1858f1f
viz: collapse PROGRAM nodes like CALL (codex) (#17438)
geohot Aug 7, 2026
baa6148
fix var of large half input (#17444)
chenyuxyz Aug 7, 2026
0c96cdc
fix prod gradients at zero (#17404)
Robertboy18 Aug 7, 2026
fca695a
clean up reduce MUL gradient (#17447)
chenyuxyz Aug 7, 2026
73e670c
c0+x<c1 -> x < c1-c0 is ints only [pr] (#17448)
chenyuxyz Aug 7, 2026
b6189db
cpu: fix eintr (#17449)
nimlgen Aug 7, 2026
1827ec5
gptoss: fix sharded invalids (#17450)
wozeparrot Aug 7, 2026
f76422b
fix cast to float _min_max [pr] (#17451)
chenyuxyz Aug 7, 2026
59b88ea
move pm_fold_cast_const [pr] (#17453)
chenyuxyz Aug 7, 2026
4a3b8f6
better _drop_valid_stmts [pr] (#17454)
chenyuxyz Aug 7, 2026
4c206a5
fix ci emu (gpt) (#17437)
nimlgen Aug 7, 2026
c0d2f9a
nolocals supports variables (#17457)
sirhcm Aug 7, 2026
9dd3b84
default llama 8b to MXFP4=1 (#17465)
Qazalin Aug 8, 2026
8c49a7a
support symbolic shapes in copy (#17461)
b1tg Aug 8, 2026
abe2256
fix symbolic sharded reshape (#17463)
b1tg Aug 8, 2026
d4d537c
add SPEC checking for the kernel graph (#17432)
geohot Aug 8, 2026
e17c21e
hcq2: timings (#17464)
nimlgen Aug 8, 2026
8c8b43d
hcq2: fix beam (#17467)
nimlgen Aug 9, 2026
566f32f
move platform tests to platform.yml (#17475)
geohot Aug 10, 2026
44f1f45
llama: custom silu kernels (#17462)
Qazalin Aug 10, 2026
2821bd6
late loss.to("CPU") in llama (#17476)
Qazalin Aug 10, 2026
8611fe2
fix hevc (#17477)
nimlgen Aug 10, 2026
66ee3cf
Merge commit '8611fe2' into HEAD
Discountchubbs Aug 10, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
190 changes: 61 additions & 129 deletions .github/actions/setup-tinygrad/action.yml

Large diffs are not rendered by default.

6 changes: 3 additions & 3 deletions .github/workflows/autogen.yml
Original file line number Diff line number Diff line change
Expand Up @@ -37,16 +37,16 @@ jobs:
llvm: 'true'
pydeps: 'pyyaml mako'
- name: Install autogen support packages
run: sudo apt-get install -y --no-install-recommends libclang-20-dev llvm-20-dev hip-dev libusb-1.0-0-dev libdrm-dev
run: sudo apt-get install -y --no-install-recommends libclang-20-dev llvm-20-dev hip-dev libusb-1.0-0-dev libdrm-dev liburing-dev
- name: Regenerate autogen files
run: |
find tinygrad/runtime/autogen -type f -name "*.py" -not -path "*/amd/*" -not -name "__init__.py" -not -name "comgr.py" -not -name "metal.py" -not -name "iokit.py" -not -name "corefoundation.py" -not -name "libclang.py" -delete
python3 -c "from tinygrad.runtime.autogen import opencl"
python3 -c "from tinygrad.runtime.autogen import cuda, nvrtc, nvjitlink, nv_570, nv_580, nv"
python3 -c "from tinygrad.runtime.autogen import cuda, nvrtc, nvjitlink, nv_570, nv_580, nv_610, nv"
python3 -c "from tinygrad.runtime.autogen import comgr_3, hsa, hip, amd_gpu, sqtt, rocprof, amdgpu_kd, amdgpu_drm"
python3 -c "from tinygrad.runtime.autogen.am import *"
python3 -c "from tinygrad.runtime.autogen.nv_regs import *"
python3 -c "from tinygrad.runtime.autogen import libc, kfd, io_uring, ib, pci, vfio"
python3 -c "from tinygrad.runtime.autogen import libc, kfd, io_uring, pci, vfio"
python3 -c "from tinygrad.runtime.autogen import llvm"
python3 -c "from tinygrad.runtime.autogen import webgpu"
python3 -c "from tinygrad.runtime.autogen import kgsl, qcom_dsp"
Expand Down
843 changes: 327 additions & 516 deletions .github/workflows/benchmark.yml

Large diffs are not rendered by default.

4 changes: 2 additions & 2 deletions .github/workflows/mlperf.yml
Original file line number Diff line number Diff line change
@@ -1,8 +1,8 @@
name: Run MLPerf Training

on:
schedule:
- cron: '5 8 * * *' # Runs at 08:05 UTC (12:05 AM Pacific Time)
#schedule:
# - cron: '5 8 * * *' # Runs at 08:05 UTC (12:05 AM Pacific Time)
push:
branches:
- update_mlperf
Expand Down
213 changes: 213 additions & 0 deletions .github/workflows/platform.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,213 @@
name: Platform Tests
env:
# increment this when downloads substantially change to avoid the internet
CACHE_VERSION: '19'
CAPTURE_PROCESS_REPLAY: ${{ github.event_name == 'pull_request' && contains(github.event.pull_request.title, '[pr]') && '1' || '0' }}
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
PYTHONPATH: ${{ github.workspace }}
CHECK_OOB: 1

on:
push:
branches:
- master
pull_request:
workflow_dispatch:

concurrency:
group: platform-${{ github.event_name }}-${{ github.event_name == 'pull_request' && github.event.pull_request.number || github.run_id }}
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

jobs:

# ****** OSX Tests ******

unittestmacos:
name: MacOS (unit)
runs-on: macos-26
timeout-minutes: 20
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: unittest-macos
deps: testing_unit
- name: Run unit tests
run: DEV=METAL python -m pytest -n=auto test/unit/ --durations=20
- name: Test tensor core ops (fake)
run: DEV=METAL DEBUG=3 TC=2 python test/backend/test_ops.py TestOps.test_gemm
- name: Test tensor core ops (real)
run: DEV=METAL DEBUG=3 python test/backend/test_ops.py TestOps.test_big_gemm
- name: Test Beam Search
run: DEV=METAL IGNORE_BEAM_CACHE=1 python3 -m pytest extra/optimization/test_beam_search.py
- name: Test Device Specific
run: DEV=METAL python3 -m pytest test/device/test_metal.py
#- name: Fuzz Test linearizer
# run: DEV=METAL DEPTH=4 FUZZ_N=50 FUZZ_MAX_SIZE=1000000 python test/external/fuzz_linearizer.py
- name: Run process replay tests
uses: ./.github/actions/process-replay

unittestmacosmock:
name: MacOS (unit, mock)
runs-on: macos-26
timeout-minutes: 20
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: unittest-macos-mock
deps: testing_unit
amd: 'true'
ocelot: 'true'
- name: Run NULL backend tests
run: SPEC=2 DEV=NULL python -m pytest -n=auto test/null/ --durations=20
- name: Run pytest (amd)
env:
DEV: MOCKKFD+AMD
FORWARD_ONLY: 1
run: |
python3 -m pytest -n=auto test/device/test_hcq.py test/test_tiny.py --durations=20
- name: Run pytest (ptx)
env:
DEV: "MOCK+NV:PTX"
FORWARD_ONLY: 1
# TODO: failing due to library loading error
CAPTURE_PROCESS_REPLAY: 0
run: |
python3 -m pytest -n=auto test/device/test_hcq.py test/test_tiny.py \
test/testextra/test_hevc.py::TestHevc::test_hevc_decode_compile --durations=20
- name: Run process replay tests
uses: ./.github/actions/process-replay

testmetal:
strategy:
fail-fast: false
matrix:
group: [1, 2]
name: MacOS (DEV=METAL) (${{ matrix.group }})
runs-on: macos-26
timeout-minutes: 20
env:
DEV: METAL
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: macos-metal
deps: testing_unit
- name: Check Device.DEFAULT and print some source
run: |
python -c "from tinygrad import Device; assert Device.DEFAULT == 'METAL'"
DEBUG=4 python test/test_tiny.py TestTiny.test_plus
- name: Run backend tests
run: python -m pytest -n=auto test/backend --durations=20 --splits 2 --group ${{ matrix.group }}
- name: Run process replay tests
uses: ./.github/actions/process-replay

testmacos:
strategy:
fail-fast: false
matrix:
dev:
- 'CPU:CLANG'
- 'CPU:LLVM'
- 'CPU:LVP'
- 'WEBGPU'

name: MacOS (DEV=${{ matrix.dev }})
runs-on: macos-26
timeout-minutes: 20
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: macos-${{ matrix.dev }}
deps: "testing_unit${{ contains(matrix.dev, 'LVP') && ' mesa' || '' }}"
llvm: ${{ contains(matrix.dev, 'LLVM') || contains(matrix.dev, 'LVP') }}
webgpu: ${{ matrix.dev == 'WEBGPU' }}
- name: Set env
run: printf "DEV=${{ matrix.dev }}${{ matrix.dev == 'CPU:CLANG' && '\nCPU_COUNT=2' || '' }}" >> $GITHUB_ENV
- name: Check Device.DEFAULT and print some source
run: |
python -c "from tinygrad import Device; from tinygrad.helpers import Target; assert Device.DEFAULT == Target.parse('${{ matrix.dev }}').device"
DEBUG=4 python test/test_tiny.py TestTiny.test_plus
- name: Run test_tiny
run: python -m pytest -n=auto test/test_tiny.py --durations=20
- name: Run process replay tests
uses: ./.github/actions/process-replay

# ****** Windows Tests ******

testwindows:
strategy:
fail-fast: false
matrix:
dev:
- 'CPU:CLANG'
- 'CPU:LLVM'
- 'CPU:X86'
- 'WEBGPU'

name: Windows (DEV=${{ matrix.dev }})
runs-on: windows-2025
timeout-minutes: 15
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: windows-${{ matrix.dev }}-minimal
deps: testing_unit
pydeps: ${{ matrix.dev == 'WEBGPU' && 'dawn-python' || '' }}
- name: Set env
shell: bash
run: printf "DEV=${{ matrix.dev }}${{ matrix.dev == 'CPU:CLANG' && '\nCPU_COUNT=2' || '' }}" >> $GITHUB_ENV
- name: Check Device.DEFAULT and print some source
shell: bash
run: |
python -c "from tinygrad import Device; from tinygrad.helpers import Target; assert Device.DEFAULT == Target.parse('${{ matrix.dev }}').device"
DEBUG=4 python test/test_tiny.py TestTiny.test_plus
- name: Run test_tiny
shell: bash
run: python -m pytest -n=auto test/test_tiny.py --durations=20


qcomclcompiletests:
name: Compile-only (QCOM CL)
runs-on: ubuntu-24.04-arm
timeout-minutes: 15
steps:
- name: Checkout Code
uses: actions/checkout@v6
- name: Setup Environment
uses: ./.github/actions/setup-tinygrad
with:
key: compile-qcomcl
deps: testing_unit
tinydreno: 'true'
- name: Set env
shell: bash
run: printf "DEV=NULL:QCOMCL:a630\nNULL_ALLOW_COPYOUT=1" >> $GITHUB_ENV
- name: Run test_ops
shell: bash
run: |
python -c "from tinygrad import Device; assert Device.DEFAULT == 'NULL'"
DEBUG=4 python3 test/backend/test_ops.py TestOps.test_add
python -m pytest -n=auto test/backend/test_ops.py --durations=20
- name: Run test_ops (IMAGE)
shell: bash
env:
IMAGE: 1
DEV: "NULL:QCOMCL:a630,IMAGE_PITCH_ALIGNMENT=64"
run: |
DEBUG=4 python test/backend/test_ops.py TestOps.test_gemm | grep read_imagef
python -m pytest -n=auto test/backend/test_ops.py --durations=20
8 changes: 7 additions & 1 deletion .github/workflows/szdiff.yml
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,15 @@ jobs:
outputs:
branchstat: ${{ steps.brstat.outputs.stat}}
steps:
- name: Check code from PR branch
- name: Check code from PR branch
uses: actions/checkout@v6
with:
repository: ${{ github.event.pull_request.head.repo.full_name }}
ref: ${{ github.event.pull_request.head.sha }}
fetch-depth: 0
# PR code is only inspected with git rev-list, never executed
allow-unsafe-pr-checkout: true
persist-credentials: false
- name: Check whether branch is up-to-date
id: brstat
run: |
Expand Down Expand Up @@ -51,6 +54,9 @@ jobs:
repository: ${{ github.event.pull_request.head.repo.full_name }}
ref: ${{ github.event.pull_request.head.sha }}
path: pr
# PR code is only line-counted by master's sz.py, never executed
allow-unsafe-pr-checkout: true
persist-credentials: false
# the base default to tinygrad master and cannot be other fork branch for security purpose
- name: Checkout code from tinygrad master
uses: actions/checkout@v6
Expand Down
Loading
Loading