Product Models Benchmarks Docs Releases Sign in Download GitHub
Benchmarks

Veizik performance, measured on real hardware.

Every row below is one measurement on a named device, under a named protocol. Each protocol lists what a row has to carry; a row missing any of it is held back rather than shown. Results are added as rows, not as prose.

Apple Silicon devices
4
comparison cells, current prose
139
competing runtimes
2
reproducible protocols
4

A further 12 comparison cells are published below, measured on an earlier set of prose that is being replaced. They are not counted in the figure above, because a cell on one set of prose and a cell on another are not two measurements of the same thing.

Comparison

Against other engines, on the same machine

Decode and Prefill against MLX and MLX bf16, on Apple M2 Pro and Apple M1 Ultra. Veizik is ahead in 134 of 139 cells, and the cells it is behind on are here too. A further 12 cells, on Apple M1 Ultra and Apple M4 Pro, were measured on the earlier prose files rather than this one; they are in their own tables below, and none of them is counted in the figure above. Qwen2.5-0.5B, Qwen2.5-1.5B and Qwen2.5-7B are marked a lower bound. On those models the engine's prefill clock stops before the last prompt step has finished, and that step is charged to decode; these figures put a whole decode step back in, which costs more than the step does. So the engine is at least this fast on both axes and may be faster: a win here is safe, while a loss, or a cell held because the two ranges overlap, may be understating the engine rather than describing it. Over Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-7B, Qwen3-1.7B. Each board below names its own unit, which direction is better, and how many runs it is a median of.

Apple M4 Pro

Model weight footprint against llama.cpp · MB, lower is better · ahead in 4 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced

ModelVeizikBaselineRatioSpreadHow each side was timed
Qwen3-1.7B968.0llama.cpp 1,276.5
Q4_K_M GGUF
0.25.3 / build 11146 (7fe450e19), Metal backend
1.32×——
Qwen2.5-0.5B343.6llama.cpp 391.9
Q4_K_M GGUF
0.25.3 / build 11146 (7fe450e19), Metal backend
1.14×——
Qwen2.5-7B4,114.0llama.cpp 4,677.1
Q4_K_M GGUF
0.25.3 / build 11146 (7fe450e19), Metal backend
1.14×——
Qwen2.5-1.5B926.9llama.cpp 980.1
Q4_K_M GGUF
0.25.3 / build 11146 (7fe450e19), Metal backend
1.06×——

Apple M2 Pro

Decode against MLX · tok/s, higher is better · ahead in 19 of 24 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B30,000 → 128110.3
109.8–110.9
MLX 74.7
74.5–74.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.48×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B24,576 → 128120.9
120.3–121.3
MLX 85.6
85.5–85.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.41×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B16,384 → 128141.7
139.8–143.2
MLX 109.5
109.4–109.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.29×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B8,192 → 128192.8
189.2–195.6
MLX 151.1
150.7–151.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.28×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B512 → 128300.6
299.0–302.5
MLX 256.2
255.9–256.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.17×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B3,968 → 128222.0
221.5–223.4
MLX 191.1
190.4–192.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.16×0.9%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B30,000 → 12867.8
67.5–68.2
MLX 59.4
59.3–59.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.14×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B24,576 → 12874.1
73.8–74.5
MLX 67.1
66.9–67.2
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.10×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B512 → 128143.8
142.5–144.4
MLX 133.4
133.2–133.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.08×0.3%both sides: the engine's own clock
Qwen3-1.7B16,384 → 12860.1
60.0–60.2
MLX 57.3
57.2–57.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.2%both sides: the engine's own clock
Qwen2.5-1.5B16,384 → 12886.3
85.8–86.5
MLX 82.6
82.5–82.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B24,576 → 12846.4
46.2–46.4
MLX 44.5
44.5–44.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.04×0.1%both sides: the engine's own clock
Qwen3-1.7B3,968 → 128106.6
106.3–106.8
MLX 102.6
102.5–102.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.04×0.2%both sides: the engine's own clock
Qwen2.5-7B30,000 → 12823.5
23.5–23.6
MLX 22.7
22.7–22.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.04×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B8,192 → 12883.9
83.6–84.2
MLX 80.9
80.8–80.9
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.04×0.2%both sides: the engine's own clock
Qwen2.5-7B24,576 → 12825.3
25.2–25.3
MLX 24.7
24.7–24.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.02×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B30,000 → 12840.3
40.2–40.3
MLX 39.4
39.3–39.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.02×0.3%both sides: the engine's own clock
Qwen2.5-1.5B512 → 128144.6
144.3–146.6
MLX 142.1
141.8–142.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.02×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B16,384 → 12828.5
28.4–28.6
MLX 28.3
28.3–28.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.01×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B8,192 → 128102.4
101.6–103.6
MLX 104.1
103.9–104.2
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
0.98×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B8,192 → 12832.3
32.1–32.5
MLX 33.1
33.1–33.2
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
0.98×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B3,968 → 128113.7
109.6–116.6
MLX 119.2
119.1–119.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
0.95×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B512 → 12838.3
38.0–38.4
MLX 40.6
40.5–40.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
0.94×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B3,968 → 12833.0
32.8–33.1
MLX 36.6
36.6–36.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
0.90×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast

Decode against MLX bf16 · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-7B512 → 12838.3
38.0–38.4
MLX bf16 12.7
12.7–12.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
3.02×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B512 → 128143.8
142.5–144.4
MLX bf16 48.9
48.8–49.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.94×0.4%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B8,192 → 12832.3
32.1–32.5
MLX bf16 11.9
11.8–11.9
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.73×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B3,968 → 12833.0
32.8–33.1
MLX bf16 12.3
12.3–12.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.68×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B512 → 128144.6
144.3–146.6
MLX bf16 54.6
54.5–54.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.65×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B16,384 → 12828.5
28.4–28.6
MLX bf16 11.2
11.2–11.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.55×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B24,576 → 12825.3
25.2–25.3
MLX bf16 10.5
10.5–10.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.40×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B3,968 → 128106.6
106.3–106.8
MLX bf16 44.5
44.4–44.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.40×0.3%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B30,000 → 12823.5
23.5–23.6
MLX bf16 10.2
10.2–10.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.31×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B3,968 → 128113.7
109.6–116.6
MLX bf16 50.8
50.7–50.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.24×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B8,192 → 128102.4
101.6–103.6
MLX bf16 48.1
48.1–48.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.13×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B8,192 → 12883.9
83.6–84.2
MLX bf16 39.7
39.6–39.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.12×0.3%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B16,384 → 12886.3
85.8–86.5
MLX bf16 43.1
43.0–43.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.00×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B512 → 128300.6
299.0–302.5
MLX bf16 150.9
150.8–151.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.99×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B24,576 → 12874.1
73.8–74.5
MLX bf16 38.4
38.4–38.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.93×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B30,000 → 12867.8
67.5–68.2
MLX bf16 35.8
35.7–35.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.90×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B16,384 → 12860.1
60.0–60.2
MLX bf16 33.1
33.0–33.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.82×0.4%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B8,192 → 128192.8
189.2–195.6
MLX bf16 109.8
109.5–110.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.76×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B3,968 → 128222.0
221.5–223.4
MLX bf16 127.0
126.4–127.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.75×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B24,576 → 12846.4
46.2–46.4
MLX bf16 28.5
28.4–28.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.63×0.3%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B16,384 → 128141.7
139.8–143.2
MLX bf16 90.7
90.6–90.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.56×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B30,000 → 12840.3
40.2–40.3
MLX bf16 26.1
26.1–26.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.54×0.2%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B24,576 → 128120.9
120.3–121.3
MLX bf16 80.4
80.2–80.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.50×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B30,000 → 128110.3
109.8–110.9
MLX bf16 74.6
74.6–74.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.48×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint

Prefill against MLX · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B512 → 1285,059.7
4,985.3–5,101.3
MLX 3,274.5
3,195.9–3,295.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.55×3.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B512 → 1281,459.1
1,450.2–1,461.4
MLX 988.6
981.0–995.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.48×1.5%both sides: the engine's own clock
Qwen2.5-1.5B512 → 1281,505.5
1,497.9–1,510.8
MLX 1,049.5
1,042.2–1,056.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.43×1.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B512 → 128333.1
332.6–333.5
MLX 232.9
231.9–233.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.43×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B3,968 → 128327.2
326.9–327.3
MLX 230.4
230.2–230.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.42×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B8,192 → 128306.8
306.8–307.0
MLX 217.9
217.7–218.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.41×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B16,384 → 128271.9
271.9–272.0
MLX 196.4
196.4–196.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.38×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B3,968 → 1281,447.8
1,446.9–1,448.3
MLX 1,048.2
1,046.6–1,050.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.38×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B3,968 → 1281,305.8
1,305.0–1,306.5
MLX 947.4
945.1–949.1
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.38×0.4%both sides: the engine's own clock
Qwen2.5-7B24,576 → 128243.6
243.6–243.7
MLX 178.3
178.3–178.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.37×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B3,968 → 1284,664.4
4,656.9–4,673.1
MLX 3,417.0
3,379.9–3,456.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.37×2.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B8,192 → 1281,272.9
1,272.5–1,273.1
MLX 937.9
936.4–939.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.36×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-7B30,000 → 128227.6
227.6–227.7
MLX 167.8
167.7–167.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.36×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B8,192 → 1281,112.0
1,111.5–1,112.5
MLX 825.7
824.4–829.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.35×0.6%both sides: the engine's own clock
Qwen2.5-0.5B8,192 → 1283,916.5
3,913.8–3,919.2
MLX 2,939.7
2,928.8–2,959.2
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.33×1.0%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B16,384 → 128880.2
879.5–880.8
MLX 662.5
661.7–662.9
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.33×0.2%both sides: the engine's own clock
Qwen2.5-1.5B16,384 → 1281,024.8
1,024.7–1,025.1
MLX 774.9
773.5–775.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.32×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen3-1.7B24,576 → 128727.0
726.7–727.4
MLX 550.3
550.1–551.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.32×0.2%both sides: the engine's own clock
Qwen3-1.7B30,000 → 128651.1
650.8–651.6
MLX 495.2
495.0–495.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.31×0.2%both sides: the engine's own clock
Qwen2.5-1.5B24,576 → 128854.5
854.3–854.7
MLX 656.1
655.5–657.1
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.30×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B30,000 → 128769.6
769.5–769.7
MLX 593.8
593.1–594.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.30×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B16,384 → 1282,947.8
2,946.4–2,948.2
MLX 2,309.8
2,304.8–2,313.9
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.28×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B24,576 → 1282,356.1
2,355.0–2,357.1
MLX 1,897.5
1,892.7–1,901.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.24×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B30,000 → 1282,071.3
2,070.2–2,071.9
MLX 1,696.5
1,694.4–1,699.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.22×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast

Prefill against MLX bf16 · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B512 → 1285,059.7
4,985.3–5,101.3
MLX bf16 3,944.8
3,917.3–4,015.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.28×2.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B512 → 1281,459.1
1,450.2–1,461.4
MLX bf16 1,218.9
1,209.7–1,228.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.20×1.6%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B30,000 → 128651.1
650.8–651.6
MLX bf16 566.9
566.5–567.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.15×0.2%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B512 → 128333.1
332.6–333.5
MLX bf16 293.7
292.8–295.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.13×0.9%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B24,576 → 128727.0
726.7–727.4
MLX bf16 641.2
640.7–641.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.13×0.1%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B512 → 1281,505.5
1,497.9–1,510.8
MLX bf16 1,328.7
1,316.6–1,337.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.13×1.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B30,000 → 128769.6
769.5–769.7
MLX bf16 692.5
691.9–693.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.11×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B16,384 → 128880.2
879.5–880.8
MLX bf16 798.3
797.6–799.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.10×0.2%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B24,576 → 128854.5
854.3–854.7
MLX bf16 779.0
778.3–779.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.10×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B30,000 → 128227.6
227.6–227.7
MLX bf16 208.8
208.7–208.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.09×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B24,576 → 1282,356.1
2,355.0–2,357.1
MLX bf16 2,170.9
2,163.0–2,176.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.09×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B30,000 → 1282,071.3
2,070.2–2,071.9
MLX bf16 1,909.1
1,903.4–1,914.9
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.08×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B16,384 → 1282,947.8
2,946.4–2,948.2
MLX bf16 2,722.1
2,718.3–2,731.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.08×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B24,576 → 128243.6
243.6–243.7
MLX bf16 225.3
225.2–225.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.08×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B16,384 → 1281,024.8
1,024.7–1,025.1
MLX bf16 951.0
949.8–951.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.08×0.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B8,192 → 1283,916.5
3,913.8–3,919.2
MLX bf16 3,667.8
3,654.3–3,670.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.07×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B16,384 → 128271.9
271.9–272.0
MLX bf16 254.8
254.7–255.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.07×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B8,192 → 1281,112.0
1,111.5–1,112.5
MLX bf16 1,047.2
1,045.9–1,048.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.06×0.3%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B8,192 → 1281,272.9
1,272.5–1,273.1
MLX bf16 1,203.8
1,202.6–1,207.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.06×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B3,968 → 1284,664.4
4,656.9–4,673.1
MLX bf16 4,413.9
4,376.2–4,429.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.06×1.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B8,192 → 128306.8
306.8–307.0
MLX bf16 291.3
291.0–291.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen3-1.7B3,968 → 1281,305.8
1,305.0–1,306.5
MLX bf16 1,241.0
1,238.9–1,243.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.3%both sides: the engine's own clock
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-7B3,968 → 128327.2
326.9–327.3
MLX bf16 312.1
311.4–312.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B3,968 → 1281,447.8
1,446.9–1,448.3
MLX bf16 1,385.3
1,383.0–1,391.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.05×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint

Apple M1 Ultra

Decode against MLX · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B16,384 → 128259.1
257.4–262.0
MLX 177.7
177.0–178.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.46×0.8%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B512 → 128377.4
355.8–385.6
MLX 272.3
271.0–273.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.39×1.0%both sides: the engine's own clock
lower bound: veizik is at least this fast
Re-measured: pre-pause environment.
Qwen2.5-0.5B8,192 → 128280.6
278.7–286.6
MLX 203.6
202.9–204.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.38×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B30,000 → 128199.0
196.9–200.3
MLX 144.5
144.2–145.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.38×1.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B24,576 → 128205.6
204.4–207.8
MLX 155.6
155.3–157.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.32×1.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B3,968 → 128305.5
304.4–313.3
MLX 238.9
238.6–239.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.28×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B3,968 → 128233.9
215.3–241.2
MLX 188.5
182.6–190.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.24×4.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B8,192 → 128201.0
200.2–204.0
MLX 164.0
163.7–166.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.23×1.8%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B512 → 128260.2
244.7–261.2
MLX 213.5
213.3–214.0
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.22×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B16,384 → 128169.1
168.1–170.5
MLX 146.0
145.8–146.8
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.16×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B24,576 → 128145.8
144.8–147.6
MLX 127.6
125.8–128.2
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.14×1.9%both sides: the engine's own clock
lower bound: veizik is at least this fast

Decode against MLX bf16 · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-1.5B512 → 128260.2
244.7–261.2
MLX bf16 126.9
126.8–127.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
2.05×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B3,968 → 128233.9
215.3–241.2
MLX bf16 117.8
94.3–118.5
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.99×25.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B8,192 → 128201.0
200.2–204.0
MLX bf16 106.8
106.6–107.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.88×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B16,384 → 128169.1
168.1–170.5
MLX bf16 99.5
98.9–99.9
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.70×1.1%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B24,576 → 128145.8
144.8–147.6
MLX bf16 90.4
90.1–91.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.61×1.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B512 → 128377.4
355.8–385.6
MLX bf16 248.5
248.2–249.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.52×0.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Re-measured: pre-pause environment.
Qwen2.5-0.5B16,384 → 128259.1
257.4–262.0
MLX bf16 172.3
171.9–175.1
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.50×1.8%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B8,192 → 128280.6
278.7–286.6
MLX bf16 193.4
192.7–193.9
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.45×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B3,968 → 128305.5
304.4–313.3
MLX bf16 222.3
222.0–222.6
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.37×0.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B30,000 → 128199.0
196.9–200.3
MLX bf16 145.9
145.7–146.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.36×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B24,576 → 128205.6
204.4–207.8
MLX bf16 156.3
155.9–156.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.31×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint

Prefill against MLX · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B512 → 1289,631.8
9,087.7–10,740.8
MLX 4,684.7
4,556.1–4,885.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
2.06×7.2%
accounted for
both sides: the engine's own clock
lower bound: veizik is at least this fast
Re-measured: pre-pause environment.
Qwen2.5-1.5B512 → 1284,032.0
3,944.9–4,038.6
MLX 2,487.6
2,448.8–2,532.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.62×3.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B3,968 → 12813,075.1
12,868.0–13,273.7
MLX 9,859.3
9,809.9–9,924.3
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.33×1.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B24,576 → 1282,795.8
2,737.8–2,800.7
MLX 2,132.9
2,123.5–2,134.6
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.31×0.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B16,384 → 1283,309.3
3,302.2–3,314.0
MLX 2,532.2
2,518.3–2,536.1
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.31×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B3,968 → 1284,469.7
4,324.8–4,478.4
MLX 3,426.6
3,404.8–3,437.7
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.30×1.0%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-1.5B8,192 → 1284,021.3
3,975.1–4,042.7
MLX 3,091.9
3,060.5–3,101.4
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.30×1.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B8,192 → 12811,327.6
11,287.5–11,399.0
MLX 9,232.8
9,173.2–9,343.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.23×1.9%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B16,384 → 1288,509.2
8,488.2–8,525.7
MLX 7,486.9
7,422.1–7,510.5
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.14×1.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B24,576 → 1286,773.1
6,764.9–6,777.1
MLX 6,231.3
6,195.5–6,248.1
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.09×0.8%both sides: the engine's own clock
lower bound: veizik is at least this fast
Qwen2.5-0.5B30,000 → 1285,895.2
5,891.9–5,899.0
MLX 5,587.5
5,569.4–5,600.1
4-bit g64, converted from the snapshot
mlx-lm/mlx 0.31.3 0.32.2
1.06×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast

Prefill against MLX bf16 · tok/s, higher is better · ahead in 10 of 10 · median of 11 runs · on Darwin, On the Origin of Species

ModelTokensVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-0.5B512 → 1289,631.8
9,087.7–10,740.8
MLX bf16 4,980.0
4,061.3–5,160.7
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.93×27.1%
accounted for
both sides: the engine's own clock
lower bound: veizik is at least this fast
Re-measured: pre-pause environment.
Qwen2.5-1.5B512 → 1284,032.0
3,944.9–4,038.6
MLX bf16 2,688.4
2,603.2–2,739.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.50×5.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B3,968 → 12813,075.1
12,868.0–13,273.7
MLX bf16 10,792.2
10,660.6–10,891.0
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.21×2.2%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B24,576 → 1282,795.8
2,737.8–2,800.7
MLX bf16 2,332.7
2,320.9–2,335.3
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.20×0.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B16,384 → 1283,309.3
3,302.2–3,314.0
MLX bf16 2,826.1
2,812.7–2,832.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.17×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B8,192 → 1284,021.3
3,975.1–4,042.7
MLX bf16 3,558.6
3,476.4–3,567.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.13×2.6%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B8,192 → 12811,327.6
11,287.5–11,399.0
MLX bf16 10,187.9
10,097.6–10,253.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.11×1.5%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-1.5B3,968 → 1284,469.7
4,324.8–4,478.4
MLX bf16 4,023.6
4,003.4–4,053.8
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.11×1.3%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B16,384 → 1288,509.2
8,488.2–8,525.7
MLX bf16 8,076.9
7,995.0–8,107.2
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.05×1.4%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint
Qwen2.5-0.5B24,576 → 1286,773.1
6,764.9–6,777.1
MLX bf16 6,621.0
6,578.9–6,627.4
bf16, the base snapshot unquantised
mlx-lm/mlx 0.31.3 0.32.2
1.02×0.7%both sides: the engine's own clock
lower bound: veizik is at least this fast
the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint

Model weight footprint against llama.cpp · MB, lower is better · ahead in 4 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced

ModelVeizikBaselineRatioSpreadHow each side was timed
Qwen3-1.7B968.0llama.cpp 1,276.5
qwen3-1.7B-Q4_K_M.gguf sha256 740b6373413…
llama.cpp 64e9bce (llama-bench, Metal)
1.32×——
Qwen2.5-0.5B343.6llama.cpp 391.9
qwen2.5-0.5B-Q4_K_M.gguf sha256 ca8c9ec09…
llama.cpp 64e9bce (llama-bench, Metal)
1.14×——
Qwen2.5-7B4,114.0llama.cpp 4,677.1
qwen2.5-7B-Q4_K_M.gguf sha256 2d6e42b4f8a…
llama.cpp 64e9bce (llama-bench, Metal)
1.14×——
Qwen2.5-1.5B926.9llama.cpp 980.1
qwen2.5-1.5B-Q4_K_M.gguf sha256 bc611b132…
llama.cpp 64e9bce (llama-bench, Metal)
1.06×——

Model weight footprint against MLX · MB, lower is better · ahead in 2 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced

ModelVeizikBaselineRatioSpreadHow each side was timed
Qwen2.5-7B4,114.0MLX 4,284.3
qwen2.5-7B-4bit (concatenated safetensors…
mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors)
1.04×——
Qwen3-1.7B968.0MLX 968.0
qwen3-1.7B-4bit-g64 (concatenated safeten…
mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors)
1.00×——
Qwen2.5-1.5B926.9MLX 868.5
qwen2.5-1.5B-4bit (concatenated safetenso…
mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors)
0.94×——
Qwen2.5-0.5B343.6MLX 278.0
qwen2.5-0.5B-4bit (concatenated safetenso…
mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors)
0.81×——

Eleven runs each, alternating between the two engines so that anything happening to the machine happens to both, and the middle run is the figure. The slowest and fastest run of each side are published beside it, so the figure is never the only thing you can see. The machine is checked quiet before each run — the GPU idle and the system's own media and screen-sharing daemons at rest — and the check is recorded on the row. Measuring through another program's work was most of the disagreement between runs on the machine these were taken on. ★A cell is published only when the result is not in question: our slowest run against their fastest, and our fastest against their slowest, both have to fall on the same side of parity. A cell where they do not is held, and that holds wins as readily as losses — it is the difference being smaller than the wobble, not the direction, that decides. Where the runs still disagree by more than 10%, the row names what makes them disagree. A wide spread that is understood and still decides the result is published with the spread shown; a wide spread nobody can account for is held. ★Our own figures wobble more than the other engine's on Apple M1 Ultra, and the row says by how much on both sides. It is the machine rather than either engine: a discrete speed state that lands on some of the work and cannot be switched off from outside. We publish the range instead of the median alone so that this is visible rather than averaged away.

Reproduce

Run it yourself.

Everything behind the Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-7B, Qwen3-1.7B 139 cells on Apple M1 Ultra and Apple M2 Pro: the prompts, the conversion, the benchmark script and what to read from the output. Each cell is a median of 11 runs per engine, alternating between them.

zsh
$ curl -fsSLO https://veizik.com/bench/veizik-bench-kit-pd1.tgz
$ shasum -a 256 veizik-bench-kit-pd1.tgz
# 7c5b91824f3328efb318b4c749efd6c50741ef166a1d17cec9219479f485cc8d
$ tar xzf veizik-bench-kit-pd1.tgz && cd repro_kit && cat README.md

279,402 bytes. The kit's own MANIFEST.sha256 covers every file inside it. The measurement log behind the published rows is served too: the raw run log and the rows built from it, so a figure here can be traced to the line it came from.

Which harness produced which rows. Each published row records the digest of the kit it was measured with, and they are not all the same kit: it was revised during the campaign. Every kit a published row names can be downloaded here.

★ This kit does not reproduce the other 12 cells on this page. Those were measured on the earlier prose files, which have no recorded source or licence — a reader cannot obtain them, which is the whole reason they are being replaced. Until each is re-measured with the kit, the figure stands for what it measured and nobody outside can check it.

Installed size

Installed size

16.3 MB

The whole installation: the CLI, four model cores and the signed core-id table. Summed from the published disk image, which is what the installer copies.

  • No dependency to install first — every library it links ships with macOS.
  • No Python, no framework, no compiler, no runtime of its own.
  • No converted or prepared copy of the model: it reads the original safetensors checkpoint.

macOS on Apple Silicon (arm64) · package-v1 · 2026-10-05

Time to a finished command

Time to a finished command

ModelDeviceTime to a finished commandTokens outProtocolRunsDateBuild
Qwen2.5-0.5BApple M4 Pro0.726 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra0.815 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro0.906 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra1.096 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra1.158 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra1.171 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M4 Pro1.312 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M4 Pro1.321 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M4 Pro1.374 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max1.425 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra1.658 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro1.772 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro1.812 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro1.833 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra2.002 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra2.116 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M4 Pro2.203 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Ultra2.498 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra2.951 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M4 Pro3.061 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro3.184 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max3.27 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra3.289 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M4 Pro3.448 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra3.659 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max4.027 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M4 Pro4.098 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro4.49 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra4.85 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro4.884 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro5.786 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max6.096 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Ultra6.101 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M1 Ultra6.345 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra6.409 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M4 Pro6.878 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro6.925 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra7.412 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro8.348 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro9.692 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra10.49 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max10.606 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Ultra10.824 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M4 Pro10.984 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro11.998 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra12.266 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max12.938 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M1 Ultra13.659 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M4 Pro14.873 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-0.5BApple M2 Pro16.175 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Ultra16.195 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro16.971 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro18.211 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max19.371 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M4 Pro21.666 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro21.831 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max22.3 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M4 Pro25.427 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M4 Pro28.245 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M1 Max29.449 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro31.307 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro31.69 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Ultra34.941 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro37.939 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max39.704 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-1.5BApple M2 Pro41.742 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M4 Pro45.1 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Ultra45.184 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen3-1.7BApple M2 Pro50.831 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max64.049 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro65.899 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M1 Max82.424 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro107.226 s128apple-wall-v1112026-10-062026.10.02-r6.1
Qwen2.5-7BApple M2 Pro138.579 s128apple-wall-v1112026-10-062026.10.02-r6.1

Timed from outside the command with a stopwatch, so everything you wait for is inside it: the process starting, the licence, the model loading, the tokens, and the exit.

The licence is leased into memory for a short window. The first run of that window pays for the lease; the runs that reuse it do not. Those are different experiences and they are different rows here, never averaged into one number.

Eleven runs each, on a quiet machine, and the middle one is shown. Anything whose runs disagreed by more than 5% is held back rather than shown. ★Every figure here is a run that reused a licence already in memory. The first run of a licence window also waits for that licence over the network, which takes seconds and varies too much from run to run to publish as a number, so those rows are held rather than shown. Those runs disagreed by 6 to 90%. What you wait for the very first time is longer than what is on this page, and this page does not yet say by how much.

Held back

Measured, not published

A result is published when it meets its protocol — every required field present, and for comparisons, both engines measured the same way and run-to-run spread inside the allowed bound. 1184 recorded results do not meet that bar, for the 13 reasons below. They stay in the dataset so nothing is lost, and appear the moment they do.

Why it is heldRowsAxesSpread
Held pending the r6.1 re-measure. (5 variants of this reason are recorded on the rows themselves)898Cross-machine output, Decode, Model weight footprint and 2 more—
One side's eleven runs disagreed by more than the 5% this protocol allows, so what the board has is a median rather than a measurement. (6 variants of this reason are recorded on the rows themselves)93Decode, Prefill, Time to a finished command5–166% measured, 5% allowed
The two engines did not read the same checkpoint: the baseline's file is the instruction-tuned or post-trained release and Veizik's is the base one. 56Decode, Prefill—
This protocol decides a cell by whether the two engines' slowest and fastest runs still agree on the winner, and that argument needs the 11 runs it was written for. (4 variants of this reason are recorded on the rows themselves)36Decode, Prefill—
The difference between the two engines is smaller than the disagreement between their own runs: on this cell our slowest run and our fastest run fall on opposite sides of the other engine's. 27Decode, Prefill—
A side's runs disagreed by more than 10% and nothing covering that side accounts for it. (21 variants of this reason are recorded on the rows themselves)22Decode, Prefill—
The baseline's generation rate was measured with an empty cache — its benchmark has no option to start from a filled one — while Veizik's was measured with the prompt already in. 16Decode—
A median needs runs to be the middle of. (2 variants of this reason are recorded on the rows themselves)16Decode, Prefill—
This is the baseline engine again, at its own default cache precision rather than the 16 bits the comparison rows use — a setting of the baseline, not of Veizik. 8Decode, Prefill—
A single run is not a measurement on this axis: with no repeat there is nothing to say how steady it was. (2 variants of this reason are recorded on the rows themselves)4Time to a finished command—
Measured with a different copy of the command than the one this build ships. 3Time to a finished command—
no generated count: the run printed no "generated" line (gen_tok empty/0; the model emitted EOS first; the shipped CLI has no ignore-EOS option), so output_tokens is 0 and the wall time is not a 128-token wall (2 variants of this reason are recorded on the rows themselves)3Decode, Time to a finished command—
Another program may have been running on this machine during these runs: none found (no in-run foreign watch in this harness; machine file activity checked for the cell window). 2Prefill—

One line per reason, not per row: the full reason, the device and the figures are recorded on every held row. Where the reason is run-to-run spread, the noise is in every engine on that board rather than only in Veizik — it needs an idle machine, not a different reading.

Method

Protocols

A protocol fixes what a result must carry before it can be published. A row missing any required field is held back rather than shown with a gap.

ProtocolWhat it measuresRequired fields
package-v1Package properties
Properties of the distributed archive itself.
artifact_id
apple-footprint-v1Apple device-memory board
How much device memory each engine's own copy of the model occupies, as each engine reports it: Veizik prints what it allocated for the weights, llama-bench prints the size of the file it loaded. Both figures are the weights alone and neither includes the growing cache a long conversation adds. There is no prompt and no timing here, so the figure does not move with the machine: the same four numbers came back identical on every Mac measured.
model_id, device_id, artifact_id, runs, baseline_engine, baseline_version, baseline_checkpoint, veizik_value, baseline_value, measured_binary_sha256, baseline_full_speed, precision_mode
apple-wall-v1Time from pressing return
The whole command, timed from the outside with a stopwatch: process start, licence, model load, generation, exit. This is the only figure on the site that is what you actually wait for. The engine boards are deliberately not this — they time the engine, and this times your experience of it.
model_id, device_id, artifact_id, runs, value, engine_key, phase, spread_pct, measured_binary_sha256, measured_from
apple-compare-v2Apple comparison board
The same comparison as the protocol above, measured on a machine checked quiet first and published with the full range of what every run did. It exists because a tight measurement and a decided result are not the same thing: a cell is shown here when the difference between the two engines is larger than the measurement's own wobble, and held when it is not — whichever way it would have gone.
model_id, device_id, artifact_id, runs, veizik_value, baseline_value, baseline_engine, baseline_version, unit, veizik_spread_pct, baseline_spread_pct, veizik_min, veizik_max, baseline_min, baseline_max, measured_binary_sha256, measured_from, checkpoint_pairing, quiet_gate

Cross-engine comparisons require the competing engine's name and exact version in the same row. Rows that carry them are published here; rows that do not are listed above as held back.