Veizik performance, measured on real hardware.
Every row below is one measurement on a named device, under a named protocol. Each protocol lists what a row has to carry; a row missing any of it is held back rather than shown. Results are added as rows, not as prose.
- Apple Silicon devices
- 4
- comparison cells, current prose
- 139
- competing runtimes
- 2
- reproducible protocols
- 4
A further 12 comparison cells are published below, measured on an earlier set of prose that is being replaced. They are not counted in the figure above, because a cell on one set of prose and a cell on another are not two measurements of the same thing.
Against other engines, on the same machine
Decode and Prefill against MLX and MLX bf16, on Apple M2 Pro and Apple M1 Ultra. Veizik is ahead in 134 of 139 cells, and the cells it is behind on are here too. A further 12 cells, on Apple M1 Ultra and Apple M4 Pro, were measured on the earlier prose files rather than this one; they are in their own tables below, and none of them is counted in the figure above. Qwen2.5-0.5B, Qwen2.5-1.5B and Qwen2.5-7B are marked a lower bound. On those models the engine's prefill clock stops before the last prompt step has finished, and that step is charged to decode; these figures put a whole decode step back in, which costs more than the step does. So the engine is at least this fast on both axes and may be faster: a win here is safe, while a loss, or a cell held because the two ranges overlap, may be understating the engine rather than describing it. Over Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-7B, Qwen3-1.7B. Each board below names its own unit, which direction is better, and how many runs it is a median of.
Apple M4 Pro
Model weight footprint against llama.cpp · MB, lower is better · ahead in 4 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced
| Model | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|
| Qwen3-1.7B | 968.0 | llama.cpp 1,276.5 Q4_K_M GGUF 0.25.3 / build 11146 (7fe450e19), Metal backend | — | — | |
| Qwen2.5-0.5B | 343.6 | llama.cpp 391.9 Q4_K_M GGUF 0.25.3 / build 11146 (7fe450e19), Metal backend | — | — | |
| Qwen2.5-7B | 4,114.0 | llama.cpp 4,677.1 Q4_K_M GGUF 0.25.3 / build 11146 (7fe450e19), Metal backend | — | — | |
| Qwen2.5-1.5B | 926.9 | llama.cpp 980.1 Q4_K_M GGUF 0.25.3 / build 11146 (7fe450e19), Metal backend | — | — |
Apple M2 Pro
Decode against MLX · tok/s, higher is better · ahead in 19 of 24 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 30,000 → 128 | 110.3 109.8–110.9 | MLX 74.7 74.5–74.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 24,576 → 128 | 120.9 120.3–121.3 | MLX 85.6 85.5–85.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 16,384 → 128 | 141.7 139.8–143.2 | MLX 109.5 109.4–109.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 8,192 → 128 | 192.8 189.2–195.6 | MLX 151.1 150.7–151.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 512 → 128 | 300.6 299.0–302.5 | MLX 256.2 255.9–256.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 3,968 → 128 | 222.0 221.5–223.4 | MLX 191.1 190.4–192.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.9% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 30,000 → 128 | 67.8 67.5–68.2 | MLX 59.4 59.3–59.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 24,576 → 128 | 74.1 73.8–74.5 | MLX 67.1 66.9–67.2 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 512 → 128 | 143.8 142.5–144.4 | MLX 133.4 133.2–133.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock | |
| Qwen3-1.7B | 16,384 → 128 | 60.1 60.0–60.2 | MLX 57.3 57.2–57.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen2.5-1.5B | 16,384 → 128 | 86.3 85.8–86.5 | MLX 82.6 82.5–82.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 24,576 → 128 | 46.4 46.2–46.4 | MLX 44.5 44.5–44.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock | |
| Qwen3-1.7B | 3,968 → 128 | 106.6 106.3–106.8 | MLX 102.6 102.5–102.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen2.5-7B | 30,000 → 128 | 23.5 23.5–23.6 | MLX 22.7 22.7–22.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 8,192 → 128 | 83.9 83.6–84.2 | MLX 80.9 80.8–80.9 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen2.5-7B | 24,576 → 128 | 25.3 25.2–25.3 | MLX 24.7 24.7–24.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 30,000 → 128 | 40.3 40.2–40.3 | MLX 39.4 39.3–39.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock | |
| Qwen2.5-1.5B | 512 → 128 | 144.6 144.3–146.6 | MLX 142.1 141.8–142.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 16,384 → 128 | 28.5 28.4–28.6 | MLX 28.3 28.3–28.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 8,192 → 128 | 102.4 101.6–103.6 | MLX 104.1 103.9–104.2 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 8,192 → 128 | 32.3 32.1–32.5 | MLX 33.1 33.1–33.2 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 3,968 → 128 | 113.7 109.6–116.6 | MLX 119.2 119.1–119.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 512 → 128 | 38.3 38.0–38.4 | MLX 40.6 40.5–40.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 3,968 → 128 | 33.0 32.8–33.1 | MLX 36.6 36.6–36.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast |
Decode against MLX bf16 · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-7B | 512 → 128 | 38.3 38.0–38.4 | MLX bf16 12.7 12.7–12.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 512 → 128 | 143.8 142.5–144.4 | MLX bf16 48.9 48.8–49.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 8,192 → 128 | 32.3 32.1–32.5 | MLX bf16 11.9 11.8–11.9 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 3,968 → 128 | 33.0 32.8–33.1 | MLX bf16 12.3 12.3–12.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 512 → 128 | 144.6 144.3–146.6 | MLX bf16 54.6 54.5–54.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 16,384 → 128 | 28.5 28.4–28.6 | MLX bf16 11.2 11.2–11.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 24,576 → 128 | 25.3 25.2–25.3 | MLX bf16 10.5 10.5–10.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 3,968 → 128 | 106.6 106.3–106.8 | MLX bf16 44.5 44.4–44.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 30,000 → 128 | 23.5 23.5–23.6 | MLX bf16 10.2 10.2–10.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 3,968 → 128 | 113.7 109.6–116.6 | MLX bf16 50.8 50.7–50.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 8,192 → 128 | 102.4 101.6–103.6 | MLX bf16 48.1 48.1–48.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 8,192 → 128 | 83.9 83.6–84.2 | MLX bf16 39.7 39.6–39.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 16,384 → 128 | 86.3 85.8–86.5 | MLX bf16 43.1 43.0–43.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 512 → 128 | 300.6 299.0–302.5 | MLX bf16 150.9 150.8–151.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 24,576 → 128 | 74.1 73.8–74.5 | MLX bf16 38.4 38.4–38.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 30,000 → 128 | 67.8 67.5–68.2 | MLX bf16 35.8 35.7–35.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 16,384 → 128 | 60.1 60.0–60.2 | MLX bf16 33.1 33.0–33.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 8,192 → 128 | 192.8 189.2–195.6 | MLX bf16 109.8 109.5–110.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 3,968 → 128 | 222.0 221.5–223.4 | MLX bf16 127.0 126.4–127.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 24,576 → 128 | 46.4 46.2–46.4 | MLX bf16 28.5 28.4–28.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 16,384 → 128 | 141.7 139.8–143.2 | MLX bf16 90.7 90.6–90.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 30,000 → 128 | 40.3 40.2–40.3 | MLX bf16 26.1 26.1–26.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 24,576 → 128 | 120.9 120.3–121.3 | MLX bf16 80.4 80.2–80.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 30,000 → 128 | 110.3 109.8–110.9 | MLX bf16 74.6 74.6–74.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint |
Prefill against MLX · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 512 → 128 | 5,059.7 4,985.3–5,101.3 | MLX 3,274.5 3,195.9–3,295.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 3.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 512 → 128 | 1,459.1 1,450.2–1,461.4 | MLX 988.6 981.0–995.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.5% | both sides: the engine's own clock | |
| Qwen2.5-1.5B | 512 → 128 | 1,505.5 1,497.9–1,510.8 | MLX 1,049.5 1,042.2–1,056.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 512 → 128 | 333.1 332.6–333.5 | MLX 232.9 231.9–233.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 3,968 → 128 | 327.2 326.9–327.3 | MLX 230.4 230.2–230.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 8,192 → 128 | 306.8 306.8–307.0 | MLX 217.9 217.7–218.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 16,384 → 128 | 271.9 271.9–272.0 | MLX 196.4 196.4–196.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 3,968 → 128 | 1,447.8 1,446.9–1,448.3 | MLX 1,048.2 1,046.6–1,050.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 3,968 → 128 | 1,305.8 1,305.0–1,306.5 | MLX 947.4 945.1–949.1 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock | |
| Qwen2.5-7B | 24,576 → 128 | 243.6 243.6–243.7 | MLX 178.3 178.3–178.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 3,968 → 128 | 4,664.4 4,656.9–4,673.1 | MLX 3,417.0 3,379.9–3,456.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 2.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 8,192 → 128 | 1,272.9 1,272.5–1,273.1 | MLX 937.9 936.4–939.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-7B | 30,000 → 128 | 227.6 227.6–227.7 | MLX 167.8 167.7–167.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 8,192 → 128 | 1,112.0 1,111.5–1,112.5 | MLX 825.7 824.4–829.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock | |
| Qwen2.5-0.5B | 8,192 → 128 | 3,916.5 3,913.8–3,919.2 | MLX 2,939.7 2,928.8–2,959.2 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.0% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 16,384 → 128 | 880.2 879.5–880.8 | MLX 662.5 661.7–662.9 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen2.5-1.5B | 16,384 → 128 | 1,024.8 1,024.7–1,025.1 | MLX 774.9 773.5–775.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen3-1.7B | 24,576 → 128 | 727.0 726.7–727.4 | MLX 550.3 550.1–551.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen3-1.7B | 30,000 → 128 | 651.1 650.8–651.6 | MLX 495.2 495.0–495.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock | |
| Qwen2.5-1.5B | 24,576 → 128 | 854.5 854.3–854.7 | MLX 656.1 655.5–657.1 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 30,000 → 128 | 769.6 769.5–769.7 | MLX 593.8 593.1–594.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 16,384 → 128 | 2,947.8 2,946.4–2,948.2 | MLX 2,309.8 2,304.8–2,313.9 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 24,576 → 128 | 2,356.1 2,355.0–2,357.1 | MLX 1,897.5 1,892.7–1,901.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 30,000 → 128 | 2,071.3 2,070.2–2,071.9 | MLX 1,696.5 1,694.4–1,699.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast |
Prefill against MLX bf16 · tok/s, higher is better · ahead in 24 of 24 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 512 → 128 | 5,059.7 4,985.3–5,101.3 | MLX bf16 3,944.8 3,917.3–4,015.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 2.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 512 → 128 | 1,459.1 1,450.2–1,461.4 | MLX bf16 1,218.9 1,209.7–1,228.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.6% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 30,000 → 128 | 651.1 650.8–651.6 | MLX bf16 566.9 566.5–567.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 512 → 128 | 333.1 332.6–333.5 | MLX bf16 293.7 292.8–295.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.9% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 24,576 → 128 | 727.0 726.7–727.4 | MLX bf16 641.2 640.7–641.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 512 → 128 | 1,505.5 1,497.9–1,510.8 | MLX bf16 1,328.7 1,316.6–1,337.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 30,000 → 128 | 769.6 769.5–769.7 | MLX bf16 692.5 691.9–693.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 16,384 → 128 | 880.2 879.5–880.8 | MLX bf16 798.3 797.6–799.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 24,576 → 128 | 854.5 854.3–854.7 | MLX bf16 779.0 778.3–779.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 30,000 → 128 | 227.6 227.6–227.7 | MLX bf16 208.8 208.7–208.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 24,576 → 128 | 2,356.1 2,355.0–2,357.1 | MLX bf16 2,170.9 2,163.0–2,176.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 30,000 → 128 | 2,071.3 2,070.2–2,071.9 | MLX bf16 1,909.1 1,903.4–1,914.9 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 16,384 → 128 | 2,947.8 2,946.4–2,948.2 | MLX bf16 2,722.1 2,718.3–2,731.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 24,576 → 128 | 243.6 243.6–243.7 | MLX bf16 225.3 225.2–225.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 16,384 → 128 | 1,024.8 1,024.7–1,025.1 | MLX bf16 951.0 949.8–951.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 8,192 → 128 | 3,916.5 3,913.8–3,919.2 | MLX bf16 3,667.8 3,654.3–3,670.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 16,384 → 128 | 271.9 271.9–272.0 | MLX bf16 254.8 254.7–255.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 8,192 → 128 | 1,112.0 1,111.5–1,112.5 | MLX bf16 1,047.2 1,045.9–1,048.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 8,192 → 128 | 1,272.9 1,272.5–1,273.1 | MLX bf16 1,203.8 1,202.6–1,207.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 3,968 → 128 | 4,664.4 4,656.9–4,673.1 | MLX bf16 4,413.9 4,376.2–4,429.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 8,192 → 128 | 306.8 306.8–307.0 | MLX bf16 291.3 291.0–291.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen3-1.7B | 3,968 → 128 | 1,305.8 1,305.0–1,306.5 | MLX bf16 1,241.0 1,238.9–1,243.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-7B | 3,968 → 128 | 327.2 326.9–327.3 | MLX bf16 312.1 311.4–312.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 3,968 → 128 | 1,447.8 1,446.9–1,448.3 | MLX bf16 1,385.3 1,383.0–1,391.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint |
Apple M1 Ultra
Decode against MLX · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 16,384 → 128 | 259.1 257.4–262.0 | MLX 177.7 177.0–178.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.8% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 512 → 128 | 377.4 355.8–385.6 | MLX 272.3 271.0–273.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.0% | both sides: the engine's own clock lower bound: veizik is at least this fast Re-measured: pre-pause environment. | |
| Qwen2.5-0.5B | 8,192 → 128 | 280.6 278.7–286.6 | MLX 203.6 202.9–204.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 30,000 → 128 | 199.0 196.9–200.3 | MLX 144.5 144.2–145.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.1% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 24,576 → 128 | 205.6 204.4–207.8 | MLX 155.6 155.3–157.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.6% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 3,968 → 128 | 305.5 304.4–313.3 | MLX 238.9 238.6–239.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 3,968 → 128 | 233.9 215.3–241.2 | MLX 188.5 182.6–190.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 4.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 8,192 → 128 | 201.0 200.2–204.0 | MLX 164.0 163.7–166.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.8% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 512 → 128 | 260.2 244.7–261.2 | MLX 213.5 213.3–214.0 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 16,384 → 128 | 169.1 168.1–170.5 | MLX 146.0 145.8–146.8 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 24,576 → 128 | 145.8 144.8–147.6 | MLX 127.6 125.8–128.2 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.9% | both sides: the engine's own clock lower bound: veizik is at least this fast |
Decode against MLX bf16 · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-1.5B | 512 → 128 | 260.2 244.7–261.2 | MLX bf16 126.9 126.8–127.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 3,968 → 128 | 233.9 215.3–241.2 | MLX bf16 117.8 94.3–118.5 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 25.7% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 8,192 → 128 | 201.0 200.2–204.0 | MLX bf16 106.8 106.6–107.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 16,384 → 128 | 169.1 168.1–170.5 | MLX bf16 99.5 98.9–99.9 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.1% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 24,576 → 128 | 145.8 144.8–147.6 | MLX bf16 90.4 90.1–91.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 512 → 128 | 377.4 355.8–385.6 | MLX bf16 248.5 248.2–249.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.4% | both sides: the engine's own clock lower bound: veizik is at least this fast Re-measured: pre-pause environment. | |
| Qwen2.5-0.5B | 16,384 → 128 | 259.1 257.4–262.0 | MLX bf16 172.3 171.9–175.1 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.8% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 8,192 → 128 | 280.6 278.7–286.6 | MLX bf16 193.4 192.7–193.9 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 3,968 → 128 | 305.5 304.4–313.3 | MLX bf16 222.3 222.0–222.6 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 30,000 → 128 | 199.0 196.9–200.3 | MLX bf16 145.9 145.7–146.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 24,576 → 128 | 205.6 204.4–207.8 | MLX bf16 156.3 155.9–156.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint |
Prefill against MLX · tok/s, higher is better · ahead in 11 of 11 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 512 → 128 | 9,631.8 9,087.7–10,740.8 | MLX 4,684.7 4,556.1–4,885.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 7.2% accounted for | both sides: the engine's own clock lower bound: veizik is at least this fast Re-measured: pre-pause environment. | |
| Qwen2.5-1.5B | 512 → 128 | 4,032.0 3,944.9–4,038.6 | MLX 2,487.6 2,448.8–2,532.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 3.4% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 3,968 → 128 | 13,075.1 12,868.0–13,273.7 | MLX 9,859.3 9,809.9–9,924.3 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 24,576 → 128 | 2,795.8 2,737.8–2,800.7 | MLX 2,132.9 2,123.5–2,134.6 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.5% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 16,384 → 128 | 3,309.3 3,302.2–3,314.0 | MLX 2,532.2 2,518.3–2,536.1 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 3,968 → 128 | 4,469.7 4,324.8–4,478.4 | MLX 3,426.6 3,404.8–3,437.7 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.0% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-1.5B | 8,192 → 128 | 4,021.3 3,975.1–4,042.7 | MLX 3,091.9 3,060.5–3,101.4 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.3% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 8,192 → 128 | 11,327.6 11,287.5–11,399.0 | MLX 9,232.8 9,173.2–9,343.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.9% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 16,384 → 128 | 8,509.2 8,488.2–8,525.7 | MLX 7,486.9 7,422.1–7,510.5 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 1.2% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 24,576 → 128 | 6,773.1 6,764.9–6,777.1 | MLX 6,231.3 6,195.5–6,248.1 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.8% | both sides: the engine's own clock lower bound: veizik is at least this fast | |
| Qwen2.5-0.5B | 30,000 → 128 | 5,895.2 5,891.9–5,899.0 | MLX 5,587.5 5,569.4–5,600.1 4-bit g64, converted from the snapshot mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast |
Prefill against MLX bf16 · tok/s, higher is better · ahead in 10 of 10 · median of 11 runs · on Darwin, On the Origin of Species
| Model | Tokens | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | 512 → 128 | 9,631.8 9,087.7–10,740.8 | MLX bf16 4,980.0 4,061.3–5,160.7 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 27.1% accounted for | both sides: the engine's own clock lower bound: veizik is at least this fast Re-measured: pre-pause environment. | |
| Qwen2.5-1.5B | 512 → 128 | 4,032.0 3,944.9–4,038.6 | MLX bf16 2,688.4 2,603.2–2,739.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 5.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 3,968 → 128 | 13,075.1 12,868.0–13,273.7 | MLX bf16 10,792.2 10,660.6–10,891.0 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 2.2% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 24,576 → 128 | 2,795.8 2,737.8–2,800.7 | MLX bf16 2,332.7 2,320.9–2,335.3 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 16,384 → 128 | 3,309.3 3,302.2–3,314.0 | MLX bf16 2,826.1 2,812.7–2,832.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 8,192 → 128 | 4,021.3 3,975.1–4,042.7 | MLX bf16 3,558.6 3,476.4–3,567.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 2.6% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 8,192 → 128 | 11,327.6 11,287.5–11,399.0 | MLX bf16 10,187.9 10,097.6–10,253.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.5% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-1.5B | 3,968 → 128 | 4,469.7 4,324.8–4,478.4 | MLX bf16 4,023.6 4,003.4–4,053.8 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.3% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 16,384 → 128 | 8,509.2 8,488.2–8,525.7 | MLX bf16 8,076.9 7,995.0–8,107.2 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 1.4% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint | |
| Qwen2.5-0.5B | 24,576 → 128 | 6,773.1 6,764.9–6,777.1 | MLX bf16 6,621.0 6,578.9–6,627.4 bf16, the base snapshot unquantised mlx-lm/mlx 0.31.3 0.32.2 | 0.7% | both sides: the engine's own clock lower bound: veizik is at least this fast the baseline read the published model's own files unquantised rather than a copy converted for it, which is the closest two engines can come to opening the same checkpoint |
Model weight footprint against llama.cpp · MB, lower is better · ahead in 4 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced
| Model | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|
| Qwen3-1.7B | 968.0 | llama.cpp 1,276.5 qwen3-1.7B-Q4_K_M.gguf sha256 740b6373413… llama.cpp 64e9bce (llama-bench, Metal) | — | — | |
| Qwen2.5-0.5B | 343.6 | llama.cpp 391.9 qwen2.5-0.5B-Q4_K_M.gguf sha256 ca8c9ec09… llama.cpp 64e9bce (llama-bench, Metal) | — | — | |
| Qwen2.5-7B | 4,114.0 | llama.cpp 4,677.1 qwen2.5-7B-Q4_K_M.gguf sha256 2d6e42b4f8a… llama.cpp 64e9bce (llama-bench, Metal) | — | — | |
| Qwen2.5-1.5B | 926.9 | llama.cpp 980.1 qwen2.5-1.5B-Q4_K_M.gguf sha256 bc611b132… llama.cpp 64e9bce (llama-bench, Metal) | — | — |
Model weight footprint against MLX · MB, lower is better · ahead in 2 of 4 · one run; the figure does not vary · on the earlier prose files, being replaced
| Model | Veizik | Baseline | Ratio | Spread | How each side was timed |
|---|---|---|---|---|---|
| Qwen2.5-7B | 4,114.0 | MLX 4,284.3 qwen2.5-7B-4bit (concatenated safetensors… mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors) | — | — | |
| Qwen3-1.7B | 968.0 | MLX 968.0 qwen3-1.7B-4bit-g64 (concatenated safeten… mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors) | — | — | |
| Qwen2.5-1.5B | 926.9 | MLX 868.5 qwen2.5-1.5B-4bit (concatenated safetenso… mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors) | — | — | |
| Qwen2.5-0.5B | 343.6 | MLX 278.0 qwen2.5-0.5B-4bit (concatenated safetenso… mlx-lm 0.31.3 / mlx 0.32.2 (weights figure is engine-independent: tensor bytes of the loaded safetensors) | — | — |
Eleven runs each, alternating between the two engines so that anything happening to the machine happens to both, and the middle run is the figure. The slowest and fastest run of each side are published beside it, so the figure is never the only thing you can see. The machine is checked quiet before each run — the GPU idle and the system's own media and screen-sharing daemons at rest — and the check is recorded on the row. Measuring through another program's work was most of the disagreement between runs on the machine these were taken on. ★A cell is published only when the result is not in question: our slowest run against their fastest, and our fastest against their slowest, both have to fall on the same side of parity. A cell where they do not is held, and that holds wins as readily as losses — it is the difference being smaller than the wobble, not the direction, that decides. Where the runs still disagree by more than 10%, the row names what makes them disagree. A wide spread that is understood and still decides the result is published with the spread shown; a wide spread nobody can account for is held. ★Our own figures wobble more than the other engine's on Apple M1 Ultra, and the row says by how much on both sides. It is the machine rather than either engine: a discrete speed state that lands on some of the work and cannot be switched off from outside. We publish the range instead of the median alone so that this is visible rather than averaged away.
Run it yourself.
Everything behind the Qwen2.5-0.5B, Qwen2.5-1.5B, Qwen2.5-7B, Qwen3-1.7B 139 cells on Apple M1 Ultra and Apple M2 Pro: the prompts, the conversion, the benchmark script and what to read from the output. Each cell is a median of 11 runs per engine, alternating between them.
- The prompts are public domain. Project Gutenberg eBook #1228, public domain, cut to exactly the token count each cell names, for both tokenizers. The kit carries them with a SHA256SUMS, so you run the same bytes we did.
- The other engine is run with its end-of-text token suppressed so that both engines generate the same number of tokens. Nothing else about it is changed: same weights, same sampler, same cache, its own generate().
- Ours is not suppressed, because the shipped command has no such switch. Every row records the tokens actually generated, and a cell where our engine stopped early shows that count rather than a forced one.
- The machine is checked quiet before each run, not once at the start, and each row carries what the check saw. A row whose own readings show the machine was busy is held back rather than shown. The runs behind these cells recorded the machine idling at 0%, and the kit now refuses to start at all if it measures above 20%.
$ curl -fsSLO https://veizik.com/bench/veizik-bench-kit-pd1.tgz $ shasum -a 256 veizik-bench-kit-pd1.tgz # 7c5b91824f3328efb318b4c749efd6c50741ef166a1d17cec9219479f485cc8d $ tar xzf veizik-bench-kit-pd1.tgz && cd repro_kit && cat README.md
279,402 bytes. The kit's own MANIFEST.sha256 covers every file inside it. The measurement log behind the published rows is served too: the raw run log and the rows built from it, so a figure here can be traced to the line it came from.
Which harness produced which rows. Each published row records the digest of the kit it was measured with, and they are not all the same kit: it was revised during the campaign. Every kit a published row names can be downloaded here.
- 7c5b91824f3328efb318b4c749efd6c50741ef166a1d17cec9219479f485cc8d — 213 rows, served here
★ This kit does not reproduce the other 12 cells on this page. Those were measured on the earlier prose files, which have no recorded source or licence — a reader cannot obtain them, which is the whole reason they are being replaced. Until each is re-measured with the kit, the figure stands for what it measured and nobody outside can check it.
Installed size
The whole installation: the CLI, four model cores and the signed core-id table. Summed from the published disk image, which is what the installer copies.
- No dependency to install first — every library it links ships with macOS.
- No Python, no framework, no compiler, no runtime of its own.
- No converted or prepared copy of the model: it reads the original safetensors checkpoint.
macOS on Apple Silicon (arm64) · package-v1 · 2026-10-05
Time to a finished command
| Model | Device | Time to a finished command | Tokens out | Protocol | Runs | Date | Build |
|---|---|---|---|---|---|---|---|
| Qwen2.5-0.5B | Apple M4 Pro | 0.726 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 0.815 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 0.906 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 1.096 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 1.158 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 1.171 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M4 Pro | 1.312 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M4 Pro | 1.321 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M4 Pro | 1.374 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 1.425 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 1.658 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 1.772 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 1.812 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 1.833 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 2.002 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 2.116 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M4 Pro | 2.203 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Ultra | 2.498 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 2.951 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M4 Pro | 3.061 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 3.184 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 3.27 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 3.289 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M4 Pro | 3.448 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 3.659 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 4.027 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M4 Pro | 4.098 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 4.49 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 4.85 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 4.884 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 5.786 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 6.096 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Ultra | 6.101 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M1 Ultra | 6.345 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 6.409 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M4 Pro | 6.878 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 6.925 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 7.412 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 8.348 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 9.692 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 10.49 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 10.606 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Ultra | 10.824 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M4 Pro | 10.984 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 11.998 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 12.266 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 12.938 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M1 Ultra | 13.659 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M4 Pro | 14.873 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-0.5B | Apple M2 Pro | 16.175 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Ultra | 16.195 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 16.971 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 18.211 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 19.371 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M4 Pro | 21.666 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 21.831 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 22.3 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M4 Pro | 25.427 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M4 Pro | 28.245 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M1 Max | 29.449 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 31.307 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 31.69 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Ultra | 34.941 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 37.939 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 39.704 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-1.5B | Apple M2 Pro | 41.742 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M4 Pro | 45.1 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Ultra | 45.184 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen3-1.7B | Apple M2 Pro | 50.831 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 64.049 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 65.899 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M1 Max | 82.424 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 107.226 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
| Qwen2.5-7B | Apple M2 Pro | 138.579 s | 128 | apple-wall-v1 | 11 | 2026-10-06 | 2026.10.02-r6.1 |
Timed from outside the command with a stopwatch, so everything you wait for is inside it: the process starting, the licence, the model loading, the tokens, and the exit.
The licence is leased into memory for a short window. The first run of that window pays for the lease; the runs that reuse it do not. Those are different experiences and they are different rows here, never averaged into one number.
Eleven runs each, on a quiet machine, and the middle one is shown. Anything whose runs disagreed by more than 5% is held back rather than shown. ★Every figure here is a run that reused a licence already in memory. The first run of a licence window also waits for that licence over the network, which takes seconds and varies too much from run to run to publish as a number, so those rows are held rather than shown. Those runs disagreed by 6 to 90%. What you wait for the very first time is longer than what is on this page, and this page does not yet say by how much.
Measured, not published
A result is published when it meets its protocol — every required field present, and for comparisons, both engines measured the same way and run-to-run spread inside the allowed bound. 1184 recorded results do not meet that bar, for the 13 reasons below. They stay in the dataset so nothing is lost, and appear the moment they do.
| Why it is held | Rows | Axes | Spread |
|---|---|---|---|
| Held pending the r6.1 re-measure. (5 variants of this reason are recorded on the rows themselves) | 898 | Cross-machine output, Decode, Model weight footprint and 2 more | — |
| One side's eleven runs disagreed by more than the 5% this protocol allows, so what the board has is a median rather than a measurement. (6 variants of this reason are recorded on the rows themselves) | 93 | Decode, Prefill, Time to a finished command | 5–166% measured, 5% allowed |
| The two engines did not read the same checkpoint: the baseline's file is the instruction-tuned or post-trained release and Veizik's is the base one. | 56 | Decode, Prefill | — |
| This protocol decides a cell by whether the two engines' slowest and fastest runs still agree on the winner, and that argument needs the 11 runs it was written for. (4 variants of this reason are recorded on the rows themselves) | 36 | Decode, Prefill | — |
| The difference between the two engines is smaller than the disagreement between their own runs: on this cell our slowest run and our fastest run fall on opposite sides of the other engine's. | 27 | Decode, Prefill | — |
| A side's runs disagreed by more than 10% and nothing covering that side accounts for it. (21 variants of this reason are recorded on the rows themselves) | 22 | Decode, Prefill | — |
| The baseline's generation rate was measured with an empty cache — its benchmark has no option to start from a filled one — while Veizik's was measured with the prompt already in. | 16 | Decode | — |
| A median needs runs to be the middle of. (2 variants of this reason are recorded on the rows themselves) | 16 | Decode, Prefill | — |
| This is the baseline engine again, at its own default cache precision rather than the 16 bits the comparison rows use — a setting of the baseline, not of Veizik. | 8 | Decode, Prefill | — |
| A single run is not a measurement on this axis: with no repeat there is nothing to say how steady it was. (2 variants of this reason are recorded on the rows themselves) | 4 | Time to a finished command | — |
| Measured with a different copy of the command than the one this build ships. | 3 | Time to a finished command | — |
| no generated count: the run printed no "generated" line (gen_tok empty/0; the model emitted EOS first; the shipped CLI has no ignore-EOS option), so output_tokens is 0 and the wall time is not a 128-token wall (2 variants of this reason are recorded on the rows themselves) | 3 | Decode, Time to a finished command | — |
| Another program may have been running on this machine during these runs: none found (no in-run foreign watch in this harness; machine file activity checked for the cell window). | 2 | Prefill | — |
One line per reason, not per row: the full reason, the device and the figures are recorded on every held row. Where the reason is run-to-run spread, the noise is in every engine on that board rather than only in Veizik — it needs an idle machine, not a different reading.
Protocols
A protocol fixes what a result must carry before it can be published. A row missing any required field is held back rather than shown with a gap.
| Protocol | What it measures | Required fields |
|---|---|---|
| package-v1 | Package properties Properties of the distributed archive itself. | artifact_id |
| apple-footprint-v1 | Apple device-memory board How much device memory each engine's own copy of the model occupies, as each engine reports it: Veizik prints what it allocated for the weights, llama-bench prints the size of the file it loaded. Both figures are the weights alone and neither includes the growing cache a long conversation adds. There is no prompt and no timing here, so the figure does not move with the machine: the same four numbers came back identical on every Mac measured. | model_id, device_id, artifact_id, runs, baseline_engine, baseline_version, baseline_checkpoint, veizik_value, baseline_value, measured_binary_sha256, baseline_full_speed, precision_mode |
| apple-wall-v1 | Time from pressing return The whole command, timed from the outside with a stopwatch: process start, licence, model load, generation, exit. This is the only figure on the site that is what you actually wait for. The engine boards are deliberately not this — they time the engine, and this times your experience of it. | model_id, device_id, artifact_id, runs, value, engine_key, phase, spread_pct, measured_binary_sha256, measured_from |
| apple-compare-v2 | Apple comparison board The same comparison as the protocol above, measured on a machine checked quiet first and published with the full range of what every run did. It exists because a tight measurement and a decided result are not the same thing: a cell is shown here when the difference between the two engines is larger than the measurement's own wobble, and held when it is not — whichever way it would have gone. | model_id, device_id, artifact_id, runs, veizik_value, baseline_value, baseline_engine, baseline_version, unit, veizik_spread_pct, baseline_spread_pct, veizik_min, veizik_max, baseline_min, baseline_max, measured_binary_sha256, measured_from, checkpoint_pairing, quiet_gate |
Cross-engine comparisons require the competing engine's name and exact version in the same row. Rows that carry them are published here; rows that do not are listed above as held back.