Product Models Benchmarks Docs Releases Sign in Download GitHub
Native AI inference

Veizik. The New Basic of AI.

Run the checkpoint you already have.

One compact native engine that opens your safetensors directly — on your own hardware, weights untouched.

Language runs today. FLUX.1 is coming next.

A compact graphite computer enclosure in a dark room, lit by a single thin light line along one edge.
zsh
$ veizik run qwen2.5-7b "The capital of France is"
 Paris

Image rendered by Veizik 1.0.0 (sealed bundle, linux x86_64) · SD3.5-medium · NVIDIA RTX 3090 · Powered by Stability AI

What it runs

One engine, three model classes.

Language, image and video share the same native runtime.

How to read this section

Each class names the hardware target it executes on today. Where the public catalogue has no entries yet, it says so rather than leaving you to guess from an absence — the engine is built for all three, and this page only claims what you can run right now.

Image

Still-image generation from text and from an input image.

NVIDIA (CUDA)

FLUX.1 then Stable Diffusion 3.5 →
The engine also runs FLUX.2, Chroma, HiDream-I1.

Video

Text-to-video and image-to-video generation.

NVIDIA (CUDA)

Runs on this target in the engine: LTX-Video, Wan 2.1, Wan 2.2, CogVideoX.

Language

Text generation, read straight from safetensors checkpoints.

Apple Silicon (Metal)

8 models in this build →

Why Veizik

Nothing was traded away for the small install.

A dependency-free runtime that opens the checkpoint you already have is usually the version you settle for when you want a simple install. This one is 16.3 MB, links nothing but libraries the operating system already ships, and writes no second copy of your model — and on 2 Macs, carrying 96 and 43 cells, measured against MLX and MLX bf16, it is ahead in 134 of 139 cells. A further 12 cells, measured on an earlier set of prose, are being replaced and kept separate on the board rather than counted here. 489 further cells measured on the build this site no longer offers are held back rather than counted, because a rate is only true of the bytes it was timed on. Qwen2.5-0.5B on Apple M1 Max, Qwen2.5-1.5B on Apple M1 Max, Qwen2.5-7B on Apple M1 Max, Qwen3-1.7B on Apple M1 Max, Qwen2.5-7B on Apple M1 Ultra, Qwen3-1.7B on Apple M1 Ultra, Qwen2.5-0.5B on Apple M4 Pro, Qwen2.5-1.5B on Apple M4 Pro, Qwen2.5-7B on Apple M4 Pro and Qwen3-1.7B on Apple M4 Pro have no published cell on this build.

What that changes in practice

Installation stays small and predictable: one signed package, and the libraries it links are the ones the operating system already ships. Deployment is the same file on every machine, so what you tested is what runs. And the weights stay where you put them — generation happens on your own hardware.

Compiled, and talking to the device

A compiled runtime talks straight to the device. What you install is the thing that runs, and everything it links ships with the operating system.

No conversion step, ever

Veizik opens supported safetensors checkpoints directly. There is no export to run, no second file to keep in step with the first: the engine builds its own in-memory representation as it loads, so a model is ready the moment the download is.

One runtime, built to carry more than language

The same engine architecture is built for language, image, and video inference instead of assembling a different framework stack for each pipeline.

Local by design

Generation runs on your hardware. Prompts, inputs, weights and outputs stay where you put them. The device is activated once; after that a licence lease is held in memory, and runs that reuse it need no connection at all.

Measured

Every comparison we ran, published — including the ones we lose.

1 other engine, 2 Macs (96 / 43 cells), 4 models, 139 measured cells. Veizik is ahead in 134. A further 12 cells, measured on an earlier set of prose that is being replaced, are kept on their own boards and counted separately. 489 further cells were measured on a build this site no longer offers, so they are held rather than counted; the board lists each one. The full board names the method used on each side of every row, and lists the results that were measured but are not comparable enough to publish.

1.55×

Prefill vs MLX — ahead on all 24 cells tested, across 4 models

Apple M2 Pro · best: Qwen2.5-0.5B · 5,059.7 vs 3,274.5 tok/s

All 139 cells, engine by engine →

A dark corridor of matte black panels with one narrow light seam running into the distance.

Measured on the released build. Reproduce every published result.

Generation happens on your hardware and the weights stay in one place. You sign in once and activate the device; after that the licence is leased briefly into memory, so the first run of a session waits for it and the runs after it do not. The model itself never leaves your machine.

Image rendered by Veizik 1.0.0 (sealed bundle, linux x86_64) · SD3.5-medium · NVIDIA RTX 3090 · Powered by Stability AI

How it works

From nothing to a first run.

The whole path, including the parts that are easy to leave out of a quickstart and then discover at the end.

  1. Install

    on your machine

    One signed package. One command in a shell, and it is ready.

    ./install.sh

  2. Create an account

    needs the network

    An email link, Google or GitHub. Self-service, nobody in the loop.

    veizik.com/signin

  3. Activate this device

    needs the network

    The command opens your browser — and prints the link too, in case it cannot. You approve there, and the machine receives a licence bound to it.

    veizik activate

  4. Pull a checkpoint

    needs the network

    One command downloads a supported safetensors checkpoint into the local cache and makes it ready to run. A gated repository needs a token stored with veizik hf login; the licence is accepted at the publisher.

    veizik pull

  5. Run

    on your machine

    Generation happens on the device. Your prompt, the weights and the output stay on the machine; the first run of a session leases the licence key into memory, and for the window the server sets — 48 hours as of 2026-10-02 — the runs that reuse it need no connection. Restarting the computer ends the window.

    veizik run

3 of the 5 steps reach the network. Generation is not one of them.

zsh
$ ./install.sh
$ veizik activate
$ veizik pull qwen2.5-0.5b
$ veizik run qwen2.5-0.5b "The capital of France is"
  Paris
Models

Runs today.

Every supported model has a memory footprint measured on this build, the context length the build declares, and a tested command path.

ModelTargetDevice memoryContext
Qwen2.5-0.5BApple Silicon328 MiB32768
Qwen2.5-1.5BApple Silicon884 MiB32768
Qwen3-1.7BApple Silicon923 MiB32768
Qwen2.5-7BApple Silicon3,923 MiB32768

Measured on the released build, not estimated from parameter count. The bar is against the largest figure the catalogue holds, so it means the same here as it does on every model page.

Coming next

FLUX.1 on NVIDIA

Still-image generation from text, run directly from the published checkpoints.

Not published yet. The engine runs it today; no download is open. Stable Diffusion 3.5 follows.

The full catalogue, with filters →

Get started

Put your hardware to work.

Install Veizik, point it at a supported checkpoint, and run locally.

Basic is over. Veizik begins.