Nemotron 3.5 Lightning
30B-A3B · Before MTP drafts
M5 Max 128 GB
188-206 tokens per second
Apple Silicon + NVIDIA CUDA
Parallel decoding for Apple Silicon and NVIDIA GPUs. OpenAI-compatible API.
curl -fsSL https://tensorfold.dev/install.sh | sh. "$HOME/.tensorfold/env"
tensorfold serve \
Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit \
--name tensorfold
Downloads the model and MTP head on first serve.
Client base URL http://127.0.0.1:8080/v1
curl http://127.0.0.1:8080/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"tensorfold","messages":[{"role":"user","content":"What is lane batching?"}],"max_tokens":256,"chat_template_kwargs":{"enable_thinking":false}}'
30B-A3B · Before MTP drafts
M5 Max 128 GB
188-206 tokens per second
With DFlash2
M5 Max 128 GB
120-124 tokens per second
M3 Ultra 256 GB
105-107 tokens per second
4-bit models, short thinking replies, earlier builds. Current-release benchmarks are pending.
We partnered with Mia and her team to optimise these recipes for TensorFold. They've done incredible work on the patches, setup and testing, and we're grateful. The numbers speak for themselves.
Visit Mia's AI LabProse decode, one request per setup. Hardware and patches differ; each recipe links to the full method and results.
A drafter proposes several tokens. The target model checks them together and commits the matching prefix. Metal kernels keep verification rows consistent with serial decoding.
Kernel notes + implementation