Apple Silicon + NVIDIA CUDA

Run LLMs,
Fast.

Parallel decoding for Apple Silicon and NVIDIA GPUs. OpenAI-compatible API.

CURLmacOS · Linux
curl -fsSL https://tensorfold.dev/install.sh | sh
Start a model
. "$HOME/.tensorfold/env"
tensorfold serve \
  Vontra/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit \
  --name tensorfold

Downloads the model and MTP head on first serve.

Client base URL http://127.0.0.1:8080/v1

Send a prompt
curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"tensorfold","messages":[{"role":"user","content":"What is lane batching?"}],"max_tokens":256,"chat_template_kwargs":{"enable_thinking":false}}'
Setup guide + other models

Earlier MLX results

Recorded measurements

Nemotron 3.5 Lightning

30B-A3B · Before MTP drafts

M5 Max 128 GB

188-206 tokens per second

Qwen3.8 27B

With DFlash2

M5 Max 128 GB

120-124 tokens per second

Qwen3.8 Flash Next

M3 Ultra 256 GB

105-107 tokens per second

4-bit models, short thinking replies, earlier builds. Current-release benchmarks are pending.

TensorFold MiaAI LAB

Built with Mia's AI Lab.

We partnered with Mia and her team to optimise these recipes for TensorFold. They've done incredible work on the patches, setup and testing, and we're grateful. The numbers speak for themselves.

Visit Mia's AI Lab

Prose decode, one request per setup. Hardware and patches differ; each recipe links to the full method and results.

How it works

A drafter proposes several tokens. The target model checks them together and commits the matching prefix. Metal kernels keep verification rows consistent with serial decoding.

Kernel notes + implementation