Run open models on your machine.
rcli pull qwen3
rcli run qwen3Chat, vision, speech, and embeddings — all local. Nothing leaves the device.
brew install runanywhereai/rcli/rclior
curl -fsSL https://raw.githubusercontent.com/RunanywhereAI/RCLI/main/install.sh | shirm https://raw.githubusercontent.com/RunanywhereAI/RCLI/main/install.ps1 | iexcurl -fsSL https://raw.githubusercontent.com/RunanywhereAI/RCLI/main/install.sh | shrcli pull qwen3 # download
rcli run qwen3 # chat
rcli run qwen3 "Hello" # one-shot
rcli serve qwen3 # OpenAI-compatible API on :8080 (macOS/Linux)rcli models list --all is the full catalog. Short names work everywhere (qwen3, llama3.2, whisper-tiny, piper, …). Any Hugging Face GGUF works too:
rcli pull hf.co/Qwen/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q8_0.ggufOne rcli binary. Catalog models already name their engine (GGUF → llama.cpp, mlx-* → MLX, Core ML → NeuRT, QNN-context → QHexRT). You normally do not pick one.
Override only when you mean it:
rcli llm generate --engine mlx -m mlx-qwen3 "Hello"
rcli run --engine qhexrt /path/to/lfm2_5_230m_HNPU "Hello"
rcli image generate --engine neurt --prompt "a red cube" --out out.png--engine accepts mlx, llamacpp, sherpa, onnx, neurt / coreml / ane, and qhexrt / qnn / npu / hexagon. If you omit it, commons picks the highest-priority registered backend that implements that primitive:
| Priority | Engine | Who wins unpinned work |
|---|---|---|
| 150 | QHexRT | Every primitive it implements, and only on a Windows ARM64 overlay binary (often the only engine in that binary) |
| 110 | MLX | Apple GPU: LLM / VLM / TTS / STT / embeddings when an mlx-* model is not already pinned |
| 100 | llama.cpp | GGUF LLM / VLM / embed / rerank |
| 100 | NeuRT | Core ML only. Stays at 100 on purpose so it never steals GGUF/MLX traffic. A Core ML bundle reaches NeuRT by framework pin, not by winning priority |
| 90 | Sherpa-ONNX | STT / TTS / VAD |
| 50 | ONNX Runtime | embeddings / VAD / diarization / segmentation |
rcli backends is the source of truth for this binary. Public bottles never list neurt or qhexrt. Those engines are private overlays, never Homebrew / GitHub Release assets.
| Backend | macOS Apple Silicon | Windows x64 | Windows ARM64 | Linux x64 |
|---|---|---|---|---|
| llama.cpp | public bottle | public bottle | — | public bottle |
| MLX (Apple GPU) | public bottle (product rcli, not rcli-cxx) |
— | — | — |
| Sherpa-ONNX | public bottle | public bottle | — | public bottle |
| ONNX Runtime | public bottle | public bottle | — | public bottle |
| NeuRT (Apple Neural Engine; Core ML is the format) | overlay rebuild | — | — | — |
| QHexRT (Qualcomm Hexagon NPU) | — | — | overlay rebuild | — |
Public Windows ARM64 kits are commons-only (no llama.cpp / ONNX / Sherpa on MSVC ARM64). Snapdragon NPU is overlay-only. x64 Windows has no Hexagon path.
Yes = this engine implements the primitive. Try = a catalog id that rcli pull / a local path can run. Overlay engines still need the matching on-disk bundle (compiled .mlmodelc tree, or *_HNPU / v81/ QNN-context dir) — a Hugging Face repo page is HTML, not a model.
| Modality | Command | llama.cpp | MLX | Sherpa | ONNX | NeuRT | QHexRT |
|---|---|---|---|---|---|---|---|
| LLM | rcli run / llm generate |
yes · smollm2, qwen3 |
yes · mlx-qwen3 |
— | — | yes · lfm2-230m-ane local Core ML tree |
yes · lfm2-230m-npu local *_HNPU |
| VLM | rcli vlm generate --image |
yes · smolvlm2 |
yes · mlx-qwen2-vl |
— | — | — | yes · internvl-1b-npu local HNPU |
| TTS | rcli tts synthesize -o out.wav |
— | yes · mlx-soprano |
yes · piper |
— | — | yes · kitten-micro-npu local HNPU |
| STT | rcli stt transcribe audio.wav |
— | yes · mlx-qwen3-asr |
yes · whisper-tiny |
— | yes · parakeet-tdt-v2-ane local Core ML |
yes · whisper-base-npu local HNPU |
| VAD | rcli vad detect audio.wav |
— | — | yes | yes · silero |
— | — |
| Embeddings | rcli embed |
yes · nemotron-3-embed |
yes · mlx-qwen3-embed |
— | yes · minilm |
— | yes · embeddinggemma-npu local HNPU |
| Rerank | rcli rerank -d … |
yes · bge-reranker |
— | — | — | — | yes · nv-rerank-npu local HNPU |
| Segmentation | rcli segment image.ppm (binary P6 PPM) |
— | — | — | yes · segformer |
— | — |
| Diarization | rcli diarize audio.wav |
— | — | — | yes · sortformer |
— | — |
| Image gen | rcli image generate --prompt … --out … |
— | — | — | — | yes · sd15 (compiled Core ML zip, not the HF repo HTML) |
yes · cosmos3-diffusion-npu local HNPU |
MLX registers with a one-line -811 then Swift callbacks install it — that warning is expected. image generate is compiled only when NeuRT is linked; --prompt and --out are required (not a positional prompt). --steps 4 is enough for a smoke PNG.
QHexRT on device also needs QAIRT matching the Hexagon skel (QNN_SDK_ROOT + ADSP_LIBRARY_PATH=…\lib\hexagon-v81\unsigned on v81). Overlay 2.47 DLLs vs a 2.41/2.48 device skel will fail to instantiate graphs. Pass the *_HNPU directory, not a GGUF. GGUF files cannot run on the ARM64 overlay binary (no llama.cpp).
Catalog models are grouped by the org that trains them. GGUF rows run on llama.cpp (macOS, Windows x64, Linux). mlx-* rows run on Apple Silicon only.
| Org | Families | Try |
|---|---|---|
| Alibaba Qwen | Qwen3, Qwen3.6, Qwen3.8 | qwen3, mlx-qwen3 |
| Meta | Llama 3.2 | llama3.2, mlx-llama3.2 |
| Gemma 4 | gemma4-e2b, mlx-gemma4-e2b |
|
| Hugging Face | SmolLM2 | smollm2 |
| Liquid AI | LFM2 | lfm2 |
| IBM | Granite 4.1 | granite4.1-3b, mlx-granite4.1-3b |
| NVIDIA | Nemotron | mlx-nemotron-nano |
| PrismML | Bonsai, Ternary-Bonsai | bonsai-1.7b, mlx-bonsai-1.7b |
| DeepGrove | Maple Preview | maple-preview, mlx-maple-preview |
| Org | Families | Try |
|---|---|---|
| Hugging Face | SmolVLM2 | smolvlm2 |
| Alibaba Qwen | Qwen2-VL | qwen2-vl, mlx-qwen2-vl |
| Liquid AI | LFM2-VL, LFM2.5-VL | lfm2-vl, mlx-lfm2.5-vl |
| Apple | FastVLM | mlx-fastvlm |
| Microsoft | Fara 1.5 (computer use) | fara |
| Meta | Muse Glimmer | muse-glimmer |
| NVIDIA | Nemotron Omni | nemotron-omni |
rcli vlm generate --model smolvlm2 --image photo.png "What is in this picture?"| Org | Families | Role | Try |
|---|---|---|---|
| OpenAI | Whisper | STT | whisper-tiny |
| NVIDIA | Parakeet, Canary, Nemotron ASR | STT | parakeet-tdt-v2 |
| Alibaba Qwen | Qwen3-ASR / Qwen3-TTS | STT / TTS (MLX) | mlx-qwen3-asr |
| rhasspy | Piper | TTS | piper |
| Supertone | Supertonic | TTS | supertonic |
| Zhipu | GLM-ASR | STT (MLX) | mlx-glm-asr |
| Silero | Silero | VAD | silero |
rcli tts synthesize "Hello from the device." -o hello.wav
rcli stt transcribe hello.wav| Org | Families | Role | Try |
|---|---|---|---|
| NVIDIA | Nemotron Embed, Llama-Nemotron Embed | embeddings | nemotron-3-embed |
| Alibaba Qwen | Qwen3 Embedding | embeddings (MLX) | mlx-qwen3-embed |
| sentence-transformers | MiniLM | embeddings | minilm |
| BAAI | BGE Reranker | rerank | bge-reranker |
| NVIDIA | Sortformer | diarization | sortformer |
| NVIDIA / Hugging Face | SegFormer | segmentation | segformer |
| Stability AI / Apple | Stable Diffusion 1.5 | image gen (NeuRT) | sd15 |
macOS Apple Silicon (public bottle): llama.cpp + MLX + Sherpa + ONNX. Pull qwen3 (GGUF) or mlx-qwen3 (GPU). Image generation is NeuRT (sd15) and only works after the private overlay is linked into product rcli.
Windows x64 (public zip): GGUF / ONNX / Sherpa. No MLX, no NeuRT, no QHexRT.
Windows ARM64 (Snapdragon): public kit has no llama.cpp/ONNX/Sherpa. The QHexRT overlay runs Hexagon NPU models from a local *_HNPU tree. Do not expect mlx-*, GGUF, or sd15 on that binary.
rcli serve is macOS and Linux.
Device round-trips are by modality, not by engine. scripts/e2e.sh always
runs scripts/e2e-modalities.sh; public CI leaves the knobs unset and skips.
On a machine that already has models:
export RUNANYWHERE_HOME=/path/to/home # already-pulled OSS models
export RCLI_E2E_MODEL_ROOTS=/path/to/hnpu # *_HNPU / *_ANE / *.mlmodelc trees
bash scripts/e2e-modalities.sh /path/to/rcli # no --engine requiredRCLI_E2E_LLM, RCLI_E2E_STT, RCLI_E2E_IMAGE, … pin one primitive. Catalog
ids (mlx-qwen3, whisper-base-npu) pin the framework; a Hugging Face repo
page is HTML, not a bundle.
rcli run / rcli chat |
chat (REPL with no prompt) |
rcli pull / rcli models download |
download |
rcli list / rcli ls |
local models (--all = catalog) |
rcli show |
one model |
rcli rm |
delete |
rcli llm generate / stream |
completion |
rcli vlm generate --image |
vision |
rcli stt transcribe |
speech → text |
rcli tts synthesize |
text → WAV |
rcli vad detect |
voice activity |
rcli embed |
embeddings |
rcli rerank |
rerank documents |
rcli image generate |
text → image (NeuRT / Apple Silicon) |
rcli serve |
OpenAI-compatible HTTP (macOS/Linux) |
rcli backends |
registered engines |
rcli info |
versions and paths |
--engine |
force mlx / llamacpp / sherpa / onnx / neurt / qhexrt |
rcli --help and rcli <command> --help cover the rest.
Stage a C++ desktop kit from runanywhere-sdks. The pin is cmake/sdk-pin.cmake (RCLI_PINNED_SDK_VERSION).
C++-only (rcli-cxx on Apple; rcli elsewhere):
cmake -B build -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_PREFIX_PATH=/path/to/kit
cmake --build build
./build/rcli version # ./build/rcli-cxx on Apple
./build/rcli backendsApple Silicon product binary is the Swift MLX host (build/rcli). Independent clones need the SDK Swift tree (RCLI_SDK_SWIFT_PATH) and RCLI_APPLE_MLX_HOST=ON (the default):
export RCLI_SDK_SWIFT_PATH=/path/to/runanywhere-sdks
cmake -B build -DCMAKE_BUILD_TYPE=Release \
-DCMAKE_PREFIX_PATH=/path/to/kit
cmake --build build
# or: scripts/build-mlx.sh build
./build/rcli version
./build/rcli backendsSee CONTRIBUTING.md.
MIT. See LICENSE.