B

BaseRT by Base Compute

The fastest LLM inference runtime for Apple Silicon

Up to 6.4x faster prefill than llama.cpp33% faster decode on Qwen3Optimised for Apple M-series chipsOpen-source runtimeIntegrated with local coding agents
Category
AI Infrastructure / ML Runtime
Business model
Open-source core runtime, probable future enterprise support, hosting, or licensing
Target audience
AI/ML engineers building on-device inference on Apple hardware, developers of local coding agents, privacy-conscious users who want to run LLMs without cloud APIs
Opportunity score
55/100

Moderate opportunity: clear pain point but monetisation is unproven, competition is fierce, and the technical moat is replicable with enough investment.

Founder verdict
MAYBE

BaseRT proves the hunger for local AI speed, but the path to a defensible business is unclear and heavily dependent on temporary hardware advantages.

What is BaseRT by Base Compute?

BaseRT is an open-source inference runtime purpose-built for Apple Silicon (M1โ€“M5) that delivers substantially faster LLM inference than popular alternatives like llama.cpp and Apple's own MLX, especially on prefill-bound workloads. The team from Base Compute (Melbourne & Berlin) positions it as the default backend for on-device AI and local coding agents, emphasising privacy, speed, and a one-command install. While benchmarks show impressive gains on certain models, monetisation details remain undisclosed, and the project appears to be in an early community-building phase.

Get the full BaseRT by Base Compute playbook

The complete reverse-engineering โ€” what to copy, what to avoid, and exactly what to build instead.

Unlock full startup intelligence โ€” $39.99