AI Pulse · Under the Radar
The play
Test self-tuning inference on your highest-volume open-model workflow, measuring real latency and device costs before changing production.
Magnitude has released an open-source inference engine designed to make locally run AI models fit the machine they are running on. Instead of relying on precompiled kernels built for broad hardware categories, it compiles and tunes those kernels on the user’s own device before the model runs. The project says this can make open models run up to twice as fast as llama.cpp, with support for Apple Silicon, NVIDIA and AMD hardware, or CPU-only machines. It runs on macOS, Windows and Linux, and connects with agents including Codex, Claude Code, Cline and others. The project’s GitHub page also says prompts, files and models stay on the machine once a model has been downloaded.
For a business owner, the practical point is not the benchmark alone. Faster local models can make private AI tools more usable for work that involves internal documents, customer information or repeatable staff tasks, without sending that material to an external model provider. The speed claim will depend on the specific hardware and model, so treat it as a reason to test rather than a promised outcome.
This is the kind of thing we build into an AI command centre, where the question is which AI work should stay inside your business, which tools staff can use, and how you measure whether the setup is actually saving time.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideTalk it through
Talk it through with Sam
30 minutes on what a Command Centre would look like for your business.
Book a callFree daily email
Get this every morning.
This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.
Free daily email
Subscribe to the daily AI Pulse
One short read every morning on what is actually happening in AI. Free.
You are in
Your first AI Pulse lands tomorrow morning. Keep an eye on your inbox.
Free Resource
Put what you just read to work
The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the Guide