Enterprise DNA
News AI News

Magnitude launches self-tuning agent engine

Inference engine startup shipped a tool that compiles and tunes its own kernels on your device so open models run up to 2x faster than llama.cpp..

Enterprise DNA |
Magnitude launches self-tuning agent engine

AI Pulse · Under the Radar

The play

Test self-tuning inference on your highest-volume open-model workflow, measuring real latency and device costs before changing production.

Magnitude has released an open-source inference engine designed to make locally run AI models fit the machine they are running on. Instead of relying on precompiled kernels built for broad hardware categories, it compiles and tunes those kernels on the user’s own device before the model runs. The project says this can make open models run up to twice as fast as llama.cpp, with support for Apple Silicon, NVIDIA and AMD hardware, or CPU-only machines. It runs on macOS, Windows and Linux, and connects with agents including Codex, Claude Code, Cline and others. The project’s GitHub page also says prompts, files and models stay on the machine once a model has been downloaded.

For a business owner, the practical point is not the benchmark alone. Faster local models can make private AI tools more usable for work that involves internal documents, customer information or repeatable staff tasks, without sending that material to an external model provider. The speed claim will depend on the specific hardware and model, so treat it as a reason to test rather than a promised outcome.

This is the kind of thing we build into an AI command centre, where the question is which AI work should stay inside your business, which tools staff can use, and how you measure whether the setup is actually saving time.

More like this?

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

Add your name (optional)

No spam. Unsubscribe any time.