// HACKER NEWS — CYBERSECURITY
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Magnitude is an open source inference engine for agents that optimizes itself for your exact hardware. It compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. One click connects the agent you already use (Pi, OpenCode, Hermes, Codex, and more). Works on Apple Silicon, NVIDIA, AMD, or nothing but a CPU.
⭐ Help us reach more developers and grow the Magnitude community. Star this repo!
The desktop app includes the magnitude CLI. No separate installation is needed.
An open source inference engine that optimizes itself for your hardware. It ships as a desktop app that runs open models and connects them to the agent you already use.
They ship kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip. See the benchmarks against llama.cpp.
Any Apple Silicon, NVIDIA, or AMD GPU, or nothing but a CPU. There is no fixed minimum. Smaller machines run smaller models, and more memory lets you run larger ones.
See the full list at magnitude.dev/models. We write optimized kernels for the most popular open-weight families, which is how we beat generalist engines.
One click connects Pi, OpenCode, Hermes, OpenClaw, Codex, Claude Code, Oh My Pi, and Cline. Anything else works through the OpenAI-compatible API.
Yes. Prompts, files, and models stay on your machine. No internet needed once a model is downloaded.