Magnitude (YC S25) – Self-optimizing inference engine for agents

Hey HN, Anders and Tom here. We're building Magnitude, an inference engine for agents that optimizes itself to run as fast as possible on your hardware. It works on Mac, Linux, and Windows on any hardware and is up to 2x faster than llama.cpp.We're both software engineers and previously built an open source browser agent to 4k+ GH stars and 100k+ downloads. We increasingly wanted to run it on local models, but found that no inference engine worked for our use case.Inference engines today all make a performance tradeoff. They are either:- Built for batched inference on datacenter hardware at the cost of single-session performance (vLLM, SGLang) - Designed for broad compatibility instead of optimizing for specific hardware (llama.cpp, Ollama) - Specialized for specific hardware or models but lacking engine completeness (oMLX, ds4)Plus none of them are designed for running agents locally. Sessions are long, several often run at once, and you still want to use your computer for other things.Magnitude is built for maximum performance on your hardware and running local agents:- On-device compilation and tuning: Kernels are written with flexible parameters that are tuned on your actual device before the model runs. This gives you broad hardware compatibility with the same performance ceiling as hardware-specific kernels.- Focus on best architectures: We write our tunable, highly efficient kernels for the most popular open-weights families. This allows us to achieve and surpass the performance of hardware or model specialized engines, without forcing ourselves to over-generalize at the cost of performance.- Dynamic memory allocation: Magnitude reserves only enough memory up front to hold model weights. As your agent sessions grow, the memory heap dynamically increases, and frees itself when agents stop. Your hardware can still be used for other stuff while agents run.- Hybrid paged attention: We borrow the best ideas from engines like SGLang to allow concurrent se

Hacker NewsToolAISource
0Sign in to voteCopy link

FL score

78

out of 100

Verdict

BUILD

high confidence

Competition

No competitor data yet

Trend

No signal yet

A self-optimizing inference engine that runs local AI agents faster and more efficiently across diverse hardware.

The pain

Current inference engines either sacrifice single-session speed, hardware optimization, or completeness, making local agent use slow or resource-heavy.

The gap

No existing engine combines broad hardware compatibility, dynamic resource management, and high single-session performance for local agents.

Build angle

Leverage device-specific kernel tuning and dynamic memory management to deliver superior local inference speed without hardware lock-in.

Strengths

  • Founders have relevant open source and engineering experience
  • Clear technical differentiation with on-device tuning
  • Addresses a growing need for local AI agent performance
  • Cross-platform support broadens potential user base

Risks

  • Market may be niche and slow to adopt new inference engines
  • Competition from established engines and hardware vendors
  • Complexity of supporting wide hardware and model families
  • Monetization strategy and willingness to pay unclear

Questions about this idea?

FlyBot reads the scoring and gives you a second opinion on “Magnitude (YC S25) – Self-optimizing inference engine for agents”.

Open FlyBot