Research · Inference

Inference that reaches more people

We build the engines that let models run on the machines people already have. That is how local, accessible intelligence becomes real, not only a research result.

How we approach inference

Our goals emphasise computationally efficient intelligence and access everywhere. Inference research is where those goals meet deployment: serve advanced models without assuming heavy GPU dependency.

Accessible everywhere

Intelligence should run where people are, including ordinary CPUs and thin networks, not only in GPU-rich data centres.

Computationally efficient

Aligned with our research goal of systems that work with minimal resources while remaining useful and reliable.

Local by design

Keep models close to users and data so control, privacy, and offline use stay practical.

Any model that matters

Inference for language, vision, and other ML workloads: the stack that turns research models into deployable systems.

Current engines

Open an engine for demos, details, and links.