Accessible everywhere
Intelligence should run where people are, including ordinary CPUs and thin networks, not only in GPU-rich data centres.
Research · Inference
We build the engines that let models run on the machines people already have. That is how local, accessible intelligence becomes real, not only a research result.
Our goals emphasise computationally efficient intelligence and access everywhere. Inference research is where those goals meet deployment: serve advanced models without assuming heavy GPU dependency.
Intelligence should run where people are, including ordinary CPUs and thin networks, not only in GPU-rich data centres.
Aligned with our research goal of systems that work with minimal resources while remaining useful and reliable.
Keep models close to users and data so control, privacy, and offline use stay practical.
Inference for language, vision, and other ML workloads: the stack that turns research models into deployable systems.