KAYTUS upgrades MotusAI to run on-premises agentic AI at scale
KAYTUS has released a significant upgrade to MotusAI, its enterprise AI platform, positioning the product as a full-stack solution for running agentic AI workloads entirely within a customer's own data centre. The Singapore-headquartered infrastructure vendor says the upgrade addresses three pressure points that organisations face as they move beyond simple LLM queries into multi-agent systems: data-sovereignty risk, latency instability under bursty agent traffic, and uncontrolled API spending.
The company frames the upgrade around what it calls "Token Factories," a term it uses to describe on-premises GPU clusters configured to produce, distribute, and govern model inference at scale. Darren Cox, GM of KAYTUS Europe, said enterprises "don't just need more GPUs, they need the ability to govern, settle, and scale token services reliably." The company says MotusAI covers the lifecycle from raw GPU capacity through to per-department chargeback and quota enforcement.
What the upgrade delivers
The headline technical additions centre on inference efficiency and resilience. KAYTUS says MotusAI now uses Prefill-Decode disaggregation, dynamic KV caching, and dynamic batching to reduce time-to-first-token and end-to-end latency for multi-turn agent interactions. Integration with popular open-source inference runtimes, including vLLM and SGLang, allows real-time autoscaling triggered by live telemetry of throughput, latency and GPU utilisation.
Benchmark figures released by KAYTUS show autoscaling activating within 21 seconds of a peak-traffic event and expanding capacity to 16 instances within two minutes. Separately, the company reports that fine-grained GPU partitioning raised hardware utilisation from 68.9% to 95.7% in internal testing. The cost-saving claim of 30–50% annual reduction in token operating expenses is attributed primarily to filtering redundant requests and enforcing departmental quotas, though KAYTUS has not published methodology or third-party validation for either set of figures.
For developer teams, MotusAI now exposes an OpenAI-compatible API gateway that allows switching between open-source and commercial models without rewriting application code, alongside native support for million-token context windows for document-analysis and code-generation workloads.
Market context and competitive positioning
The on-premises AI inference market has grown sharply as enterprise buyers in regulated sectors weigh the data-residency implications of routing sensitive workloads through public cloud LLM APIs. Financial services, healthcare and government organisations in particular are under pressure from regulators in multiple jurisdictions to demonstrate that personal and commercially sensitive data does not leave controlled environments. KAYTUS cites deployments with an unnamed overseas fintech running eight GPU servers, a Japanese GPU cloud operator serving more than 30 enterprise clients, and a Southeast Asian NeoCloud operator that chose the integrated KAYTUS hardware-software stack over an unnamed international competitor.
KAYTUS competes in a crowded field that includes purpose-built inference orchestration from vendors such as NVIDIA NIM, and open-source stacks assembled around vLLM or TGI, as well as broader enterprise AI platform plays from larger infrastructure names. The company's differentiator, as articulated in this release, is the combination of inference orchestration with chargeback, multi-tenancy, and liquid-cooling hardware from a single vendor, which may appeal to operators building shared GPU-cloud services rather than single-tenant enterprise deployments.
Regulatory read-across
Digital sovereignty frameworks are hardening across KAYTUS's key markets. The EU AI Act's obligations for high-risk AI systems, DORA requirements for financial-sector resilience in Europe, and equivalent mandates in Asia-Pacific are making on-premises or private-cloud deployment a compliance default rather than an option for many regulated buyers. NIS2 obligations around incident reporting and supply-chain security add further weight to the argument for keeping inference processing inside the enterprise perimeter. How MotusAI's self-healing failover and audit logging map to these frameworks will likely be an early question from enterprise procurement teams. KAYTUS has not yet published specific compliance certifications such as SOC 2 or ISO 27001 against the platform.