Tencent Hy3 goes global with API pricing and multi-platform access

Tencent's 295-billion-parameter Hy3 model is now broadly available internationally via WorkBuddy, Miora and TokenHub, priced from $0.13 per million input tokens.

Brightly lit data center corridor with symmetrical rows of server racks featuring copper piping, numerous green and blue indicator lights, and square white ceiling lights.

Tencent has extended international access to Hy3, its latest large language model, one month after the model's formal launch on 6 July. The rollout makes the model available through three consumer and enterprise products, a direct API, and an expanding list of third-party developer platforms, with API pricing on OpenRouter starting at US$0.1288 per million input tokens and US$0.5336 per million output tokens.

Built on a hybrid Mixture-of-Experts architecture that blends fast and slow reasoning modes, Hy3 carries 295 billion total parameters but activates only 21 billion at inference time, a design choice that is intended to reduce compute cost per task while maintaining competitive output quality. The model supports a context window of up to 256,000 tokens and is available under an Apache 2.0 licence, permitting commercial use and downstream modification.

The product stack

Tencent is routing Hy3 through three distinct surfaces. WorkBuddy, described by the company as China's most widely used AI agent workspace, is offering free access until 31 August 2026 and the company says internal evaluations showed a task-success rate above 90 per cent alongside a 34 per cent reduction in average task-completion time compared with the previous Hy model. Tencent Design Miora, an AI-native creative studio targeting design and content teams, uses the model to generate production-ready graphics, video, 3D assets and user interfaces from natural-language briefs. Tencent Cloud TokenHub functions as a Model-as-a-Service aggregation layer, providing intelligent routing across Hy3 and third-party models through a single API endpoint with usage governance and flexible billing.

On third-party platforms, Hy3 is already integrated with Hermes, Kilo, Cline, OpenClaw, OpenCode and Cherry Studio, and is listed on Hugging Face and ModelScope. The company says the model recorded more than 68 times as many API calls as its predecessor model and reached the top of OpenRouter's global LLM usage leaderboard within a week of launch, though those figures are self-reported and were not independently verified at the time of publication.

Poshu Yeung, Senior Vice President of Tencent Cloud International, said the platform was designed to "give enterprises the control and reliability they need to move from pilots to scaled deployments."

Market context

The global enterprise LLM market is increasingly fragmented between US hyperscaler offerings, open-weight models from Meta and Mistral, and a growing tier of Chinese vendors, of which Tencent, Alibaba (Qwen) and Baidu (Ernie) are the most prominent internationally. Hy3 competes principally on cost efficiency and inference-time parameter activation ratios: at 21 billion active parameters it is positioned by Tencent as comparable in output quality to dense models two to five times larger, a claim that is consistent with the broader MoE design rationale but has not yet been validated by independent third-party benchmarking.

For enterprise buyers in regions such as the Gulf Cooperation Council, where a number of large-scale AI pilot programmes are moving towards production decisions in 2026, the availability of a commercially permissive, cost-competitive model with an enterprise token-management layer addresses a practical procurement gap. At the same time, data-residency and sovereignty requirements in markets such as the EU and parts of the Middle East mean that Tencent will need to demonstrate clear data-processing boundaries for regulated workloads, particularly as the EU AI Act's general-purpose AI model obligations continue to phase in.

Tencent Cloud's infrastructure spans 66 availability zones across 23 regions supported by more than 3,200 acceleration nodes, giving the company the physical footprint to support low-latency inference globally. The next milestones to watch are formal independent benchmark results, named enterprise customer wins outside China, and clarity on data-residency commitments for regulated industries.