NEAR AI Launched as Inference Provider on OpenRouter
The platform now hosts the multimodal GLM 5.3 Flash model, which features a 1 million token context window.
Updated on Sept. 28, 2026 in Artificial Intelligence

Live Poll
Do you trust private AI providers to keep your personal data and prompts completely secure?
NEAR AI Cloud has launched as a new inference provider on OpenRouter, giving users access to the Z.ai-developed GLM 5.3 Flash model. The deployment leverages hardware-based confidential computing and a strict zero data retention policy.
Why it matters
The integration expands access to high-capacity multimodal models while emphasizing privacy through cryptographically attested execution. This move brings secure, scalable inference options to the growing ecosystem of models hosted on OpenRouter.
The GLM 5.3 Flash model utilizes a multimodal mixture-of-experts architecture, which routes tasks to specific sub-networks to optimize efficiency. Inference is priced at $0.15 per million input tokens and $0.50 per million output tokens.
The players
NEAR AI
An AI inference provider focusing on confidential computing and zero-retention data policies.
OpenRouter
A centralized hub that routes AI model requests across various providers and hosted architectures.
Z.ai
The research organization responsible for developing the GLM series of multimodal AI models.
The details
NEAR AI Cloud secures inference by running processes inside Trusted Execution Environments (TEEs) — hardware-isolated areas of a processor that prevent unauthorized access to data in use. It utilizes Intel TDX and NVIDIA confidential computing solutions to maintain this security boundary. For every request, the platform generates a cryptographic attestation report, a digital signature confirming that the code was executed within a verified secure environment.
Timeline
August 26, 2026: The GLM 5.3 Flash model launched publicly.
September 28, 2026: NEAR AI Cloud went live on the OpenRouter platform.
The Tech Race
This development follows the established pattern of modular model routing where inference providers compete on security, cost, and architecture. It marks a push to bring hardware-attested, confidential computing to the broader model-as-a-service market.
Developers and users can now access the GLM 5.3 Flash model directly through OpenRouter using the provided token pricing. The integration is immediately available for workflows requiring confidential computing and auditability of model execution.
The takeaway
The move highlights a growing industry focus on verifiable privacy in model inference, moving beyond simple output quality. Watch for future performance benchmarks comparing the active 18 billion parameters of this mixture-of-experts model against denser, non-moe alternatives.
Further reading
For broader trends in model accessibility and hosting, see the latest updates in Artificial Intelligence.
Source note: This article includes information reported by Crypto Briefing.
Live Poll
Do you trust private AI providers to keep your personal data and prompts completely secure?







