Jul 30, 2026
Enterprise

Protopia and Rafay GPU partnership targets multi-tenant AI factories

Protopia AI and Rafay Systems are pairing token-metered model access with prompt privacy controls for shared enterprise GPU infrastructure.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 3 min read

Protopia and Rafay GPU partnership targets multi-tenant AI factories
Photo: SiliconANGLE

Protopia AI and Rafay Systems are combining their products to make shared GPU infrastructure more usable for enterprises that want token-metered AI access without exposing private context. The Protopia Rafay GPU effort pairs Rafay’s Token Factory with Protopia’s Stained Glass technology, a privacy layer meant to let operators serve multiple customers from the same hardware without assigning dedicated GPUs to each tenant.

The companies did not disclose commercial terms, customer names, revenue impact, deployment counts or performance overhead. The pitch, described by Protopia Chief Executive Eiman Ebrahimi and Rafay Systems co-founder and Chief Executive Haseeb Budhani in a theCUBE interview at the NYSE Wired: Robotics & AI Infra Leaders event, is aimed at a common problem for AI infrastructure providers: expensive GPU capacity can be reserved on paper while still sitting underused in practice.

Ebrahimi said operators often focus on adding more supply, while existing capacity may be over-allocated and not consistently busy. In that case, he said, providers are missing revenue from infrastructure they have already bought or leased.

What are Protopia and Rafay offering for shared GPUs?

Rafay’s Token Factory, which the company made generally available in April, gives infrastructure operators a way to provide metered, serverless access to open-source AI models and related services. Instead of carving out GPUs for each enterprise customer, operators can sell access by tokens through a more abstracted service model.

Protopia’s contribution is data isolation for those shared environments. Its Stained Glass technology changes prompts and contextual data before they reach the target model, so the model can use a transformed representation without receiving the original plain text, according to Ebrahimi.

That design is meant to sit upstream of inference and work alongside zero-data-retention policies. The claim is narrower than broad AI security marketing: Protopia says it reduces the exposure of readable enterprise data if information is mishandled or leaked during model use. The companies did not provide independent benchmark results or third-party validation in the discussion.

Why does multi-tenancy matter for enterprise AI infrastructure?

Multi-tenancy is central to the economics of AI infrastructure because GPU utilization determines whether an operator can turn high fixed costs into profitable usage. SiliconANGLE reported that the GPU-as-a-service market is projected to reach $7.36 billion in 2026, a sign that the category is moving beyond one-off hardware rental toward consumption-based services.

Budhani described the combined system as giving enterprises a private slice of shared GPU capacity. He said providers can make more complete use of their infrastructure, while customers may get lower pricing because the hardware is shared across tenants instead of reserved in isolated blocks.

The hard questions are still operational. The companies did not say how much utilization improves, how pricing changes versus dedicated GPU access, which models are supported in production, or how customers should measure any tradeoff between privacy transformation and model quality.

Budhani also argued that security features can increase provider margins because enterprise buyers may pay more for protected access. That may be true for regulated customers, but adoption will depend on evidence that shared inference can meet enterprise privacy, latency and cost requirements at the same time.

This story draws on original reporting from SiliconANGLE.

More from Enterprise

All Enterprise →