Skip to content
Hyperfluid 2.0 is live console.hyperfluid.cloud/signup
Private AI platform

Your models run on your own GPUs.

An OpenAI-compatible endpoint, inside your perimeter

Serve AI models on your own cluster's GPUs: generation, embeddings and document OCR, ready to use with the vLLM and HuggingFace runtimes, or through optional external models. A single OpenAI-compatible endpoint, per-team API keys and consumption tracking, all the way to air-gapped networks.

Hyperfluid console: AI model serving, API keys and consumption tracking

Specifications

Model serving
On your cluster's GPUs
Shared models
Generation, embeddings, OCR
Runtimes
vLLM, HuggingFace, external models
Single endpoint
OpenAI-compatible
API keys
Per team
Consumption
Tracked in the console

Use cases

Serve your models on your GPUs

Deploy recent open generation models on your own cluster's GPUs and expose them to all your applications through a single OpenAI-compatible endpoint.

OpenAI-compatible

Embeddings and OCR ready to use

Shared models cover text generation, embeddings and document OCR, with nothing to provision. External models are available as an option.

Shared models

Your AI assistants on your data

Plug your AI assistants and agents into your tables through the Data APIs MCP endpoint, and ask questions in natural language with the Dashboards copilot.

Via MCP

In action

Plugging an internal application into AI

A product team wants to add generation without sending data outside its perimeter

  1. Pick a shared model or serve one on your cluster's GPUs
  2. Create an API key dedicated to the team in the console, with its scope and expiry
  3. Point the application at the single OpenAI-compatible endpoint
  4. Track the key's consumption in the console, by model and by key

The application uses AI, data stays inside your perimeter

Querying your data without writing SQL

A business analyst wants to explore their governed tables from the console

  1. Open Dashboards in the console
  2. Ask the copilot the question in natural language
  3. The copilot proposes a widget, with its query and a preview on your real data
  4. Pin the widget to the dashboard or ask for an adjustment

The answer is built on your governed tables, without exporting a single record

Key benefits

  • Your models run on your GPUs, your data stays inside your perimeter
  • A single OpenAI-compatible endpoint for all your applications
  • API keys per team, consumption tracked in the console
  • Self-hosted models up to air-gap, external models as an option

Ready to put AI to work on your data?

Discover how Hyperfluid sovereign AI leverages your information assets, without compromising them.