Skip to content
Hyperfluid 2.0 is live console.hyperfluid.cloud/signup
Federated lakehouse

All your sources, one SQL query.

Break down organizational silos

Transform your scattered data sources into unified analytics power. Built on Trino and Apache Iceberg, Data Hub federates your managed PostgreSQL and your Iceberg tables on S3 object storage into a single queryable base, and is fed through the Open Data, Google Drive and Email connectors.

Hyperfluid Data Hub interface: SQL engine and connected source catalogs

Specifications

Federation
Managed PostgreSQL, Iceberg on S3
SQL Engine
Distributed Trino
Ingestion
Open Data, Google Drive, Email
Cross-source
Multi-source joins
Storage
S3 buckets created from the console
Medallion
Bronze, silver, gold Iceberg schemas
Dataset Branching
Roadmap

Use cases

Cross-source analytics

Join your managed PostgreSQL with your Iceberg tables in a single query

Multiple sources, one SQL query

Complex KPI calculation

Advanced analytics on distributed datasets

No ETL pipeline to maintain

Unified lakehouse

Open data into Iceberg, Google Drive files and email attachments into S3: connectors feed your queryable lakehouse

No ingestion code to write

In action

The impossible join

An analyst needs to cross customer data stored in PostgreSQL with product events stored in Iceberg tables

  1. Declare the managed PostgreSQL catalog and the Iceberg catalog in the SQL engine
  2. Write a single SQL query joining both sources
  3. SELECT * FROM postgres.customers JOIN iceberg.events...
  4. Read the results without copying or moving the data

A multi-source join with no prior extraction and no intermediate copy

Consolidated reporting

A finance team consolidates its monthly reporting from its PostgreSQL database and its Iceberg tables

  1. Federate the managed PostgreSQL database and the Iceberg catalog in the Trino SQL engine
  2. Write a query covering both sources in the SQL editor
  3. The Trino engine distributes the computation across the cluster
  4. Save the query and replay it at every close

Consolidated reporting, queryable in SQL, with no pipeline to maintain

A lakehouse fed from the console

A data team feeds its lakehouse with open data datasets and collects its business files from Google Drive and an email inbox

  1. Create the lakehouse S3 buckets and an Iceberg catalog from the console
  2. Import an open data source into an Iceberg table of that catalog, one off or daily
  3. Connect the Google Drive and IMAP email connectors to drop files and attachments into the buckets
  4. Query the Iceberg tables in standard SQL from the console editor

A lakehouse fed from the console, with no ingestion code to write

Powered by open standards

SQL engine

Trino

Distributed and stateless SQL engine. Query your connected sources with standard SQL, without moving your data.

Table format

Apache Iceberg

Open source table format for data lakes: ACID transactions, schema evolution, compaction and snapshot expiry driven from the console.

Connectors

Multi-source federation

PostgreSQL
Managed database, read only
Apache Iceberg
Lakehouse tables
S3
Object storage, exposed through Iceberg
MySQL
Relational databases (coming soon)
MongoDB
NoSQL (coming soon)
Elasticsearch
Search (coming soon)

Ingestion connectors

Open Data
23 public sources, one off or daily ingestion into your Iceberg tables
Google Drive
File synchronization into your S3 buckets, with optional text extraction
Email
Attachments from an IMAP mailbox dropped into your S3 buckets

Open data, files and emails feed your Iceberg tables and S3 buckets with no pipeline to write. More federation sources are coming soon.

Key benefits

  • Query your connected sources with standard SQL
  • No ETL required: analyze in place
  • Distributed computation for massive scale
  • Apache Iceberg for ACID transactions

Ready to unify your data sources?

Discover how Data Hub can eliminate your data silos.