Data Hub
Break down organizational silos
Transform your scattered data sources into unified analytics power. Built on Trino and Apache Iceberg, Data Hub federates PostgreSQL, your Iceberg tables and S3 object storage into a single queryable base, and is fed through the Open Data, Google Drive and Email connectors.
Powered by open standards
Trino
Distributed and stateless SQL engine. Query your connected sources with standard SQL, without moving your data.
Apache Iceberg
Open source table format for data lakes. ACID transactions, time travel, schema evolution and optimal performance.
Connectors
Multi-source federation
PostgreSQL
Relational databases
Apache Iceberg
Lakehouse tables
S3
Object storage
MySQL
Relational databases (coming soon)
MongoDB
NoSQL (coming soon)
Elasticsearch
Search (coming soon)
Ingestion connectors
Open Data
23 public datasets, one-click import into Iceberg
Google Drive
File synchronization
IMAP attachments into your S3 storage
Open data, files and emails feed your Iceberg tables and S3 buckets with no pipeline to write. More federation sources are coming soon.
Specifications
Federation
PostgreSQL, S3, Iceberg
SQL Engine
Distributed Trino
Ingestion
Open Data, Google Drive, Email
Cross-source
Multi-source joins
S3 Buckets
Bronze, silver, gold in one click
Dataset Branching
Roadmap
Use cases
Cross-source analytics
Join PostgreSQL with your Iceberg tables in a single query
Multiple sources, one SQL queryComplex KPI calculation
Advanced analytics on distributed datasets
No ETL pipeline to maintainUnified lakehouse
Open data into Iceberg, Google Drive files and email attachments into S3: connectors feed your queryable lakehouse
Ingestion without pipelinesIn action
The impossible join
Sarah needs to join customer data (PostgreSQL) with product events (Iceberg tables)
- 1 Traditional approach: Extract → Transform → Load (weeks of work)
- 2 Hyperfluid approach: One SQL query across both sources
- 3 SELECT * FROM postgres.customers JOIN iceberg.events...
- 4 Results in seconds, without moving data
Complex analytics that took weeks, now in real-time
The KPI revolution
The finance team consolidates its monthly reporting from its PostgreSQL database, Iceberg tables and S3 exports.
- 1 Federate PostgreSQL, Iceberg and S3 in the Trino SQL engine
- 2 Write SQL query covering all sources
- 3 Trino engine distributes computation automatically
- 4 Complex revenue calculation with correct attribution
Consolidated reporting, queryable in SQL, with no pipeline to maintain.
Dataset branching
Branch your datasets to experiment safely (coming soon)
- 1 Create a dataset branch for Q3 analysis
- 2 Experiment with data transformations safely
- 3 Compare results with main branch
- 4 Merge successful changes or abandon experiments
Safe data experimentation without breaking production
Key benefits
Ready to unify your data sources?
Discover how Data Hub can eliminate your data silos.
Request a demo