← Explore
TOPIC

#big-data

Open source repositories tagged with #big-data, ranked by health score.

apache
apache/hive
Java
97
health

Apache Hive

★ 6.0k
apache
apache/flink
Java
94
health

Apache Flink

★ 26.4k
apache
apache/ozone
Java
92
health

Scalable, reliable, distributed storage system optimized for data analytics and object store workloads.

★ 1.3k
apache
apache/calcite
Java
92
health

Apache Calcite

★ 5.2k
Eventual-Inc
Eventual-Inc/Daft
Rust
87
health

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

★ 5.8k
StarRocks
StarRocks/starrocks
Java
87
health

The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.

★ 12.2k
apache
apache/paimon
Java
87
health

Apache Paimon is a lake format that enables building a Realtime Lakehouse Architecture with Flink and Spark for both streaming and batch operations.

★ 3.4k
prestodb
prestodb/presto
Java
87
health

The official home of the Presto distributed SQL query engine for big data

★ 16.8k
apache
apache/hugegraph
Java
87
health

A graph database that supports more than 100+ billion data, high performance and scalability (Include OLTP Engine & REST-API & Backends)

★ 3.2k
trinodb
trinodb/trino
Java
86
health

Official repository of Trino, the distributed SQL query engine for big data, formerly known as PrestoSQL (https://trino.io)

★ 13.3k
ClickHouse
ClickHouse/ClickHouse
C++
86
health

ClickHouse® is a real-time analytics database management system

★ 50.3k
apache
apache/datafusion
Rust
84
health

Apache DataFusion SQL Query Engine

★ 9.4k
linkedin
linkedin/openhouse
Java
84
health

Open Control Plane for Tables in Data Lakehouse

★ 399
apache
apache/beam
Java
83
health

Apache Beam is a unified programming model for Batch and Streaming data processing.

★ 8.7k
ytsaurus
ytsaurus/ytsaurus
C++
83
health

YTsaurus is a scalable and fault-tolerant open-source big data platform.

★ 2.2k
apache
apache/fluss
Java
79
health

Apache Fluss is a streaming storage built for real-time analytics.

★ 2.2k