dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
arrow-select
Counted from the Cargo.toml manifests of the 32 indexed repositories that declare arrow-select as a dependency — not download counts. Dependency data last verified 2026-09-09.
Crates that show up unusually often in arrow-select projects. The most distinctive pairings rank first — crates these projects use far more than the average indexed Rust project does, not just crates that are popular everywhere. Each percentage is the share of arrow-select projects that also use it.
Developer-friendly OSS embedded retrieval library for multimodal AI. Search More; Manage Less.
A data visualization and analytics component, especially well-suited for large and/or streaming datasets.
Data Agent Ready Warehouse : One for Analytics, Search, AI, Python Sandbox. — rebuilt from scratch. Unified architecture on your S3.
Simple, Elastic-quality search for Postgres
Open Lakehouse Format for Multimodal AI. Convert from Parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..
High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale
Official Rust implementation of Apache Arrow
A native Rust library for Delta Lake, with bindings into Python
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
Parseable is an observability datalake built from first principles.
Tonbo is an embedded database for serverless and edge runtimes.
Apache Iceberg
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Local-first ETL/ELT studio: a drag-and-drop visual pipeline designer that compiles to SQL and runs on DuckDB. Tiny desktop app, no servers, git-friendly workspaces.
Lakehouse native graph engine with git-style workflows
Intuitive Data Workflows
Scalable graph analytics database powered by a multithreaded, vectorized temporal engine, written in Rust
Protocol and libraries for sending and receiving OpenTelemetry data using Apache Arrow
SciRS2 - Scientific Computing and AI in Rust., providing SciPy-compatible APIs while leveraging Rust's performance, safety, and concurrency features. Unlike traditional scientific libraries
The native Rust implementation for Apache Hudi, with C++ & Python API bindings.
A typed, polyglot, functional language
Apache Paimon Rust The rust implementation of Apache Paimon.
Embeddable spreadsheet engine — parse, evaluate & mutate Excel workbooks from Rust, Python, or the browser. Arrow-powered, 320+ functions.
We're back! Now firing notebooks out of a t-shirt gun.
DuckLake took Flight. Welcome to SwanLake.
On-device property graph database. Schema-as-code. One CLI → One Folder. No Server. Think: DuckDB for graphs.
Library for bringing distributed capabilities to Apache DataFusion
Graph database native to the cloud. Embedded, multi-tenant, built on object storage.
Lossless storage and search for AI agent sessions, across every agentic client.
Open-source streaming SQL engine written in Rust using Apache Arrow and DataFusion. Supports continuous queries, temporal stream joins, tumbling/session windows, and CDC/Kafka connectors. Lightweight, embeddable, and sub-microsecond latency
A framework to manage data, continuously