Ariadne logo

Find the thread through your data lake.

Ariadne builds lightweight file-level indexes over Parquet, CSV, and JSON files so Spark joins read only the files they actually need to.

CI Maven Central License Spark Scala

Why Ariadne?

If you're joining a DataFrame against a large collection of files, Spark normally has to read all of them. Ariadne builds lightweight indexes over those files so that only the relevant ones are loaded — skipping files that can't contain matching rows.

Like Ariadne from Greek mythology, this library helps you navigate your data labyrinth. And just like Ariadne, I hope you'll one day betray it — like Theseus — and move on to something better, such as Apache Iceberg or Delta Lake proper. But in the meantime, if your data lake is more of a data swamp, Ariadne can help.

A quick look

import dev.cjfravel.ariadne.Index
import dev.cjfravel.ariadne.Index._  // brings in the df.join(index, ...) implicit

// Point Ariadne at any Hadoop-accessible storage location
spark.conf.set("spark.ariadne.storagePath", "abfss://ariadne@mystorage.dfs.core.windows.net/index")

// Build an index over your data files
val orders = Index("orders", orderSchema, "parquet")
orders.addIndex("customer_id")
orders.addFile(orderFiles: _*)
orders.update

// Use it in a normal Spark join — only matching files are loaded
val result = customerDf.join(orders, Seq("customer_id"), "inner")

Where to next?

For users

You want to add Ariadne to your project and start indexing files.

For contributors

You want to hack on Ariadne itself, or understand how it works inside.

📚 API Reference — full Scaladoc for every public class, trait, and method. Generated from the source.