Why Ariadne?
If you're joining a DataFrame against a large collection of files, Spark normally has to read all of them. Ariadne builds lightweight indexes over those files so that only the relevant ones are loaded — skipping files that can't contain matching rows.
Like Ariadne from Greek mythology, this library helps you navigate your data labyrinth. And just like Ariadne, I hope you'll one day betray it — like Theseus — and move on to something better, such as Apache Iceberg or Delta Lake proper. But in the meantime, if your data lake is more of a data swamp, Ariadne can help.
A quick look
import dev.cjfravel.ariadne.Index
import dev.cjfravel.ariadne.Index._ // brings in the df.join(index, ...) implicit
// Point Ariadne at any Hadoop-accessible storage location
spark.conf.set("spark.ariadne.storagePath", "abfss://ariadne@mystorage.dfs.core.windows.net/index")
// Build an index over your data files
val orders = Index("orders", orderSchema, "parquet")
orders.addIndex("customer_id")
orders.addFile(orderFiles: _*)
orders.update
// Use it in a normal Spark join — only matching files are loaded
val result = customerDf.join(orders, Seq("customer_id"), "inner")
Where to next?
For users
You want to add Ariadne to your project and start indexing files.
- Getting Started — install & basic usage
- Index Types — pick the right one for your data
- Usage Guide — joins, catalog, SQL, inspection
- Configuration — every Spark conf knob
- Maintenance — compact, vacuum, delete files
- Troubleshooting — errors & fixes
For contributors
You want to hack on Ariadne itself, or understand how it works inside.
- Contributing — dev setup, code style, PR process
- Architecture — traits, data flows, internals
- Security Policy — reporting vulnerabilities
- CLA — Contributor License Agreement
📚 API Reference — full Scaladoc for every public class, trait, and method. Generated from the source.