Getting Started
Install Ariadne, point it at a storage location, build your first index, and use it in a Spark join.
Installation
Ariadne ships two build lines: Spark 3.5 / Delta 3.2 (Scala 2.12, Java 11) and Spark 4.1 / Delta 4.1 (Scala 2.13, Java 21). Add the dependency that matches the Spark runtime on your cluster.
Maven
<!-- Spark 3.5 / Delta 3.2 (Scala 2.12, Java 11) -->
<dependency>
<groupId>dev.cjfravel</groupId>
<artifactId>ariadne-spark35_2.12</artifactId>
<version>0.1.10-beta</version>
</dependency>
<!-- Spark 4.1 / Delta 4.1 (Scala 2.13, Java 21) -->
<dependency>
<groupId>dev.cjfravel</groupId>
<artifactId>ariadne-spark41_2.13</artifactId>
<version>0.1.10-beta</version>
</dependency>
SBT
// Spark 3.5 / Delta 3.2 (Scala 2.12)
libraryDependencies += "dev.cjfravel" %% "ariadne-spark35" % "0.1.10-beta"
// Spark 4.1 / Delta 4.1 (Scala 2.13)
libraryDependencies += "dev.cjfravel" %% "ariadne-spark41" % "0.1.10-beta"
Match your Spark line. The spark35 artifacts target Spark 3.5 / Delta 3.2 on Scala 2.12 and Java 11; the spark41 artifacts target Spark 4.1 / Delta 4.1 on Scala 2.13 and Java 21. Pick the artifact whose Scala version matches your cluster's Spark.
Configure storage
Before creating any indexes, tell Ariadne where to keep its metadata and index data. The path must be Hadoop-accessible (S3, ADLS, HDFS, or a local path for testing).
spark.conf.set(
"spark.ariadne.storagePath",
s"abfss://$container@$account.dfs.core.windows.net/ariadne"
)
That's the only required setting. See Configuration for the full list of tuning knobs.
Your first index
Three steps: declare the index, tell it which columns to index and which files to track, then build it.
import dev.cjfravel.ariadne.Index
// 1. Declare the index (name, schema of source files, file format)
val index = Index("orders", orderSchema, "parquet")
// 2. Decide what to index and which files to track
index.addIndex("customer_id")
index.addFile(
"abfss://lake@mystorage.dfs.core.windows.net/orders/2024-q1.parquet",
"abfss://lake@mystorage.dfs.core.windows.net/orders/2024-q2.parquet",
"abfss://lake@mystorage.dfs.core.windows.net/orders/2024-q3.parquet"
)
// 3. Build — scans the files and writes the index to Delta
index.update
Adding files later is fine. Call addFile any time and re-run update — Ariadne only scans files it hasn't seen before.
Index name rules. The name becomes a directory under storagePath — it must be a valid Hadoop file name.
Joining with the index
Once built, the index acts as a stand-in for your data files in a Spark join. There are two equivalent directions:
import dev.cjfravel.ariadne.Index._ // brings in the df.join(index, …) implicit
// DataFrame on the left
val result = customerDf.join(index, Seq("customer_id"), "inner")
// Index on the left
val result = index.join(customerDf, Seq("customer_id"), "inner")
The join returns the minimal complete dataset for the join: every row whose key appears in customerDf is found, and indexed rows whose keys are absent from customerDf are pruned away. Under the hood Ariadne uses the index to skip files that can't contain any of the keys in customerDf.
Pruning happens at file granularity, not row granularity. If a file contains even one matching key, the whole file is read, so unmatched rows that happen to share that file ride along in the result. Which unmatched rows survive therefore depends on how rows are distributed across files, and can change after compact() or after new files are added. Filter on your join keys if you need row-exact results.
Because of this, join types whose answer is defined by unmatched index-side rows are rejected with an UnsupportedJoinTypeException rather than returning a file-layout-dependent result. Which types those are depends on the direction:
| Call | Index side | Supported | Rejected |
|---|---|---|---|
index.join(df, …) | left | inner, left_semi, right, right_outer | left, left_outer, left_anti, full, full_outer, outer |
df.join(index, …) | right | inner, left_semi, left, left_outer, left_anti | right, right_outer, full, full_outer, outer |
Aliases and casing are normalized, so left_outer, leftouter and LEFT_OUTER are all treated alike. To join against every row of the indexed dataset regardless of matches, read the data files directly instead of going through the index.
Reconnecting to an existing index
Once an index exists on disk, you can reconnect to it without restating the schema or format:
val index = Index("orders") // reads metadata, infers everything else
The schema, format, indexed columns, and read options are all loaded from the stored metadata.
What to read next
Index Types
Choose the right index for your data shape — regular, bloom, range, temporal, computed, or exploded.
Compare typesUsage Guide
Column selection, file lookups, JSON, schema evolution, the Ariadne catalog, and Spark SQL access.
Read the guideConfiguration
Every spark.ariadne.* setting, what it does, and when you'd change it.