Usage Guide
Everything you'll do with an existing Ariadne index — joins, column selection, file lookups, catalog access, and SQL.
Joining
An Ariadne index participates in normal Spark joins, in either direction. What the index guarantees is the minimal complete dataset for the join: a matching indexed value is never missed, and indexed rows whose keys are absent from the other DataFrame are pruned away. It does not promise the same row set you would get by joining the raw files — discarding non-matching data is the point of an index.
import dev.cjfravel.ariadne.Index._ // brings in the df.join(index, …) implicit
// DataFrame on the left
val a = customerDf.join(index, Seq("customer_id"), "inner")
// Index on the left
val b = index.join(customerDf, Seq("customer_id"), "inner")
Granularity: files, not rows
The index maps values to files. Pruning therefore drops whole files that cannot contain any join key; it never drops individual rows within a file that survives. Practical consequences:
- A file containing one matching key is read in full, so unmatched rows sharing that file appear in the result.
- Which unmatched rows appear depends on physical file layout, so a row count can change after
compact(), after adding files, or after re-partitioning the source data — even though the underlying data is unchanged. - Matched rows are always exact and stable. Only the incidental extras vary.
- Add an explicit predicate on the join keys if you need row-exact output that is independent of file layout.
Supported join types
Because unmatched index-side rows are a file-layout artifact, join types whose result is defined by those rows have no stable meaning and are rejected with an UnsupportedJoinTypeException. Which types are rejected depends on which side the index is on:
| Call | Index side | Supported | Rejected |
|---|---|---|---|
index.join(df, …) | left | inner, left_semi, right, right_outer | left, left_outer, left_anti, full, full_outer, outer |
df.join(index, …) | right | inner, left_semi, left, left_outer, left_anti | right, right_outer, full, full_outer, outer |
The two directions are mirror images: whichever side the index occupies, the outer join that preserves your DataFrame is available and the one that preserves the index is not. Aliases and casing are normalized, so left_outer, leftouter and LEFT_OUTER behave identically.
cross is accepted but not listed above. Both entry points always take join columns, so Spark attaches an equi-join condition and the result is identical to inner — it is not a separate capability.
// Keep every customerDf row, matched or not — index on the left
index.join(customerDf, Seq("customer_id"), "right")
// Same intent with the index on the right
customerDf.join(index, Seq("customer_id"), "left")
// Rejected: asks for indexed rows that have no match in customerDf
index.join(customerDf, Seq("customer_id"), "left_anti") // UnsupportedJoinTypeException
To work with every row of the indexed dataset regardless of matches, read the data files directly rather than going through the index.
DataFrame API only. The Spark SQL catalog behaves differently: it pre-prunes only for INNER equi-joins and leaves every other join type as a plain full scan, so SQL joins keep standard Spark semantics for all join types — at the cost of no pruning.
Column selection
Reduce I/O by reading only the columns you need from the source files. select mutates and returns the same Index, so it chains cleanly into a join:
val result = index
.select("user_id", "name", "email")
.join(queryDf, Seq("user_id"), "inner")
Not thread-safe. select() mutates internal state on the Index instance. Don't share one instance across threads while calling select.
File lookups without joining
Sometimes you just want to know which files would be read for a given set of keys — for audit, prefetching, or building your own job graph.
val files: Set[String] = index.locateFiles(Map(
"user_id" -> Array("u1", "u2", "u3"),
"region" -> Array("us-west")
))
The map keys are indexed column names; values are arrays of candidate values. Multi-column lookups use AND semantics — a file must match on every key.
JSON & format read options
Pass format-specific read options at index creation. They're persisted in metadata and applied every time the source files are read.
val jsonIndex = Index(
"events",
jsonSchema,
"json",
readOptions = Map("multiLine" -> "true")
)
jsonIndex.addExplodedFieldIndex("users", "id", "user_id")
jsonIndex.addFile("events.json")
jsonIndex.update
Same mechanism works for CSV ("header" -> "true", "delimiter" -> "|", etc.).
Schema evolution
If your source schema changes — a column is added, types widen — reconnect with the new schema and explicitly opt in to the mismatch:
val index = Index("myIndex", newSchema, "parquet", allowSchemaMismatch = true)
The hard requirement: every previously indexed column must still exist in the new schema. Otherwise you'll get an IndexNotFoundInNewSchemaException and need to remove and rebuild.
Ariadne index catalog
The IndexCatalog object is a stateless directory of every index under your storagePath. Useful for discovery, cleanup, and tooling.
import dev.cjfravel.ariadne.IndexCatalog
IndexCatalog.list() // Seq[String] of all index names
IndexCatalog.exists("myIndex")
IndexCatalog.describe("myIndex") // IndexSummary: types, file count, format
IndexCatalog.describeAll()
IndexCatalog.toDF().show() // tabular view of all indexes
IndexCatalog.get("myIndex") // reconnect — same as Index("myIndex")
IndexCatalog.remove("myIndex") // delete an index and all its data
// Clean up indexes that reference a deleted file
IndexCatalog.findIndexes("abfss://data@mystorage.dfs.core.windows.net/old.parquet").foreach { name =>
IndexCatalog.get(name).deleteFiles("abfss://data@mystorage.dfs.core.windows.net/old.parquet")
}
Spark SQL catalog
Ariadne can expose every index as a Spark SQL table. Register the catalog and SQL extension on your SparkConf before the SparkContext is created (this is a hard requirement for spark.sql.extensions):
val conf = new SparkConf()
.set("spark.sql.catalog.ariadne",
"dev.cjfravel.ariadne.catalog.AriadneCatalog")
.set("spark.sql.extensions",
"dev.cjfravel.ariadne.catalog.AriadneSparkExtension")
Every index under storagePath is automatically discoverable — no per-index registration:
SHOW TABLES IN ariadne;
DESCRIBE ariadne.customers;
SELECT * FROM ariadne.customers WHERE id = 123;
SELECT *
FROM ariadne.customers c
JOIN orders o ON c.id = o.customerid;
Read-only catalog. Index creation, updates, and deletion all happen through the Scala API — not through SQL DDL.