Joining

An Ariadne index participates in normal Spark joins, in either direction. What the index guarantees is the minimal complete dataset for the join: a matching indexed value is never missed, and indexed rows whose keys are absent from the other DataFrame are pruned away. It does not promise the same row set you would get by joining the raw files — discarding non-matching data is the point of an index.

import dev.cjfravel.ariadne.Index._  // brings in the df.join(index, …) implicit

// DataFrame on the left
val a = customerDf.join(index, Seq("customer_id"), "inner")

// Index on the left
val b = index.join(customerDf, Seq("customer_id"), "inner")

Granularity: files, not rows

The index maps values to files. Pruning therefore drops whole files that cannot contain any join key; it never drops individual rows within a file that survives. Practical consequences:

Supported join types

Because unmatched index-side rows are a file-layout artifact, join types whose result is defined by those rows have no stable meaning and are rejected with an UnsupportedJoinTypeException. Which types are rejected depends on which side the index is on:

CallIndex sideSupportedRejected
index.join(df, …)leftinner, left_semi, right, right_outerleft, left_outer, left_anti, full, full_outer, outer
df.join(index, …)rightinner, left_semi, left, left_outer, left_antiright, right_outer, full, full_outer, outer

The two directions are mirror images: whichever side the index occupies, the outer join that preserves your DataFrame is available and the one that preserves the index is not. Aliases and casing are normalized, so left_outer, leftouter and LEFT_OUTER behave identically.

cross is accepted but not listed above. Both entry points always take join columns, so Spark attaches an equi-join condition and the result is identical to inner — it is not a separate capability.

// Keep every customerDf row, matched or not — index on the left
index.join(customerDf, Seq("customer_id"), "right")

// Same intent with the index on the right
customerDf.join(index, Seq("customer_id"), "left")

// Rejected: asks for indexed rows that have no match in customerDf
index.join(customerDf, Seq("customer_id"), "left_anti")   // UnsupportedJoinTypeException

To work with every row of the indexed dataset regardless of matches, read the data files directly rather than going through the index.

DataFrame API only. The Spark SQL catalog behaves differently: it pre-prunes only for INNER equi-joins and leaves every other join type as a plain full scan, so SQL joins keep standard Spark semantics for all join types — at the cost of no pruning.

Column selection

Reduce I/O by reading only the columns you need from the source files. select mutates and returns the same Index, so it chains cleanly into a join:

val result = index
  .select("user_id", "name", "email")
  .join(queryDf, Seq("user_id"), "inner")

Not thread-safe. select() mutates internal state on the Index instance. Don't share one instance across threads while calling select.

File lookups without joining

Sometimes you just want to know which files would be read for a given set of keys — for audit, prefetching, or building your own job graph.

val files: Set[String] = index.locateFiles(Map(
  "user_id"   -> Array("u1", "u2", "u3"),
  "region"    -> Array("us-west")
))

The map keys are indexed column names; values are arrays of candidate values. Multi-column lookups use AND semantics — a file must match on every key.

JSON & format read options

Pass format-specific read options at index creation. They're persisted in metadata and applied every time the source files are read.

val jsonIndex = Index(
  "events",
  jsonSchema,
  "json",
  readOptions = Map("multiLine" -> "true")
)
jsonIndex.addExplodedFieldIndex("users", "id", "user_id")
jsonIndex.addFile("events.json")
jsonIndex.update

Same mechanism works for CSV ("header" -> "true", "delimiter" -> "|", etc.).

Schema evolution

If your source schema changes — a column is added, types widen — reconnect with the new schema and explicitly opt in to the mismatch:

val index = Index("myIndex", newSchema, "parquet", allowSchemaMismatch = true)

The hard requirement: every previously indexed column must still exist in the new schema. Otherwise you'll get an IndexNotFoundInNewSchemaException and need to remove and rebuild.

Ariadne index catalog

The IndexCatalog object is a stateless directory of every index under your storagePath. Useful for discovery, cleanup, and tooling.

import dev.cjfravel.ariadne.IndexCatalog

IndexCatalog.list()                 // Seq[String] of all index names
IndexCatalog.exists("myIndex")
IndexCatalog.describe("myIndex")    // IndexSummary: types, file count, format
IndexCatalog.describeAll()
IndexCatalog.toDF().show()          // tabular view of all indexes
IndexCatalog.get("myIndex")         // reconnect — same as Index("myIndex")
IndexCatalog.remove("myIndex")      // delete an index and all its data

// Clean up indexes that reference a deleted file
IndexCatalog.findIndexes("abfss://data@mystorage.dfs.core.windows.net/old.parquet").foreach { name =>
  IndexCatalog.get(name).deleteFiles("abfss://data@mystorage.dfs.core.windows.net/old.parquet")
}

Spark SQL catalog

Ariadne can expose every index as a Spark SQL table. Register the catalog and SQL extension on your SparkConf before the SparkContext is created (this is a hard requirement for spark.sql.extensions):

val conf = new SparkConf()
  .set("spark.sql.catalog.ariadne",
       "dev.cjfravel.ariadne.catalog.AriadneCatalog")
  .set("spark.sql.extensions",
       "dev.cjfravel.ariadne.catalog.AriadneSparkExtension")

Every index under storagePath is automatically discoverable — no per-index registration:

SHOW TABLES IN ariadne;
DESCRIBE ariadne.customers;

SELECT * FROM ariadne.customers WHERE id = 123;

SELECT *
FROM ariadne.customers c
JOIN orders o ON c.id = o.customerid;

Read-only catalog. Index creation, updates, and deletion all happen through the Scala API — not through SQL DDL.