Removing files

When source files are archived, replaced, or deleted, remove them from the index so they're not considered during joins:

index.deleteFiles(
  "abfss://lake@mystorage.dfs.core.windows.net/orders/old_data_2023.parquet",
  "abfss://lake@mystorage.dfs.core.windows.net/orders/archived_events.parquet"
)

This removes the corresponding rows from the main index, every large-index table for that column, and the staging table if one exists. It also removes the file from the tracked file list so a future update doesn't try to re-scan it.

Pair with IndexCatalog.findIndexes if you don't know which indexes reference a given file:

val path = "abfss://data@mystorage.dfs.core.windows.net/old_file.parquet"
IndexCatalog.findIndexes(path).foreach { name =>
  IndexCatalog.get(name).deleteFiles(path)
}

Compaction

Repeated update calls accumulate small Delta files. Run compact() periodically to consolidate them:

index.compact()

This runs Delta OPTIMIZE on the main index table and every large-index table. Read performance improves; write performance is unaffected.

Vacuum

Delta keeps old file versions for time-travel. After compaction, those old versions become unreachable but still occupy storage. vacuum removes files older than the retention window:

index.vacuum()                       // default: 168 hours (7 days)
index.vacuum(retentionHours = 72)    // 3-day retention

Be conservative with short retention. Any active read transaction older than the retention window can fail. The Delta default of 7 days exists for good reason — only shorten it if you know nothing in your environment time-travels.

Auto-compaction

For long-running jobs that do many small update calls, enable auto-compaction so you don't have to schedule it separately:

spark.conf.set("spark.ariadne.autoCompactThreshold", "10")

After every 10 update batches, Ariadne runs compact() automatically. The batch counter is persisted in index metadata, so it accumulates correctly across Spark jobs.

If autoCompactThreshold is unset and the counter reaches 50, Ariadne logs a warning recommending compaction. Manual compact() also resets the counter.

Removing an index entirely

Delete the index and all its data:

Index.remove("myIndex")          // delete and remove from catalog
IndexCatalog.remove("myIndex")   // same — equivalent

This deletes the metadata file, all Delta tables (main index, staging, every large-index table), and the tracked file list. The underlying source data files are not touched.

Pruning metrics

Every join logs how effective the index was at pruning files. This is the easiest way to confirm Ariadne is actually doing its job:

Index pruning: loaded 42 of 10000 files (1.23 GB of 45.67 GB) — 97% data pruned

The file counts are exact. The byte figures are not: the total-bytes denominator comes from a running total maintained across update and deleteFiles, so it can drift slightly if files change size in place, which makes the percentage approximate. If you're seeing 0% pruned consistently, your join probably isn't on an indexed column, or every key matches every file — see Troubleshooting.

For more detail, enable debug logging:

spark.conf.set("spark.ariadne.debug", "true")