Required

KeyDefaultDescription
spark.ariadne.storagePath (required) Where Ariadne stores index data. Must be Hadoop-accessible (S3, ADLS, HDFS, or a local path for testing).

Sizing & performance

KeyDefaultDescription
spark.ariadne.largeIndexLimit 500000 Distinct values a single file contributes to a column before that (file, column) pair switches to a separate "large index" Delta table (with row-per-value layout instead of an array column). This is a per-file threshold, not a total across the column.
spark.ariadne.stagingConsolidationThreshold 50 Number of update batches between staging consolidations during update. Lower = more frequent merges, less risk of large staging table; higher = fewer merges, faster overall but more to redo on failure.
spark.ariadne.indexRepartitionCount (not set) Repartition index metadata to N partitions before explode operations in joins. Set this to avoid FetchFailedException on very large indexes. Typical values: 100–500.
spark.ariadne.repartitionDataFiles false Also repartition the loaded data files during joins (not just index metadata). Enable when source files themselves are large enough to shuffle-fail.
spark.ariadne.autoCompactThreshold (not set) Auto-compact (Delta OPTIMIZE) after this many update batches. The counter is persisted across Spark jobs.
spark.ariadne.autoBloomFpr 0.01 False-positive rate for the automatic bloom filters built on columns that exceed largeIndexLimit. This also determines how many distinct query-side values are worth probing: a filter prunes (1 − fpr)n of non-matching files, so past roughly ln(0.01) / ln(1 − fpr) values (458 at the default 1% FPR) it prunes less than 1% of non-matching files and the pre-filter is skipped automatically. Skipping widens the candidate file set but never changes results. That bound is derived, not configurable.

Concurrency & locks

KeyDefaultDescription
spark.ariadne.lockTimeout 1800 Seconds before a lock is considered stale and auto-heals (default: 30 min).
spark.ariadne.lockRetryInterval 60 Base retry interval (seconds) for lock acquisition, using exponential backoff.
spark.ariadne.lockMaxWait 3600 Maximum seconds to wait for a lock before throwing IndexLockException (default: 1 hr).
spark.ariadne.lockRefreshInterval 1 Refresh the lock's lastRefreshedAt timestamp every N batches during long update runs.

Observability

KeyDefaultDescription
spark.ariadne.debug false Log detailed join diagnostics — timing breakdown, file sizes, physical plans. Useful when tuning, noisy in production.

Spark SQL catalog

KeyDefaultDescription
spark.sql.catalog.ariadne (not set) Set to dev.cjfravel.ariadne.catalog.AriadneCatalog to expose indexes as Spark SQL tables.
spark.sql.extensions (not set) Add dev.cjfravel.ariadne.catalog.AriadneSparkExtension to enable JOIN optimization. Must be set on the SparkConf before the SparkContext is created.

Concurrency model

Ariadne uses file-based locks under each index directory to coordinate writes between Spark jobs. Locks are acquired and released automatically by the API methods that need them.

Two locks exist per index:

.filelist.lock
Held during addFile while the tracked-file list is updated.
.update.lock
Held during update, deleteFiles, compact, and vacuum — anything that mutates Delta tables.

If a job crashes mid-operation and leaves a stale lock behind, the next job to attempt acquisition will auto-heal it after lockTimeout seconds (default 30 min). During long-running updates, the held lock is periodically refreshed so honest activity isn't mistaken for a crash.

Single-instance, single-thread for mutation. An Index object is not safe to mutate from multiple threads. For concurrent mutation across jobs, use separate Index instances — they'll coordinate via the file locks. Read-only operations on the same instance from multiple threads are safe so long as nothing is mutating concurrently.