Configuration
Every Ariadne setting is a standard Spark configuration property. Only storagePath is required — the rest have sensible defaults.
Required
| Key | Default | Description |
|---|---|---|
spark.ariadne.storagePath |
(required) | Where Ariadne stores index data. Must be Hadoop-accessible (S3, ADLS, HDFS, or a local path for testing). |
Sizing & performance
| Key | Default | Description |
|---|---|---|
spark.ariadne.largeIndexLimit |
500000 |
Distinct values a single file contributes to a column before that (file, column) pair switches to a separate "large index" Delta table (with row-per-value layout instead of an array column). This is a per-file threshold, not a total across the column. |
spark.ariadne.stagingConsolidationThreshold |
50 |
Number of update batches between staging consolidations during update. Lower = more frequent merges, less risk of large staging table; higher = fewer merges, faster overall but more to redo on failure. |
spark.ariadne.indexRepartitionCount |
(not set) | Repartition index metadata to N partitions before explode operations in joins. Set this to avoid FetchFailedException on very large indexes. Typical values: 100–500. |
spark.ariadne.repartitionDataFiles |
false |
Also repartition the loaded data files during joins (not just index metadata). Enable when source files themselves are large enough to shuffle-fail. |
spark.ariadne.autoCompactThreshold |
(not set) | Auto-compact (Delta OPTIMIZE) after this many update batches. The counter is persisted across Spark jobs. |
spark.ariadne.autoBloomFpr |
0.01 |
False-positive rate for the automatic bloom filters built on columns that exceed largeIndexLimit. This also determines how many distinct query-side values are worth probing: a filter prunes (1 − fpr)n of non-matching files, so past roughly ln(0.01) / ln(1 − fpr) values (458 at the default 1% FPR) it prunes less than 1% of non-matching files and the pre-filter is skipped automatically. Skipping widens the candidate file set but never changes results. That bound is derived, not configurable. |
Concurrency & locks
| Key | Default | Description |
|---|---|---|
spark.ariadne.lockTimeout |
1800 |
Seconds before a lock is considered stale and auto-heals (default: 30 min). |
spark.ariadne.lockRetryInterval |
60 |
Base retry interval (seconds) for lock acquisition, using exponential backoff. |
spark.ariadne.lockMaxWait |
3600 |
Maximum seconds to wait for a lock before throwing IndexLockException (default: 1 hr). |
spark.ariadne.lockRefreshInterval |
1 |
Refresh the lock's lastRefreshedAt timestamp every N batches during long update runs. |
Observability
| Key | Default | Description |
|---|---|---|
spark.ariadne.debug |
false |
Log detailed join diagnostics — timing breakdown, file sizes, physical plans. Useful when tuning, noisy in production. |
Spark SQL catalog
| Key | Default | Description |
|---|---|---|
spark.sql.catalog.ariadne |
(not set) | Set to dev.cjfravel.ariadne.catalog.AriadneCatalog to expose indexes as Spark SQL tables. |
spark.sql.extensions |
(not set) | Add dev.cjfravel.ariadne.catalog.AriadneSparkExtension to enable JOIN optimization. Must be set on the SparkConf before the SparkContext is created. |
Concurrency model
Ariadne uses file-based locks under each index directory to coordinate writes between Spark jobs. Locks are acquired and released automatically by the API methods that need them.
Two locks exist per index:
.filelist.lock- Held during
addFilewhile the tracked-file list is updated. .update.lock- Held during
update,deleteFiles,compact, andvacuum— anything that mutates Delta tables.
If a job crashes mid-operation and leaves a stale lock behind, the next job to attempt acquisition will auto-heal it after lockTimeout seconds (default 30 min). During long-running updates, the held lock is periodically refreshed so honest activity isn't mistaken for a crash.
Single-instance, single-thread for mutation. An Index object is not safe to mutate from multiple threads. For concurrent mutation across jobs, use separate Index instances — they'll coordinate via the file locks. Read-only operations on the same instance from multiple threads are safe so long as nothing is mutating concurrently.