Index Types
Six index types cover different data shapes. A column can have exactly one type — choose based on cardinality and the kind of query you run against it.
Cheat sheet
| Type | Best for | False positives? |
|---|---|---|
| Regular | Low-to-medium cardinality columns (regions, statuses, categories) | No |
| Bloom | High-cardinality keys (user IDs, transaction IDs) | Yes (configurable) |
| Range | Naturally ordered columns (dates, amounts) | No |
| Temporal | Slowly-changing entities (latest-version semantics) | No |
| Computed | Indexing a derived value via SQL expression | Inherited |
| Exploded field | Joining on elements inside nested array columns | No |
Mutually exclusive per column. Each column gets exactly one index type. Attempting to add a second type throws IllegalArgumentException. Different columns can use different types within the same index — see Combining types.
Regular index
The simplest option. Tracks the distinct values present in each file, then skips files that can't contain any of your join keys — exact pruning, no false positives.
index.addIndex("region")
index.addIndex("status")
Use when: you need exact (no-false-positive) pruning. Regular indexes handle very large numbers of distinct values per column, but the larger that set grows the more work update does to build the index — when you don't need exact pruning, prefer a bloom index.
Auto-bloom kicks in for you. As soon as a single file contributes spark.ariadne.largeIndexLimit (default 500,000) or more distinct values to an index column, Ariadne marks that column auto-bloom and builds a bloom filter for every file in it, to keep lookups fast. The limit is per (file, column) pair, not a total across the whole column — a column with millions of distinct values spread thinly across many files never crosses it. This covers regular, computed, exploded-field, and temporal index columns (not range, and not binary-typed columns). You don't have to switch to addBloomIndex manually unless you want to control the FPR.
Bloom filter index
A space-efficient probabilistic index for high-cardinality columns. Bloom filters may occasionally include a file that doesn't actually match (false positive) but will never miss a file that does.
index.addBloomIndex("user_id", fpr = 0.01) // 1% FPR (default)
index.addBloomIndex("session_id", fpr = 0.001) // 0.1% — tighter, larger filter
Use when: the column is high-cardinality (IDs, hashes) and storing every distinct value as an array would be expensive.
False positives mean extra files get read. A 1% FPR is a good default. Tighten it (e.g. 0.001) only when over-reading is more expensive than the larger filter.
Range index
Tracks the minimum and maximum value of a column per file. Files whose range doesn't overlap any query value are skipped entirely.
index.addRangeIndex("event_date")
index.addRangeIndex("amount")
Use when: your data is naturally ordered or batched along a column — daily Parquet drops, monthly extracts, amount-bucketed files. Range indexes are cheap to maintain and very effective on date/time columns.
Temporal index
For columns where the same entity appears in multiple files over time (e.g. nightly snapshots, slowly-changing dimensions), and you only want the latest version. During joins, Ariadne automatically deduplicates — keeping only the row with the latest timestamp per key.
index.addTemporalIndex("user_id", "updated_at")
// If user_id=1 exists in both january.parquet and june.parquet,
// only the June row is returned from the join.
Use when: you have append-only history files but want point-in-time-latest semantics on read. The timestamp column must be present in every file you index.
When a temporal column passes spark.ariadne.largeIndexLimit it gets an auto-bloom filter like any other index type, which lets queries skip whole files before reading the large index. Pruning never changes which version wins: only files that hold none of the queried values are skipped, so they could not have supplied the latest row anyway.
Computed index
Index a value derived from a SQL expression — the computed column doesn't have to physically exist in your files. Useful for partition-like prefixes, hash buckets, or date-component breakdowns.
index.addComputedIndex("category", "substring(Id, 1, 4)")
index.addComputedIndex("year", "year(event_timestamp)")
index.addComputedIndex("bucket", "abs(hash(user_id)) % 16")
Use when: the natural join key is a transformation of a stored column.
Exploded field index
Index elements inside array columns. Without this, joining on a nested array field would require reading every file.
// Index users[].id as a virtual column named "user_id"
index.addExplodedFieldIndex("users", "id", "user_id")
// Index tags[].name as "tag_name"
index.addExplodedFieldIndex("tags", "name", "tag_name")
You join on the alias name (user_id, tag_name), not the original nested path.
Use when: you have JSON/Parquet with embedded arrays of structs and want lookups against fields inside those arrays.
Combining types on one index
Different columns within the same index can use different types. On a multi-column join, a file must satisfy all indexed conditions to be loaded.
val index = Index("events", eventSchema, "parquet")
index.addRangeIndex("event_date") // prune by date overlap
index.addBloomIndex("user_id", 0.01) // probabilistic filter on user IDs
index.addIndex("region") // exact filter on region
index.addFile(eventFiles: _*)
index.update
val result = index.join(
queryDf,
Seq("event_date", "user_id", "region"),
"inner"
)
Stack types deliberately. A common, effective combo is range on the partition column + bloom on the join key + regular on the low-cardinality filter. Each type narrows the file set on a different axis.