Cheat sheet

Type Best for False positives?
Regular Low-to-medium cardinality columns (regions, statuses, categories) No
Bloom High-cardinality keys (user IDs, transaction IDs) Yes (configurable)
Range Naturally ordered columns (dates, amounts) No
Temporal Slowly-changing entities (latest-version semantics) No
Computed Indexing a derived value via SQL expression Inherited
Exploded field Joining on elements inside nested array columns No

Mutually exclusive per column. Each column gets exactly one index type. Attempting to add a second type throws IllegalArgumentException. Different columns can use different types within the same index — see Combining types.

Regular index

The simplest option. Tracks the distinct values present in each file, then skips files that can't contain any of your join keys — exact pruning, no false positives.

index.addIndex("region")
index.addIndex("status")

Use when: you need exact (no-false-positive) pruning. Regular indexes handle very large numbers of distinct values per column, but the larger that set grows the more work update does to build the index — when you don't need exact pruning, prefer a bloom index.

Auto-bloom kicks in for you. As soon as a single file contributes spark.ariadne.largeIndexLimit (default 500,000) or more distinct values to an index column, Ariadne marks that column auto-bloom and builds a bloom filter for every file in it, to keep lookups fast. The limit is per (file, column) pair, not a total across the whole column — a column with millions of distinct values spread thinly across many files never crosses it. This covers regular, computed, exploded-field, and temporal index columns (not range, and not binary-typed columns). You don't have to switch to addBloomIndex manually unless you want to control the FPR.

Bloom filter index

A space-efficient probabilistic index for high-cardinality columns. Bloom filters may occasionally include a file that doesn't actually match (false positive) but will never miss a file that does.

index.addBloomIndex("user_id", fpr = 0.01)     // 1% FPR (default)
index.addBloomIndex("session_id", fpr = 0.001) // 0.1% — tighter, larger filter

Use when: the column is high-cardinality (IDs, hashes) and storing every distinct value as an array would be expensive.

False positives mean extra files get read. A 1% FPR is a good default. Tighten it (e.g. 0.001) only when over-reading is more expensive than the larger filter.

Range index

Tracks the minimum and maximum value of a column per file. Files whose range doesn't overlap any query value are skipped entirely.

index.addRangeIndex("event_date")
index.addRangeIndex("amount")

Use when: your data is naturally ordered or batched along a column — daily Parquet drops, monthly extracts, amount-bucketed files. Range indexes are cheap to maintain and very effective on date/time columns.

Temporal index

For columns where the same entity appears in multiple files over time (e.g. nightly snapshots, slowly-changing dimensions), and you only want the latest version. During joins, Ariadne automatically deduplicates — keeping only the row with the latest timestamp per key.

index.addTemporalIndex("user_id", "updated_at")

// If user_id=1 exists in both january.parquet and june.parquet,
// only the June row is returned from the join.

Use when: you have append-only history files but want point-in-time-latest semantics on read. The timestamp column must be present in every file you index.

When a temporal column passes spark.ariadne.largeIndexLimit it gets an auto-bloom filter like any other index type, which lets queries skip whole files before reading the large index. Pruning never changes which version wins: only files that hold none of the queried values are skipped, so they could not have supplied the latest row anyway.

Computed index

Index a value derived from a SQL expression — the computed column doesn't have to physically exist in your files. Useful for partition-like prefixes, hash buckets, or date-component breakdowns.

index.addComputedIndex("category", "substring(Id, 1, 4)")
index.addComputedIndex("year",     "year(event_timestamp)")
index.addComputedIndex("bucket",   "abs(hash(user_id)) % 16")

Use when: the natural join key is a transformation of a stored column.

Exploded field index

Index elements inside array columns. Without this, joining on a nested array field would require reading every file.

// Index users[].id as a virtual column named "user_id"
index.addExplodedFieldIndex("users", "id", "user_id")

// Index tags[].name as "tag_name"
index.addExplodedFieldIndex("tags", "name", "tag_name")

You join on the alias name (user_id, tag_name), not the original nested path.

Use when: you have JSON/Parquet with embedded arrays of structs and want lookups against fields inside those arrays.

Combining types on one index

Different columns within the same index can use different types. On a multi-column join, a file must satisfy all indexed conditions to be loaded.

val index = Index("events", eventSchema, "parquet")
index.addRangeIndex("event_date")     // prune by date overlap
index.addBloomIndex("user_id", 0.01)  // probabilistic filter on user IDs
index.addIndex("region")              // exact filter on region
index.addFile(eventFiles: _*)
index.update

val result = index.join(
  queryDf,
  Seq("event_date", "user_id", "region"),
  "inner"
)

Stack types deliberately. A common, effective combo is range on the partition column + bloom on the join key + regular on the low-cardinality filter. Each type narrows the file set on a different axis.