Contributing to Ariadne
Thanks for your interest. This page covers everything you need to file a bug, set up a dev environment, and submit a high-quality pull request.
Reporting bugs & requesting features
๐ Bug report
Open a GitHub issue and include:
- A clear description of the problem
- Steps to reproduce
- Expected vs. actual behavior
- Spark version, Scala version, Ariadne version
โจ Feature request
Open a feature request describing:
- The use case or problem you're trying to solve
- Your proposed solution, if you have one
If AI helped you write it, review it. AI assistance is welcome for drafting bug reports and feature requests, but every issue must reflect genuine human review before it's posted.
- AI-drafted bug reports must include a working reproduction โ a minimal, runnable example that demonstrates the problem on a current release. Speculative bugs that an AI inferred from reading the code, without a repro, will be closed.
- AI-drafted feature requests must articulate a real use case you've actually hit โ not a hypothetical capability gap the AI surfaced. Be ready to discuss it.
- Either way, you're accountable for the content. If you can't answer follow-up questions, the issue isn't ready.
Development setup
Prerequisites
- Java 11
JAVA_HOME=/usr/lib/jvm/java-11-openjdkor equivalent for the deployed Synapse Spark 3.5 line; use Java 21 for Fabric Spark 4.1.- Maven 3.x
- Used for build, test, packaging.
- scalafmt / scalafix / scalastyle
- Enforced by the build. Run
mvn scalafmt:formatto format; scalafix (RemoveUnused+OrganizeImports) and scalastyle (adapted from Apache Spark) run in theverifyphase, somvn verify(and CI) fail on any violation.
Don't upgrade runtime families independently. Spark 3.5 / Scala 2.12.17 / Java 11 track the deployed Azure Synapse runtime, while Spark 4.1 / Scala 2.13.17 / Java 21 track Fabric Runtime 2.0.
Git hooks (recommended)
scalafix and scalastyle run in the verify phase, behind a full compile and test run, so a stray unused import is only reported after a complete CI cycle. The same checks take a few seconds locally. Install the pre-commit hook to run them before the commit instead:
dev/scripts/install-git-hooks.sh
This points core.hooksPath at the version-controlled dev/hooks directory, so the hooks stay in sync with the repository rather than being copied into .git/hooks and going stale. Use --check to see the current state and --uninstall to remove it.
The pre-commit hook exits immediately unless the commit stages Scala sources. When it does, it runs scalafmt:format and re-stages what it reformatted, then compiles and runs scalafix and scalastyle. Files with both staged and unstaged changes are never re-staged automatically โ the hook reports them and stops, because git add would sweep in work you deliberately left out of the commit. scalafmt formats every source directory configured in pom.xml rather than only the staged files, so if it rewrites a Scala file that is not part of your commit the hook names that file and stops, leaving you to stage or revert it.
git commit --no-verify- Skip the hook for one commit.
ARIADNE_SKIP_HOOKS=1- Disable the hook entirely.
ARIADNE_HOOK_SKIP_SCALAFIX=1- Skip only the compile-dependent scalafix step, which is the slowest part.
Build & test
# Run all tests
mvn test
# Run a single test suite
mvn test -Dsuites="dev.cjfravel.ariadne.IndexTests"
# Run one test within a suite (space separator; @ means exact-name match)
mvn test -Dsuites="dev.cjfravel.ariadne.IndexTests @my test name"
# Drop the @ to match any test whose name contains the string
mvn test -Dsuites="dev.cjfravel.ariadne.IndexTests my test"
# Package (produces the shaded JAR)
mvn package
The test phase also runs dev/scripts/readme-has-version.sh, which fails if the version in pom.xml isn't reflected in the docs. Always update both when bumping the version.
Code style
- Format with scalafmt (
mvn scalafmt:format). Imports are organized and unused code removed by scalafix (RemoveUnused+OrganizeImports), and scalastyle (adapted from Apache Spark) enforces static checks. All three run in the build and CI, so a violation fails the build. Install the pre-commit hook to catch violations before they reach CI. - Standard Scala naming:
camelCasefor methods/vals,PascalCasefor types. - Expression-oriented style โ prefer
if/elseexpressions overreturn. - Use
import scala.collection.JavaConverters._(the project is Scala 2.12;CollectionConvertersis 2.13+). - Every public class, trait, object, and method must have Scaladoc.
logger.warn(...)for normal operational messages โ this is intentional and matches how Spark surfaces logs in cluster environments. Reservedebugfor verbose/low-value entries.- No silent
catchblocks. Always log the exception before swallowing or rethrowing.
Backwards compatibility
Ariadne aims to be backwards compatible between releases โ existing indexes should keep working when users upgrade, and the public Scala API shouldn't break under feet. Treat that as the default expectation when proposing changes.
That said, Ariadne is still in beta. If a breaking change clearly benefits the project more than the disruption costs, it can land โ but only with the maintainer's explicit sign-off. Don't assume a breaking change will be accepted just because it's cleaner; open an issue or discussion to align first.
Shading & relocation
Ariadne bundles Guava and Gson, relocated under dev.cjfravel.ariadne.shaded.*, so they can't collide with whatever versions the host Spark runtime ships. Relocation is easy to break in two different directions, and both have caused real bugs here.
Adding a bundled dependency
A new relocated dependency needs three coordinated edits, not one: a <relocation> entry in the shade plugin config, a matching <artifactSet><include> entry, and a reject_entry assertion in dev/scripts/package-contents-tests.sh confirming the original package is absent from the packaged jar. Miss one and the package ships unrelocated and collides at runtime โ which is exactly how com.google.thirdparty, error_prone_annotations and j2objc-annotations each had to be retrofitted after the fact.
Staying safe when consumers relocate Ariadne
Downstream builds sometimes shade Ariadne itself into their own namespace. The Maven Shade Plugin rewrites bytecode references but leaves the Scala ScalaSignature pickle byte-for-byte identical, so a relocated class loads under its new package while its pickle still describes the old one. Any type resolved through Scala reflection then can't be found, usually surfacing as ENCODER_NOT_FOUND during an index update.
CI cannot catch this. Ariadne's own build never relocates dev.cjfravel.ariadne, so a change that breaks under downstream relocation passes every check and fails only in a consumer's shaded build. It has to be caught by review.
- Never use an Ariadne-defined case class as a Spark
Datasetor encoder element type. Use aDataFramewith an explicitStructType, or build the encoder explicitly withEncoders.tuple,Encoders.STRINGand friends. - Prefer the two-argument
udaf(aggregator, inputEncoder)overload โ the single-argument one derives the input encoder reflectively. .as[T],.toDS()andemptyDataset[T]are safe only for standard-library types such asStringand tuples of primitives, which relocation never rewrites.Encoders.javaSerialization[T]andEncoders.kryo[T]are safe โ they take aClassTag, which compiles to aclassOfreference that shade does rewrite. OnlyTypeTag-based derivation is unsafe.- Never hardcode
dev.cjfravel.ariadnein a string used for lookup (Class.forName, Spark extension or config keys, service-loader entries). Relocation rewrites the class, not the string. Package names inside Scaladoc examples are fine.
To verify a change by hand, package target/classes into a jar, install it locally, then shade it in a scratch project with a <relocation> onto a different prefix and exercise the affected code path.
Submitting a pull request
- Fork the repository and branch from
main. - Make your changes, ensuring
mvn testpasses. - Update documentation if your changes affect the public API.
- If you changed any public class, trait, method, or scaladoc, no extra step is needed โ the API reference is generated and published automatically from
main. Preview it locally withdev/scripts/build-docs-site.sh. - Don't change the version in
pom.xmlor the docs โ versioning is managed by the maintainer. - Open a pull request with a clear description.
Contributor License Agreement. By submitting a PR you agree to the terms of the CLA. PRs cannot be merged until the CLA is accepted.
AI-assisted contributions
AI assistance (Copilot, ChatGPT, etc.) is welcome โ but every PR must reflect genuine human understanding:
- Review and comprehend all AI-generated code before submitting. If you can't explain what it does and why, it's not ready.
- Purely AI-generated PRs with no human review will be closed. We need contributors who can discuss their changes, respond to feedback, and reason about edge cases.
- Attribute AI assistance honestly. There's no stigma โ just be transparent.
The goal is quality contributions from people who understand what they're submitting, regardless of how they got there.
Questions
Open a GitHub Discussion.