Engines
Reble’s engine interface is one method: execute a model’s SQL against pinned input snapshots and write the result to a branch. Two engines implement it today — DuckDB (default, embedded, zero setup) and Spark (embedded PySpark, for transforms that outgrow a single DuckDB process). Same contract either way: scoped runs, tag-pinned inputs, provenance in the snapshot summary, branch writes, main untouched until promote.
Picking
Section titled “Picking”DuckDB is the right default. Reads stream out-of-core through
iceberg_scan and spill under memory_limit — measured, a 1M-row
lifecycle over S3 runs in ~30s from a laptop (numbers).
It starts in milliseconds and has no JVM.
Spark is for the jobs DuckDB genuinely can’t do: joins whose working
set exceeds single-node spill, or transforms that lean on Spark UDFs you
already have. pip install 'reble[spark]', then:
compute_policy: prefer: spark # or per-run: reble run --engine sparkor override per run: reble run --engine spark.
Spark configuration
Section titled “Spark configuration”engines.spark in reble.yml:
engines: spark: master: local[*] # default — no cluster required app_name: reble settings: # raw Spark conf passthrough packages: org.apache.iceberg:iceberg-spark-runtime-3.5_2.12:1.6.1packages defaults to the Iceberg Spark runtime matched to PySpark 3.5
(plus iceberg-aws-bundle for Glue catalogs, plus a JDBC driver for sql
catalogs). First run downloads the jars (ivy cache) — later runs are
cached.
Catalogs
Section titled “Catalogs”The Spark engine maps your existing reble.yml catalog to a Spark catalog and points both engines at the same backend, so branches and tags created by either side are visible to both:
| reble.yml type | Spark catalog | notes |
|---|---|---|
glue |
Glue catalog | aws-bundle jar included |
hive |
Hive catalog | |
rest / polaris / nessie |
REST catalog | same uri |
sql |
JDBC catalog | postgresql only — SQLite holds a database-level lock that Spark’s connection pool and pyiceberg would fight over; a server database lets both writers interleave safely |
Reads resolve through Spark version travel — tbl VERSION AS OF <snapshot_id> — the pinned-snapshot equivalent of DuckDB’s
iceberg_scan(snapshot_from_id=…). Writes go through Iceberg’s
DataFrameWriterV2 with snapshot-property.* options, which is the only
path that carries provenance into the snapshot summary on Spark 3.5.
Configuration keys
Section titled “Configuration keys”Both engines take a free-form settings block under engines in reble.yml.
engines.duckdb
| Key | Default | Meaning |
|---|---|---|
read_mode |
auto |
auto uses iceberg_scan (streaming, out-of-core). arrow materializes through PyArrow instead. |
memory_limit |
— | DuckDB’s memory ceiling, e.g. 4GB. Reads spill past it rather than failing. |
temp_directory |
.reble/spill |
Where spilled data goes. |
settings |
{} |
Raw SET passthrough. Overrides Reble’s S3 auto-configuration, so use it deliberately. |
engines.spark
| Key | Default | Meaning |
|---|---|---|
master |
local[*] |
Spark master URL. |
app_name |
reble |
Application name shown in the Spark UI. |
settings |
{} |
Raw Spark conf. A packages key here replaces the default Iceberg runtime jars. |
Pick the engine with compute_policy.prefer, or per run with
reble run --engine spark. See Configuration.
Limits
Section titled “Limits”- Diffs still run on DuckDB regardless of engine. Diffing is read-only
compute; DuckDB spills under
memory_limit. If real workloads hit that wall, a Spark diff path is a straightforward follow-up. - The Spark engine writes unpartitioned tables, like the DuckDB engine — v0 model outputs are flat full-refresh builds.
- A JVM (17+) is required. macOS/Linux x86_64/arm64 both work.
- Configuration — where these keys live in
reble.yml. - Performance — measured numbers for both engines.