Skip to content

TEST DO NOT MERGE - Delta pipeline and cdf#12582

Draft
felipepessoto wants to merge 26 commits into
apache:mainfrom
felipepessoto:delta_pipeline_and_cdf
Draft

TEST DO NOT MERGE - Delta pipeline and cdf#12582
felipepessoto wants to merge 26 commits into
apache:mainfrom
felipepessoto:delta_pipeline_and_cdf

Conversation

@felipepessoto

Copy link
Copy Markdown
Contributor

What changes are proposed in this pull request?

How was this patch tested?

Was this patch authored or co-authored using generative AI tooling?

malinjawi and others added 19 commits June 16, 2026 22:36
Delta change-data-feed reads enter Spark as a CDCReader.DeltaCDFRelation,
which does not have the FileSourceScanExec + DeltaParquetFileFormat shape
that Gluten's Delta scan offload recognizes. Add a Gluten Delta planner
strategy (wired from VeloxDeltaComponent) that recognizes DeltaCDFRelation,
expands it through Delta's own CDF batch planning, and rewrites the
projection/filter attributes onto the expanded plan, so the existing Delta
scan offload path can plan the underlying CDF file scans as
DeltaScanTransformer.

CDF over a deletion-vector-enabled table is kept on Spark: native CDF lacks
the DV-aware row-level reconciliation, so an offloaded CDF scan would emit
still-live rows as `delete` change rows. Re-planning Delta's analyzed CDF
batch plan also drops that reconciliation on some Delta versions (Delta 2.4 /
Spark 3.4), where the remove side would surface every row of a logically
removed file as a `delete` change row. The planner strategy therefore declines
to intercept CDF reads whose table has deletion vectors enabled and leaves them
to Delta's own DeltaCDFRelation scan, which reconciles DVs correctly on every
supported Delta version; DeltaScanTransformer additionally guards both CDF scan
sides -- the add side (CdcAddFileIndex) and the remove side
(TahoeRemoveFileIndex) -- as a backstop. Normal (non-CDF) DV scans are
unaffected and continue to apply the DV natively.

Addresses apache#12195.
Inspect AddFile and RemoveFile actions in the requested CDF version range instead of relying on the table property, which only controls future DV writes.\n\nCover both DV-enabled ranges without DV actions and existing DV actions after the property is disabled.
…es baseline

Run delta-io/delta's `spark` ScalaTest suite against a Gluten Velox bundle in CI
and gate the results against a committed baseline so the many expected Delta-on-
Gluten failures stay manageable and can be fixed incrementally without letting
currently-passing tests silently regress.

What it adds (.github/workflows/util/delta-spark-ut/):
- delta_spark_ut.yml: builds the native lib + Gluten bundle, then runs the Delta
  spark suite sharded by suite into 4 shards x 4 forked test JVMs (~16-way), and
  gates each shard against the baseline.
- compare-test-results.py: the gate. Per shard, regressions (failed not in the
  baseline) fail the build; newly-passing baselined tests are flagged so the
  baseline can be tightened. Also supports seed/aggregate modes.
- known-failures.txt: the committed baseline of expected failures.
- setup-delta.sh: clones Delta, injects the Gluten bundle, patches
  DeltaSQLCommandTest, and force-fails the two DeletionVectorsSuite 2B-row tests
  whose native row-index materialization OOM-kills the runner and hangs the shard.
- README.md: how the pipeline, gating and baseline-refresh work.

The workflow also carries a hang watchdog that thread-dumps and kills a wedged
fork, and tunes the per-fork heap (2G) and off-heap (2G) to fit the ~16G runner.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
…line

Delta's data-skipping, limit-push-down, column-pruning and scan-metric tests
collect file-source scans by matching the concrete `FileSourceScanExec` case
class. Under the Gluten Velox bundle the scan is offloaded to
DeltaScanTransformer, a sibling that implements the same `FileSourceScanLike`
interface but is not FileSourceScanExec, so the match misses and the scan
looks absent. This surfaced as `scala.MatchError: List()` (~56
DataSkipping*/DeltaLimitPushDown* tests), empty generated-column partition
filters (~45 OptimizeGeneratedColumnSuite tests) and broken column-pruning /
scan-metric checks across the Delete, Update, Merge, DeletionVectors and
RowId suites and the TestsStatistics helper.

Gluten copies `partitionFilters` and the other accessors these tests read
verbatim onto the offloaded scan, so results are identical to vanilla -- only
the test's `case` match breaks. Fix it by cherry-picking the two merged
upstream Delta commits that widen these matches to the shared
`FileSourceScanLike` interface (behavior-preserving for vanilla, which also
implements it):

  * delta-io/delta#7104 -- ScanReportHelper.collectScans
  * delta-io/delta#7105 -- the remaining 9 test sources, its follow-up

Both are merged on Delta master but land after the ref this workflow builds
against (v4.2.0), so setup-delta.sh cherry-picks them onto the shallow
checkout. Each fetches the fix commit at depth 2 (commit + parent) so
cherry-pick can compute the parent->fix diff, and uses `cherry-pick -n` so no
committer identity is required. Once the pinned DELTA_REF advances to include
a commit its cherry-pick becomes a clean no-op and that block can be removed.

The cherry-picks run before the DeletionVectorsSuite 2B-row force-fail step:
that step sed-injects fail() into DeletionVectorsSuite.scala, which
delta-io/delta#7105 also edits, and git cherry-pick refuses to apply onto a
working tree with uncommitted changes to a file it touches (exit 128).

Refresh known-failures.txt from run 28299900971 (the delta-spark-aggregate job
output), which ran all 19073 tests across 16 shards: removes 187 now-passing
tests with 0 regressions, 963 -> 776. ~147 come from the fixes above
(DataSkipping*, DeltaLimitPushDown*, OptimizeGeneratedColumnSuite, MergeInto*,
RowIdSuite); the remaining ~40 are other suites that now pass (e.g.
HiveConvertToDeltaSuite, BitmapAggregatorE2ESuite). Verified against the
per-shard ran/failed lists: every baseline entry was observed this run (0
stale), so nothing was dropped due to a crashed or incomplete shard.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Make delta_spark_ut.yml a reusable workflow (on: workflow_call) and call it from
velox_backend_x86.yml so the Delta tests reuse the native lib + arrow jars that
workflow already builds, instead of duplicating the build-native-lib-centos-7
job. GitHub artifacts cannot be shared across workflows, so the only way to
reuse the artifact is to run the Delta jobs in the same workflow run.

delta_spark_ut.yml keeps a workflow_dispatch trigger for standalone manual runs
(its build-native-lib-centos-7 job is gated to that case and skipped when
called); the pull_request trigger is removed so the suite no longer double-runs.
velox_backend_x86.yml gains an arrow-jars upload on its native build and a
delta-spark-ut job that calls the reusable workflow. That job runs on every
velox trigger like the other spark-test jobs, since core/velox/substrait/cpp
changes can affect Delta query offload.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address PR review feedback:

- setup-delta.sh: replace the shallow-clone + full-clone fallback (which ran a
  destructive `rm -rf "$DELTA_DIR"`) with a single `git init` + shallow
  `fetch --depth 1 origin "$DELTA_REF"` + `checkout FETCH_HEAD`. This resolves a
  tag, branch, or commit SHA uniformly (`git clone --branch` rejects SHAs),
  drops the dead fallback branch, and removes the unguarded recursive delete.

- compare-test-results.py: in enforce mode, a missing/typoed --known-failures
  path made load_entries() return an empty set, silently degrading to seed mode
  and passing the gate without enforcing regressions. Treat a missing baseline
  file as a configuration error (exit 2); an existing-but-empty file is still
  allowed and legitimately seeds.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Address PR review feedback with four robustness fixes:

- compare-test-results.py (enforce/seed): raise NoReportsError and exit 2 when no
  JUnit <testsuite> elements are parsed, instead of warning and returning empty
  sets. Otherwise a misconfiguration (wrong reports dir, broken reporter, suites
  crashing before writing XML) yields zero failures -> zero regressions -> a
  silent green gate.
- compare-test-results.py (aggregate): exit 2 before writing baseline-out when no
  per-shard failures-*.txt / ran-*.txt inputs are found. The gate-list download is
  continue-on-error and aggregate runs with if: always(), so missing artifacts
  would otherwise produce an empty baseline that could be committed, wiping
  known-failures.txt.
- setup-delta.sh: pass the Delta ref after `--` in git fetch so a ref starting
  with `-` can't be misread as a git option (the script is workflow_dispatch-
  runnable with a user-supplied ref).
- velox_backend_x86.yml: drop secrets: inherit from the reusable Delta UT call.
  delta_spark_ut.yml references no secrets, so inheriting them needlessly forwards
  all caller secrets to a workflow that clones and runs external code.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Some Delta-on-Gluten MERGE tests that write deletion vectors fail
non-deterministically: the same bundle passes them on one CI run and
fails them on the next (native RoaringBitmapArray addSafe aborting on an
invalid Long.MAX_VALUE row index). Such tests cannot live in
known-failures.txt -- baselining them reds the gate on every run where
they pass, and leaving them out reds it on every run where they fail.

Add a flaky-tests.txt quarantine list read by the gate. A quarantined
test is neutral: it never counts as a regression when it fails nor as
now-passing when it passes, and is excluded from the regenerated
baseline (aggregate mode). The suite portion of each entry is an fnmatch
glob so one line covers a root-cause family across generated suite
variants (e.g. *DVs*Suite); the test name is matched exactly.

Seed the list with the DV-merge family behind the native row-index bug.
This is an interim measure -- entries should be removed once that bug is
fixed in the native backend so the tests are enforced again.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
…port fail-fast, baseline count

Fixes three open review comments on the Delta Spark UT pipeline:

1. spark_version was redundant with the hard-coded GLUTEN_BUNDLE_SPARK_VERSION /
   GLUTEN_SPARK_PROFILE, and a workflow_dispatch run could set them out of sync so
   the Delta tests ran against a mismatched Gluten bundle. Make spark_version the
   single source of truth: the Gluten bundle profile (-Pspark-<v>), the bundle jar
   name and Delta's -DsparkVersion are all derived from it, so a mismatch is now
   impossible (no guard needed). Scala 2.13 / JDK 17 stay pinned.

2. Corrupt/truncated JUnit reports (compare-test-results.py). parse_reports
   previously warned and skipped any XML that failed to parse, so a report
   truncated by a killed/OOM'd fork could silently drop a suite's failures and
   let the gate go green on partial data. A TEST-*.xml that fails to parse now
   raises CorruptReportError (exit 2); other XML matched by the broad target/**
   glob is still skipped.

3. Hard-coded failure count in the baseline header (known-failures.txt). The
   header stated a fixed total that drifts every time the baseline is refreshed.
   Drop the number and point at the entry count / delta-spark-aggregate summary
   as the authoritative source instead.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The aggregate job downloads per-shard gate lists with continue-on-error and the
aggregator only hard-failed when *no* inputs were found. If a shard died before
writing its gate lists (OOM / watchdog) or its artifact failed to download, the
job still regenerated a baseline that silently omitted that shard's failures,
shrinking known-failures.txt and reddening the next run.

Add --expected-shards to compare-test-results.py: in aggregate mode it counts
shards that produced a complete failures-*/ran-* pair and refuses to aggregate
(exit 2, before writing --baseline-out) when fewer than expected are present. A
shard counts only when BOTH files exist, so a partial download is caught too.
The workflow passes --expected-shards ${{ env.DELTA_NUM_SHARDS }}; omitting the
flag (default 0) disables the check, preserving prior behavior.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
… body

Two readability/cleanup changes to the Delta Spark UT pipeline:

1. Drop redundant Arrow jar handling. Since apache#12244
   (2026-06-09) Gluten depends on vanilla Apache Arrow instead of the
   custom 15.0.0-gluten rename, so the Arrow jars resolve from Maven
   Central. Remove the arrow-jar staging in both native builds, the
   velox-arrow-jars / delta-spark-ut-arrow-jars uploads, the caller's
   arrow_jars_artifact input, and the bundle job's "Download Arrow jars"
   step. The bundle build's Maven cache still holds Arrow across runs and
   resolves it from Central on a miss.

2. Extract the shard test body into run-delta-tests.sh. The "Run Delta
   spark module tests" step embedded ~190 lines of shell (hang watchdog,
   sbt invocation with JVM/heap tuning, cgroup memory forensics, report
   check, known-failures gate). Move that whole body into
   util/delta-spark-ut/run-delta-tests.sh so the workflow step is just an
   `env:` block plus a one-line script call, shrinking delta_spark_ut.yml
   by ~180 lines. The move is faithful: the script body is the previous
   inline block with only the GitHub `${{ }}` expressions replaced by env
   vars (matrix.shard -> SHARD_ID which is already job-level env;
   spark_version/update_baseline/fail_on_fixed added to the step env).
   Verified by a textual diff against the old body and by end-to-end runs
   with a stubbed sbt (baseline-only -> green, regression -> red, seed ->
   green).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The Gluten/JDK17 test JVM flags (--add-opens + the Netty reflection
property, mirroring root pom.xml's extraJavaTestArgs) lived only in the
workflow step's env: block, so a developer running the Delta suite
locally had no shared source for them.

Move them into util/delta-spark-ut/java-test-args.sh, which exports
JAVA_TOOL_OPTIONS. run-delta-tests.sh sources it (via a BASH_SOURCE-
relative path so it also works for local runs), and the workflow env
block no longer sets JAVA_TOOL_OPTIONS. A developer can source the same
file before `sbt spark/test` to get identical flags; README documents it.

Faithful: the sourced value is byte-identical to the previous env-block
value (verified), and an end-to-end run with a stubbed sbt confirms the
forked test JVM still receives all 18 flags with JAVA_TOOL_OPTIONS unset
in the environment.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The native Delta DV bitmap row-index bug (RoaringBitmapArray aborting on a
Long.MAX_VALUE row index during a MERGE that writes deletion vectors) is
intermittent and lands on a DIFFERENT *DVs*Suite MERGE test each run, so
listing every test in flaky-tests.txt was whack-a-mole (6 entries and
counting).

Add error-signature quarantine: the gate now records each failed test's
JUnit <failure>/<error> text, and flaky-error-patterns.txt holds regexes
matched against it. A failure matching a pattern is treated as flaky
regardless of which test it hit -- neither counted as a regression nor
written to the shard's failures list (so it can't leak into the
regenerated baseline). Seeded with the RoaringBitmapArray signature.

This is more precise than a name glob: a different real failure in the
same DV suite is still caught, because only failures carrying the
signature are ignored. The six *DVs*Suite name entries are removed from
flaky-tests.txt (superseded); the name mechanism stays for flakes without
a distinctive error signature.

Verified against a real failing shard report: the DV failure is a
regression without the patterns file and quarantined (gate green) with it.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Run 29042495519 regressed on another DV-merge test whose failure was NOT
the RoaringBitmapArray "exceeds max representable value" signature but a
second, distinct native error from the same DV bitmap aggregation --

  Reason: Delta bitmap row index cannot be negative: -6254810385378525259
  Expression: value >= 0
  Function: addRowIndex  File: .../delta/DeltaBitmapAggregator.cc:44

vs the previously-seen too-large variant (value <= kMaxRepresentableValue,
RoaringBitmapArray.cpp addSafe). Same root cause (garbage row index into
the DV bitmap aggregator), different bounds check and different message.

Add a second explicit pattern for it rather than broadening the existing
one -- the two are distinct native error strings, so keeping each pattern
tightly bound to its message avoids masking an unrelated failure that
merely mentions a bitmap row index. Verified against the real failing
report: the negative-index failure is a regression without the pattern and
quarantined with it; the suite's five other (non-bitmap) baseline failures
and benign controls are not matched.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
The Velox Variant/UDT type-offload gap was fixed upstream, so 74 Delta UT
tests (65 Variant + 9 user-defined-type) that were expected failures now pass
and were tripping the fail-on-fixed gate as "now-passing".

Remove them from known-failures.txt (810 -> 736). The 6 remaining
`variant auto compact` entries are a different root cause (auto-compact
metrics) and correctly stay in the baseline.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098
DeltaFastDropFeatureSuite "We do not create redundant DV tombstones after
cloning isShallowClone: true" now passes consistently (green on 2026-07-13 and
2026-07-14 runs); the earlier FileNotFoundException on a deletion-vector .bin
during the DROP FEATURE parallel OPTIMIZE no longer reproduces. Keeping it in
the baseline trips the fail-on-fixed gate on every run, so remove it (736 ->
735).

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 66a9e40f-3ac8-45be-8fee-a606a22fa098
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@felipepessoto
felipepessoto force-pushed the delta_pipeline_and_cdf branch from 8c21e81 to 8b3c652 Compare July 20, 2026 19:33
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@felipepessoto
felipepessoto force-pushed the delta_pipeline_and_cdf branch from 8b3c652 to 1bde8e7 Compare July 20, 2026 19:36
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

felipepessoto and others added 3 commits July 24, 2026 23:41
…Delta-scoped)

Delta's upstream CDF filter-pushdown tests (DeltaCDC*Suite "filters should be
pushed down" / "select individual column should push down filters" / "filters
with special characters in name should be pushed down") assert the executed plan
prints `PushedFilters: [*IsNotNull(id), *LessThan(id,5)]`.

The leading `*` is Spark's RowDataSourceScanExec convention for a filter the
source fully evaluates itself (so Spark needs no separate post-scan Filter).
Gluten offloads the CDF read to DeltaScanTransformer (a FileSourceScanLike),
which pushes every conjunct into the native scan via PushDownFilterToScan and
applies them as exact row-level filters -- the paired FilterExecTransformer is a
no-op (FilterExecTransformerBase.isNoop), the same end state the `*` advertises.
But the inherited FileSourceScanLike rendering leaves PushedFilters unmarked, so
the assertion fails once the scan is offloaded.

Post-process DeltaScanTransformer's rendered plan string (`simpleString`) to
mark the PushedFilters entries with `*`. `metadata` is a lazy val and cannot be
super-overridden, so the marking is applied to the rendered node string. This is
Delta-scoped: only Delta scans are affected; other file-source scans are
unchanged.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 02286396-395f-41db-ad89-ad69a192cded
…string (Delta-scoped)"

This reverts commit 2dbb578. The Delta-scoped marking is undone so the next
commit can apply the equivalent marking consistently across all Gluten
file-source scans (not just Delta) in the shared FileSourceScanExecTransformerBase.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 02286396-395f-41db-ad89-ad69a192cded
… scans

Consistent, all-scans version of the CDF filter-pushdown rendering fix (the
Delta-scoped variant was added then reverted in the previous two commits).

Move the marking into the shared FileSourceScanExecTransformerBase so every
Gluten file-source scan -- not just Delta -- renders `PushedFilters` with a `*`
on each entry, matching Spark's RowDataSourceScanExec convention for a filter the
source fully evaluates itself. Gluten pushes every conjunct into the native scan
via PushDownFilterToScan and applies them as exact row-level filters, so the
paired FilterExecTransformer is a no-op (FilterExecTransformerBase.isNoop) -- the
same end state the `*` advertises.

`metadata` is a lazy val and cannot be super-overridden, so the mark is applied
to the rendered plan string in both paths that print it:
  - simpleString (used by executedPlan.toString, e.g. Delta's CDF UTs), and
  - verboseStringWithOperatorId (used by the formatted-explain plan-stability
    golden files).

Golden plans updated accordingly (TPC-H approved-plan, TPC-DS plan-stability, and
gluten-tpch-plan-stability across Spark 3.3/3.4/3.5/4.0/4.1): 6264 PushedFilters
entries across 998 files now carry `*`. The goldens contain no DataSourceV2
BatchScan or vanilla FileSourceScanExec PushedFilters, so every entry originates
from a FileSourceScanExecTransformer and marking all of them reproduces the new
rendering exactly. The authoritative way to refresh them remains re-running the
plan-stability suites with GLUTEN_UPDATE=1; the edits here were produced by the
same deterministic marking the code applies.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: 02286396-395f-41db-ad89-ad69a192cded
@github-actions github-actions Bot added the CORE works for Gluten Core label Jul 24, 2026
@github-actions

Copy link
Copy Markdown

Run Gluten Clickhouse CI on x86

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants