Triage pass over the open requires-triage queue, per the project Bug Triage Guide.
- Date: 2026-07-20
- Total issues processed: 20 (15 triaged, 5 skipped, 0 failed)
- Type counts: 11 bugs, 4 enhancements
- Priority counts applied:
priority:critical 3, priority:high 1, priority:medium 2, priority:low 5
- Guide: docs/source/contributor-guide/bug_triage.md
Labels have already been applied and requires-triage removed from each issue listed under "Bugs" and "Enhancements". A reviewer should spot-check the calls and close this issue when satisfied. To correct a label, edit the affected issue directly.
Bugs
priority:critical
- Spark 4.2: Comet native collect_set does not normalize NaN / -0.0 (SPARK-57298) (#4966)
- Area labels:
area:aggregation, spark sql tests
- Rationale: Spark 4.2 normalizes NaN and -0.0 in
collect_set; Comet's native aggregate does not, so results silently diverge from Spark — decision-tree step 1 (silent wrong results).
- cast string to boolean: trim ISO control bytes to match Spark's UTF8String.trimAll (#4959)
- Area labels:
area:expressions
- Rationale: reporter demonstrates Comet casts
" \x01true\x01 " to null where Spark casts to true — silent divergence in a widely-used cast path (matches the guide's "incorrect cast behavior" critical example).
- Native implementation of
get_json_object returns last value for duplicate keys, Spark returns first (#4947)
- Area labels:
area:expressions
- Rationale: for
{"a":1,"a":2} with path $.a Comet returns 2 while Spark returns 1; silent wrong result on a documented Spark semantic — decision-tree step 1.
priority:high
- Comet throws
java.lang.NegativeArraySizeException (#4944)
- Area labels:
area:scan
- Rationale: protobuf
AbstractMessageLite.toByteArray overflows to a negative size while serializing native Iceberg scan partitions on a large real workload; visible crash rather than wrong result — decision-tree step 2.
priority:medium
- Spark 4.2: ANSI arithmetic overflow tests fail under Comet (#4967)
- Area labels:
area:expressions
- Rationale:
CometExpressionSuite ANSI-overflow tests are assume-gated off on Spark 4.2 pending investigation of behavior differences; today's workaround is the guard, but the underlying divergence could affect ANSI-mode users — functional gap with workaround per decision-tree step 3.
- Spark 4.2:
OneRowRelation in Union branches forces Union and downstream aggregates off Comet (TPC-DS q77a) (#4949)
- Area labels:
area:scan
- Rationale: unsupported
OneRowRelation scan cascades fallback across a whole Union + downstream aggregates on Spark 4.2's q77a plan; correctness is preserved via Spark, but Comet acceleration is lost — decision-tree step 3 (performance regression with workaround).
priority:low
- Spark 4.2: native Iceberg REST catalog scan test fails under Comet (#4969)
- Area labels:
area:scan
- Rationale:
CometIcebergNativeSuite "REST catalog with native Iceberg scan" is assume-skipped on Spark 4.2 because no Iceberg spark-runtime is published for 4.2 and connector.catalog.View changed shape — test-only failure per decision-tree step 4.
- Spark 4.2: BloomFilter tests fail under Comet (#4968)
- Area labels:
area:expressions
- Rationale: Comet's own
CometExecRuleSuite / CometExec3_4PlusSuite BloomFilter cases are assume-skipped on Spark 4.2 pending investigation; test-only failure per decision-tree step 4.
- Spark 4.2: UnionCodegenSuite partitioning-aware test inspects UnionExec that Comet replaces (#4965)
- Area labels:
spark sql tests
- Rationale:
SPARK-56482 test reaches into UnionExec via collectFirst, which Comet replaces with CometUnionExec; test-harness assumption failure, no user-facing bug — decision-tree step 4.
- Spark 4.2: segment-tree window metrics unavailable under CometWindowExec (#4964)
- Area labels:
spark sql tests, area:expressions
- Rationale: new Spark 4.2 segment-tree window suites read a metric off a Spark
Window node that Comet replaces; test-only failure per decision-tree step 4.
- Spark 4.2: SQL Last Attempt Metric (SLAM) not propagated through Comet operators (#4963)
- Area labels:
spark sql tests
- Rationale: SLAM is a new Spark 4.2 metrics/telemetry surface; Comet doesn't propagate it through native operators, breaking the new suites but not affecting query results — decision-tree step 4.
Enhancements
- Config naming audit: rename inconsistent spark.comet.* keys before 1.0 (#4978)
- Area labels: none
- Rationale: umbrella tracker proposing six rename clusters via the
.withAlternative(...) deprecation-alias mechanism; a coordinated cleanup, not a defect.
- Optimize
WideDecimalBinaryExpr: skip the null-masking pass in non-ANSI mode when nothing overflows (#4943)
- Area labels:
area:expressions
- Rationale: proposes a
contains(&i128::MAX) fast path so the null-masking pass is skipped when no value overflowed; bit-identical output, purely a perf optimization.
- [EPIC] Optimize native scalar expressions used in TPC-DS (#4936)
- Area labels:
area:expressions
- Rationale: umbrella tracking benchmark-first tuning of hot scalar expressions with existing tests as the correctness gate; new performance work, not a defect.
- feat: Support Apache Uniffle remote shuffle service for Comet native shuffle (#4913)
- Area labels:
area:shuffle
- Rationale: proposes a pluggable partition writer / shuffle block reader abstraction with Uniffle as the first remote-shuffle backend — a new capability, not a defect.
Escalations to consider
- Spark 4.2: ANSI arithmetic overflow tests fail under Comet (#4967)
- Currently
priority:medium on the assumption the divergence surfaces only with spark.sql.ansi.enabled; if investigation shows Comet either throws where Spark doesn't or (worse) silently produces the pre-4.2 result in ANSI mode, that's the guide's escalation trigger for priority:critical (silent wrong result in ANSI). Reviewer decision after root-cause.
Skipped — needs more info
- Codegen dispatch for YearMonth and DayTimeInterval with shuffle and nested element (#4975)
- Reproducer is provided but the "Expected behavior" and "Additional context" sections are empty; without a stated symptom (crash vs. wrong result vs. missed native path) the priority is unclear.
bug label was applied by the reporter — reviewer should confirm and prioritize after root-cause.
- (Informational) AI driven performance analysis of scan implementation across plain Spark, Velox, and Comet (#4842)
- Long informational analysis (still open from previous triage). Not a single bug report or feature request; contains ~9 candidate remediation items that would each become their own issue. Reviewer's call whether to split it, apply a
discussion/documentation label, or close.
- Bug triage results: 2026-07-13 (#4905)
- Prior triage-summary meta-issue awaiting human review and closure; leaving
requires-triage in place so it stays visible.
- Bug triage results: 2026-07-06 (#4838)
- Prior triage-summary meta-issue awaiting human review and closure; leaving
requires-triage in place.
- Bug triage results: 2026-06-29 (#4751)
- Prior triage-summary meta-issue awaiting human review and closure; leaving
requires-triage in place.
Triage pass over the open
requires-triagequeue, per the project Bug Triage Guide.priority:critical3,priority:high1,priority:medium2,priority:low5Labels have already been applied and
requires-triageremoved from each issue listed under "Bugs" and "Enhancements". A reviewer should spot-check the calls and close this issue when satisfied. To correct a label, edit the affected issue directly.Bugs
priority:critical
area:aggregation,spark sql testscollect_set; Comet's native aggregate does not, so results silently diverge from Spark — decision-tree step 1 (silent wrong results).area:expressions" \x01true\x01 "tonullwhere Spark casts totrue— silent divergence in a widely-used cast path (matches the guide's "incorrect cast behavior" critical example).get_json_objectreturns last value for duplicate keys, Spark returns first (#4947)area:expressions{"a":1,"a":2}with path$.aComet returns2while Spark returns1; silent wrong result on a documented Spark semantic — decision-tree step 1.priority:high
java.lang.NegativeArraySizeException(#4944)area:scanAbstractMessageLite.toByteArrayoverflows to a negative size while serializing native Iceberg scan partitions on a large real workload; visible crash rather than wrong result — decision-tree step 2.priority:medium
area:expressionsCometExpressionSuiteANSI-overflow tests areassume-gated off on Spark 4.2 pending investigation of behavior differences; today's workaround is the guard, but the underlying divergence could affect ANSI-mode users — functional gap with workaround per decision-tree step 3.OneRowRelationin Union branches forces Union and downstream aggregates off Comet (TPC-DS q77a) (#4949)area:scanOneRowRelationscan cascades fallback across a wholeUnion+ downstream aggregates on Spark 4.2's q77a plan; correctness is preserved via Spark, but Comet acceleration is lost — decision-tree step 3 (performance regression with workaround).priority:low
area:scanCometIcebergNativeSuite"REST catalog with native Iceberg scan" isassume-skipped on Spark 4.2 because no Iceberg spark-runtime is published for 4.2 andconnector.catalog.Viewchanged shape — test-only failure per decision-tree step 4.area:expressionsCometExecRuleSuite/CometExec3_4PlusSuiteBloomFilter cases areassume-skipped on Spark 4.2 pending investigation; test-only failure per decision-tree step 4.spark sql testsSPARK-56482test reaches intoUnionExecviacollectFirst, which Comet replaces withCometUnionExec; test-harness assumption failure, no user-facing bug — decision-tree step 4.spark sql tests,area:expressionsWindownode that Comet replaces; test-only failure per decision-tree step 4.spark sql testsEnhancements
.withAlternative(...)deprecation-alias mechanism; a coordinated cleanup, not a defect.WideDecimalBinaryExpr: skip the null-masking pass in non-ANSI mode when nothing overflows (#4943)area:expressionscontains(&i128::MAX)fast path so the null-masking pass is skipped when no value overflowed; bit-identical output, purely a perf optimization.area:expressionsarea:shuffleEscalations to consider
priority:mediumon the assumption the divergence surfaces only withspark.sql.ansi.enabled; if investigation shows Comet either throws where Spark doesn't or (worse) silently produces the pre-4.2 result in ANSI mode, that's the guide's escalation trigger forpriority:critical(silent wrong result in ANSI). Reviewer decision after root-cause.Skipped — needs more info
buglabel was applied by the reporter — reviewer should confirm and prioritize after root-cause.discussion/documentationlabel, or close.requires-triagein place so it stays visible.requires-triagein place.requires-triagein place.