Skip to content

Bug triage results: 2026-07-20 #4980

Description

@andygrove

Triage pass over the open requires-triage queue, per the project Bug Triage Guide.

  • Date: 2026-07-20
  • Total issues processed: 20 (15 triaged, 5 skipped, 0 failed)
  • Type counts: 11 bugs, 4 enhancements
  • Priority counts applied: priority:critical 3, priority:high 1, priority:medium 2, priority:low 5
  • Guide: docs/source/contributor-guide/bug_triage.md

Labels have already been applied and requires-triage removed from each issue listed under "Bugs" and "Enhancements". A reviewer should spot-check the calls and close this issue when satisfied. To correct a label, edit the affected issue directly.

Bugs

priority:critical

  • Spark 4.2: Comet native collect_set does not normalize NaN / -0.0 (SPARK-57298) (#4966)
    • Area labels: area:aggregation, spark sql tests
    • Rationale: Spark 4.2 normalizes NaN and -0.0 in collect_set; Comet's native aggregate does not, so results silently diverge from Spark — decision-tree step 1 (silent wrong results).
  • cast string to boolean: trim ISO control bytes to match Spark's UTF8String.trimAll (#4959)
    • Area labels: area:expressions
    • Rationale: reporter demonstrates Comet casts " \x01true\x01 " to null where Spark casts to true — silent divergence in a widely-used cast path (matches the guide's "incorrect cast behavior" critical example).
  • Native implementation of get_json_object returns last value for duplicate keys, Spark returns first (#4947)
    • Area labels: area:expressions
    • Rationale: for {"a":1,"a":2} with path $.a Comet returns 2 while Spark returns 1; silent wrong result on a documented Spark semantic — decision-tree step 1.

priority:high

  • Comet throws java.lang.NegativeArraySizeException (#4944)
    • Area labels: area:scan
    • Rationale: protobuf AbstractMessageLite.toByteArray overflows to a negative size while serializing native Iceberg scan partitions on a large real workload; visible crash rather than wrong result — decision-tree step 2.

priority:medium

  • Spark 4.2: ANSI arithmetic overflow tests fail under Comet (#4967)
    • Area labels: area:expressions
    • Rationale: CometExpressionSuite ANSI-overflow tests are assume-gated off on Spark 4.2 pending investigation of behavior differences; today's workaround is the guard, but the underlying divergence could affect ANSI-mode users — functional gap with workaround per decision-tree step 3.
  • Spark 4.2: OneRowRelation in Union branches forces Union and downstream aggregates off Comet (TPC-DS q77a) (#4949)
    • Area labels: area:scan
    • Rationale: unsupported OneRowRelation scan cascades fallback across a whole Union + downstream aggregates on Spark 4.2's q77a plan; correctness is preserved via Spark, but Comet acceleration is lost — decision-tree step 3 (performance regression with workaround).

priority:low

  • Spark 4.2: native Iceberg REST catalog scan test fails under Comet (#4969)
    • Area labels: area:scan
    • Rationale: CometIcebergNativeSuite "REST catalog with native Iceberg scan" is assume-skipped on Spark 4.2 because no Iceberg spark-runtime is published for 4.2 and connector.catalog.View changed shape — test-only failure per decision-tree step 4.
  • Spark 4.2: BloomFilter tests fail under Comet (#4968)
    • Area labels: area:expressions
    • Rationale: Comet's own CometExecRuleSuite / CometExec3_4PlusSuite BloomFilter cases are assume-skipped on Spark 4.2 pending investigation; test-only failure per decision-tree step 4.
  • Spark 4.2: UnionCodegenSuite partitioning-aware test inspects UnionExec that Comet replaces (#4965)
    • Area labels: spark sql tests
    • Rationale: SPARK-56482 test reaches into UnionExec via collectFirst, which Comet replaces with CometUnionExec; test-harness assumption failure, no user-facing bug — decision-tree step 4.
  • Spark 4.2: segment-tree window metrics unavailable under CometWindowExec (#4964)
    • Area labels: spark sql tests, area:expressions
    • Rationale: new Spark 4.2 segment-tree window suites read a metric off a Spark Window node that Comet replaces; test-only failure per decision-tree step 4.
  • Spark 4.2: SQL Last Attempt Metric (SLAM) not propagated through Comet operators (#4963)
    • Area labels: spark sql tests
    • Rationale: SLAM is a new Spark 4.2 metrics/telemetry surface; Comet doesn't propagate it through native operators, breaking the new suites but not affecting query results — decision-tree step 4.

Enhancements

  • Config naming audit: rename inconsistent spark.comet.* keys before 1.0 (#4978)
    • Area labels: none
    • Rationale: umbrella tracker proposing six rename clusters via the .withAlternative(...) deprecation-alias mechanism; a coordinated cleanup, not a defect.
  • Optimize WideDecimalBinaryExpr: skip the null-masking pass in non-ANSI mode when nothing overflows (#4943)
    • Area labels: area:expressions
    • Rationale: proposes a contains(&i128::MAX) fast path so the null-masking pass is skipped when no value overflowed; bit-identical output, purely a perf optimization.
  • [EPIC] Optimize native scalar expressions used in TPC-DS (#4936)
    • Area labels: area:expressions
    • Rationale: umbrella tracking benchmark-first tuning of hot scalar expressions with existing tests as the correctness gate; new performance work, not a defect.
  • feat: Support Apache Uniffle remote shuffle service for Comet native shuffle (#4913)
    • Area labels: area:shuffle
    • Rationale: proposes a pluggable partition writer / shuffle block reader abstraction with Uniffle as the first remote-shuffle backend — a new capability, not a defect.

Escalations to consider

  • Spark 4.2: ANSI arithmetic overflow tests fail under Comet (#4967)
    • Currently priority:medium on the assumption the divergence surfaces only with spark.sql.ansi.enabled; if investigation shows Comet either throws where Spark doesn't or (worse) silently produces the pre-4.2 result in ANSI mode, that's the guide's escalation trigger for priority:critical (silent wrong result in ANSI). Reviewer decision after root-cause.

Skipped — needs more info

  • Codegen dispatch for YearMonth and DayTimeInterval with shuffle and nested element (#4975)
    • Reproducer is provided but the "Expected behavior" and "Additional context" sections are empty; without a stated symptom (crash vs. wrong result vs. missed native path) the priority is unclear. bug label was applied by the reporter — reviewer should confirm and prioritize after root-cause.
  • (Informational) AI driven performance analysis of scan implementation across plain Spark, Velox, and Comet (#4842)
    • Long informational analysis (still open from previous triage). Not a single bug report or feature request; contains ~9 candidate remediation items that would each become their own issue. Reviewer's call whether to split it, apply a discussion/documentation label, or close.
  • Bug triage results: 2026-07-13 (#4905)
    • Prior triage-summary meta-issue awaiting human review and closure; leaving requires-triage in place so it stays visible.
  • Bug triage results: 2026-07-06 (#4838)
    • Prior triage-summary meta-issue awaiting human review and closure; leaving requires-triage in place.
  • Bug triage results: 2026-06-29 (#4751)
    • Prior triage-summary meta-issue awaiting human review and closure; leaving requires-triage in place.

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Fields

    No fields configured for issues without a type.

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions