Skip to content

Gold pipelines seem not to be self consistent when evaluated against shipped gold DBs #5

Description

@Tejas2507

I found several DAComp-DE tasks where the shipped gold package does not appear to reproduce the shipped gold DuckDB database when re-run under the official evaluator flow.

According to the DAComp-DE evaluation README, CFS runs pred_dir/<example_id>/run.py and compares the generated DuckDB database against the gold database, while CS runs SQL in the gold environment for table-level checks.

If the repository's own gold package is used as the prediction, I would expect it to reproduce the shipped gold DB or score very close to 100, unless nondeterminism is explicitly documented.

However, several gold packages appear not to be self-consistent.

Observed aggregate self-test results:

  • Impl gold CFS average: ~58.45
  • Impl gold CS average: ~74.31
  • Evol gold CFS average: ~39.73
  • Evol gold CS average: ~59.33

Example impl cases

Task Failure type Observed behavior Possible Reason
impl-007, impl-010, impl-015 DuckDB attach conflict in gold run.py Fresh gold rebuild produces a 268 KB shell DB; 0/76 gold tables are present. Gold run.py finds raw tables but raw import fails repeatedly with Cannot attach "source_db" because the DB is already attached as app_reporting_start.
dacomp-de-impl-028 Raw schema discovery bug Fresh gold rebuild produces a 268 KB shell DB; 0/16 gold tables are present. Gold run.py logs Found 0 raw tables, then staging fails on missing raw.channel_basic_a_2, raw.video, etc. The source tables exist under the raw schema, so this appears to be schema-unaware table discovery.
dacomp-de-impl-022 CS evaluator cannot find nested SQL paths Full pipeline mostly works: 44/44 tables present and 41/44 row counts match. But CS score is only ~14. Gold SQL files live under nested paths such as sql/intermediate/customer/ and sql/intermediate/calendar/. run.py can execute them, but CS appears to expect flat paths like sql/intermediate/{table}.sql, causing false CS failures.
dacomp-de-impl-002 Runtime date drift from current_date Rebuilt DB is larger than reference; int_pendo__calendar_spine grows from 1,747 to 1,981 rows, and daily marts grow by ~13-21%. Gold SQL uses current_date + interval '1 week' for the calendar spine. The frozen gold max date is 2025-10-18, while a fresh rebuild extends to 2026-06-09, so gold-as-pred no longer reproduces the shipped reference DB.

Example Evol cases

Task Observed behavior Possible Reason
dacomp-de-evol-004 Rebuilt output drifts from frozen gold; daily spines extend to the current run date and mart naming differs. The output depends on runtime date, so the same gold package does not reproduce the frozen reference DB.
dacomp-de-evol-005 marts.customer360__trend_analysis grows from about 90k rows to about 6.7M rows on rebuild. The rebuild is affected by time-window drift and join amplification, so gold-as-pred no longer matches frozen gold.
dacomp-de-evol-039 CFS is ~99 but CS is ~9. The full pipeline works, but CS appears unable to locate nested SQL files under paths like sql/intermediate/billing/..., even though run.py can execute them.
dacomp-de-evol-048 Rebuilt DB is an empty 268 KB shell while CS is 100. Gold run.py imports 0 raw tables because raw tables live in the raw schema; the shipped gold DB is populated, but a fresh rebuild is empty.

Why this matters

These cases suggest that some evaluation failures may be caused by gold package reproducibility issues or evaluator path assumptions, not only by candidate pipeline quality.

In particular:

  • CFS can score 0 when the official gold run.py itself rebuilds an empty database.
  • CS can score very low when SQL files are nested in folders that run.py supports but the CS evaluator does not appear to discover.
  • Runtime date usage can cause frozen gold DBs and freshly rebuilt DBs to diverge over time.

Could you clarify whether these behaviors are expected, and whether there are any known limitations or assumptions in the evaluation framework regarding reproducibility, nested SQL file discovery, schema discovery, or runtime-date-dependent pipelines? In particular, should the shipped gold packages be expected to reproduce the shipped gold databases when evaluated under the official workflow?

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions