Skip to content

feat: add opt-in readable column labels for SQL results - #26022

Open
haochunchang wants to merge 7 commits into
apache:mainfrom
haochunchang:feat/25903-pretty-column-names
Open

haochunchang wants to merge 7 commits into
apache:mainfrom
haochunchang:feat/25903-pretty-column-names

Conversation

@haochunchang

@haochunchang haochunchang commented Oct 4, 2026 •

Copy link
Copy Markdown

Which issue does this PR close?

Rationale for this change

An unaliased column in a SQL result is named with the planner's internal lookup key, so headers include type wrappers, table qualifiers, and default window frames:

Query item Column name Label with this PR (option on)
t.a + 1 t.a + Int64(1) a + 1
(a + b) * c t.a + t.b * t.c (a + b) * c
avg(c) FILTER (WHERE c > 5) avg(t.c) FILTER (WHERE t.c > Int64(5)) avg(c) FILTER (WHERE c > 5)
row_number() OVER (ORDER BY a) row_number() ORDER BY [t.a ASC NULLS LAST] RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW row_number() OVER (ORDER BY a)

The name is also the key that the planner uses to find columns, so it can't simply be shortened (#2027, #5174, #10274). This PR implements the metadata direction proposed in this comment on #25903: column names don't change, and each output column of a top-level query gets a readable label in its Arrow field metadata. DataFrame::show() and datafusion-cli table output display the label as the header.

Because no name changes, code that reads result columns by name keeps working, and two columns can share a label.

What changes are included in this PR?

  • New option datafusion.sql_parser.column_labels, default false. When it's false, planning reads one config value and does nothing else.
  • Metadata keys in datafusion_common::metadata:
    • datafusion.label (COLUMN_LABEL_KEY): the label, such as a + 1.
    • datafusion.label_of (COLUMN_LABEL_OF_KEY): the column name that the label applies to. Field metadata stays on a field after an alias, so after with_column_renamed("sum(t.b)", "total") the field still carries sum(b). A label is shown only while the field name equals label_of.
    • display_name, schema_with_display_names, and batches_with_display_names return the name to show for a field. The batch helper shares the arrays without copying them, returns unlabeled batches unchanged, and keeps the row count of zero-column batches.
  • Labeling step (datafusion/sql/src/column_labels.rs), called from the Statement::Query branch of sql_statement_to_plan:
    • Traces each output column, by field index, down to the expression that produced it (through Projection, Aggregate, Window, SubqueryAlias, Filter, Sort, Limit, Distinct::All, Union, and ON joins), and renders it with Expr::human_display.
    • Renders aggregate functions as name(args) with their DISTINCT, null treatment, FILTER, and ORDER BY parts. Ordered-set aggregates use WITHIN GROUP (ORDER BY ...), such as percentile_cont(0.5) WITHIN GROUP (ORDER BY c).
    • Renders window functions as name(args) OVER (...) and leaves out default parts: the default frame (same tie check as the planner), constant sort keys, ASC, and the default NULLS ordering.
    • Aliases each labeled expression of the top Projection to its own name with the label metadata. It adds a projection only when the top node is something else, such as Sort or Limit.
  • A column gets no label when:
    • it has an explicit alias;
    • it's a column reference typed in the outermost SELECT;
    • it can't be traced (USING/NATURAL joins, UNNEST, VALUES, recursive CTEs, subquery expressions, DISTINCT ON);
    • its label equals its name.
  • Display: DataFrame::to_string(), show(), and show_limit(), and the datafusion-cli table format, show labels. datafusion-cli CSV, TSV, JSON, and NDJSON output keep column names, because programs parse them and JSON keys must be unique.
  • Expr::human_display fixes (in SqlDisplay): renders scalar functions (coalesce(NULL, 1)), adds parentheses by operator precedence ((a + b) * c, a - (b - c)), prints CUBE instead of ROLLUP for GroupingSet::Cube, keeps binary operands of NOT and negation grouped (NOT (a > 1 AND b > 2), (- (a + b))), wraps IN lists in parentheses, quotes string literals ('foo', 'it''s'), and prints ESCAPE instead of CHAR for LIKE and SIMILAR TO.
  • Projection pushdown for projections that set metadata (datafusion/physical-plan/src/projection.rs): remove_unnecessary_projections used to stop at any projection whose output metadata differs from what its expressions derive. Every labeled query has one on top, so with the option on, struct field access wasn't pushed into the scan and filters and joins lost their embedded projections. Such a projection is now split: the expressions are computed in a projection with derived metadata, which is pushed down as usual, and a column-only projection on top re-applies the metadata. The plan is kept as is when the projection only passes columns through or when the computation can't move. This applies to any projection that sets field metadata, not only labels.
  • New Sort::human_display and SortListDisplay render sort keys in SQL syntax for labels. UdafHumanDisplayBuilder is unchanged, so physical EXPLAIN text for aggregates stays the same.
  • Docs: config reference, plus a "Column labels" section in output-field-name-semantic.md.

Not affected

CREATE VIEW, CREATE TABLE AS, INSERT ... SELECT, COPY, DESCRIBE, and the DataFrame API add no labels. Column names never change, so references such as ORDER BY "t.a + Int64(1)" and client lookups such as df["sum(t.a)"] keep working.

Decisions where the proposal was silent

  • count(*) OVER (...): the planner plans it as count(1) and aliases it back to count(*) .... The label keeps count(*).
  • Window frame tie check for a window over an aggregate uses the aggregate's input schema, because that's the schema the planner used to choose the default frame.

What is the testing strategy for this PR?

  • datafusion/core/tests/sql/column_labels.rs lists name => label for each output column of one query per row in the proposal's tables, and runs each query. It also covers:
    • the option off;
    • ordered-set aggregates (WITHIN GROUP);
    • ORDER BY by name and by position, HAVING, DISTINCT, DISTINCT ON, UNNEST, USING joins, set operations, recursive CTEs, subqueries, and windows ordered by a unique key;
    • duplicate labels;
    • no labels on CTAS and views;
    • labels in every batch from collect(), with ORDER BY ... LIMIT still planned as TopK;
    • DataFrame::to_string() headers, select_columns by name, and with_column_renamed dropping the stale label.
  • datafusion-cli/src/print_format.rs: print_column_labels checks table output (with and without rows) and CSV output. test_print_batches_zero_column_batch_with_rows checks that a zero-column batch prints in table format.
  • datafusion_common::metadata unit tests for display_name and batches_with_display_names, including zero-column batches.
  • The sqllogictest suite passes with no expectation changes, apart from the new option in information_schema.slt. With the option forced on locally, projection_pushdown.slt and parquet_nested_schema_pruning.slt plans match the option-off plans apart from one column-only ProjectionExec on top.
  • datafusion/expr unit tests and doctests for precedence, scalar functions, CUBE, NOT and negation grouping, IN lists, string literals, ESCAPE, and Sort::human_display.
  • datafusion/physical-plan/src/projection.rs unit tests: a projection that sets metadata is split and pushed below a filter; an identity projection that sets metadata stays; a projection whose input can't take it stays.

sqllogictest doesn't show field metadata, so the label cases are Rust tests instead of .slt files.

Benchmarks

datafusion/core/benches/sql_planner.rs at the current head, option off vs on, same build, run back to back on an M-series laptop. The option was set through DATAFUSION_SQL_PARSER_COLUMN_LABELS, read by a local, uncommitted change to the bench's session setup (SessionConfig::from_env()). Clickbench cases were left out locally.

Benchmark Off On Change
logical_select_all_from_1000 8.31 ms 8.57 ms +3.1%
physical_select_all_from_1000 21.03 ms 21.47 ms +2.1%
physical_select_aggregates_from_200 6.94 ms 8.84 ms +27%
logical_wide_aggregate_100_exprs 1.31 ms 1.50 ms +22%
physical_plan_tpch_all 28.03 ms 26.67 ms -4.9%
physical_plan_tpcds_all 460.32 ms 451.23 ms -2.0%

SELECT * gets no labels, so it plans the same with the option on or off. The small select_all and TPC-H/TPC-DS changes, in both directions, are run-to-run variance on this laptop.

The cost on wide aggregates comes from the labeling aliases: about 10 µs per column for physical planning and 2 µs per column for logical planning. With the option off, optimize_projections removes the projection above the Aggregate because it only passes columns through. With aliases that carry metadata, the projection has to stay, through the optimizer and as a ProjectionExec that reuses the input arrays.

Are there any user-facing changes?

  • New config option datafusion.sql_parser.column_labels, default false. With the default, nothing changes.
  • With the option on, result fields carry label metadata. Code that compares a result batch with a schema it built itself, such as MemTable::try_new(schema, batches), can fail because Schema::contains compares field metadata. Using df.schema() avoids this.
  • With the option on, a labeled query's physical plan ends in a column-only ProjectionExec that carries the label metadata.
  • Expr::human_display output changes for NOT, negation, IN lists, string literals, and ESCAPE, as listed above.
  • New public API (additive): Sort::human_display and SortListDisplay in datafusion_expr::expr; COLUMN_LABEL_KEY, COLUMN_LABEL_OF_KEY, column_label_metadata, display_name, schema_with_display_names, and batches_with_display_names in datafusion_common::metadata.

Follow-up: turn the option on in datafusion-cli, where #2027 first reported unreadable headers.

🤖 Generated with Claude Code

@github-actions github-actions Bot added documentation Improvements or additions to documentation sql SQL Planner logical-expr Logical plan and expressions core Core DataFusion crate sqllogictest SQL Logic Tests (.slt) common Related to common crate labels Oct 4, 2026
@haochunchang haochunchang changed the title feat: add opt-in readable column names for SQL results feat: add opt-in readable column labels for SQL results Oct 4, 2026
@haochunchang
haochunchang force-pushed the feat/25903-pretty-column-names branch from 35ef561 to 0e39f1b Compare October 8, 2026 12:23
haochunchang and others added 7 commits October 10, 2026 20:01
Unaliased SQL output columns are named with the planner's lookup key,
such as `t.a + Int64(1)`, which includes type wrappers and table
qualifiers. Add `datafusion.sql_parser.pretty_column_names` (default
false). When true, each output column of a top-level query gets a
readable label, such as `a + 1`, applied as an alias on top of the
planned query. Names inside the plan don't change, so name resolution,
views, CTAS, INSERT, COPY, DESCRIBE and the DataFrame API keep today's
names. A column keeps its current name when its label collides with
another column or when its origin can't be traced.

Also improve `Expr::human_display`: render scalar functions, add
parentheses by operator precedence, fix CUBE printing as ROLLUP, and
render aggregate FILTER and ORDER BY clauses in SQL syntax. This
changes some physical plan text even when the option is off.

Closes apache#25903

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Labeling stacked a new projection on top of the query plan. Every
optimizer pass then walked it until merge_consecutive_projections
folded it into the query's own projection, which made planning a
200-aggregate query about 38% slower with the option on. When the top
node is already a projection, alias its expressions in place, and add
a projection only for other top nodes, such as Sort or Limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Renaming result columns broke code that reads them by name and needed
collision rules for duplicate labels. Labels now go in the field
metadata keys `datafusion.label` and `datafusion.label_of`, and column
names don't change. `DataFrame::show` and datafusion-cli table output
display the label. `label_of` keeps a stale label from showing after
a column is renamed.

Rename the option to `datafusion.sql_parser.column_labels`. Replace
the EXPLAIN-based sqllogictest file with Rust tests, because
sqllogictest can't show field metadata.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`batches_with_display_names` rebuilt every batch with
`RecordBatch::try_new`, which rejects a batch without columns.
`DataFrame::show`, `to_string`, `show_limit` and datafusion-cli table
output call it even with the option off, so a query such as
`SELECT FROM t` failed to print. Return the batch unchanged when no
column has a label, which also skips a needless rebuild, and keep the
row count when relabeling.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Aggregate labels came from `UdafHumanDisplayBuilder`. Rendering its
FILTER and ORDER BY clauses in SQL syntax for labels also changed
physical EXPLAIN text with the option off and hid literal types such
as `Float32(1)`. It also broke `reverse_expr`: it still looks for
`ORDER BY [`, so a reversed `first_value` showed as `last_value` with
the old sort direction.

Restore the builder and render aggregate labels in the labeling step,
as window labels already are. Ordered-set aggregates now get labels
such as `percentile_cont(0.5) WITHIN GROUP (ORDER BY c)`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`remove_unnecessary_projections` gave up on any projection whose output
metadata differs from what its expressions derive. With
`column_labels` on, every labeled query has such a projection on top,
so struct field access stopped being pushed into the scan and filters
and joins lost their embedded projections.

Split such a projection instead: compute the expressions in a
projection with derived metadata, push it down as usual, and re-apply
the metadata in a column-only projection on top. Keep the plan as is
when the projection only passes columns through, or when the
computation can't move.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some labels read as a different or ambiguous expression:
`NOT (a > 1 AND b > 2)` showed as `NOT a > 1 AND b > 2`, `-(a + b)`
as `(- a + b)`, `a IN (1, 2)` as `a IN 1, 2`, and string literals had
no quotes, so `SELECT 1, '1'` gave two `1` headers. LIKE and SIMILAR TO
also printed their escape character as `CHAR '\'`.

Group binary operands of NOT and negation, wrap IN lists in
parentheses, quote string literals, and print `ESCAPE`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@haochunchang
haochunchang force-pushed the feat/25903-pretty-column-names branch from fbb8554 to eb93ba8 Compare October 10, 2026 13:02
@github-actions github-actions Bot added the physical-plan Changes to the physical-plan crate label Oct 10, 2026
@haochunchang
haochunchang marked this pull request as ready for review October 10, 2026 13:03
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

common Related to common crate core Core DataFusion crate documentation Improvements or additions to documentation logical-expr Logical plan and expressions physical-plan Changes to the physical-plan crate sql SQL Planner sqllogictest SQL Logic Tests (.slt)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Proposal] Readable default column names for SQL query results

1 participant