Repository navigation
feat: add opt-in readable column labels for SQL results - #26022
Open
haochunchang wants to merge 7 commits into
Open
haochunchang wants to merge 7 commits into
haochunchang wants to merge 7 commits into
Conversation
haochunchang
force-pushed
the
feat/25903-pretty-column-names
branch
from
October 8, 2026 12:23
35ef561 to
0e39f1b
Compare
Unaliased SQL output columns are named with the planner's lookup key, such as `t.a + Int64(1)`, which includes type wrappers and table qualifiers. Add `datafusion.sql_parser.pretty_column_names` (default false). When true, each output column of a top-level query gets a readable label, such as `a + 1`, applied as an alias on top of the planned query. Names inside the plan don't change, so name resolution, views, CTAS, INSERT, COPY, DESCRIBE and the DataFrame API keep today's names. A column keeps its current name when its label collides with another column or when its origin can't be traced. Also improve `Expr::human_display`: render scalar functions, add parentheses by operator precedence, fix CUBE printing as ROLLUP, and render aggregate FILTER and ORDER BY clauses in SQL syntax. This changes some physical plan text even when the option is off. Closes apache#25903 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Labeling stacked a new projection on top of the query plan. Every optimizer pass then walked it until merge_consecutive_projections folded it into the query's own projection, which made planning a 200-aggregate query about 38% slower with the option on. When the top node is already a projection, alias its expressions in place, and add a projection only for other top nodes, such as Sort or Limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Renaming result columns broke code that reads them by name and needed collision rules for duplicate labels. Labels now go in the field metadata keys `datafusion.label` and `datafusion.label_of`, and column names don't change. `DataFrame::show` and datafusion-cli table output display the label. `label_of` keeps a stale label from showing after a column is renamed. Rename the option to `datafusion.sql_parser.column_labels`. Replace the EXPLAIN-based sqllogictest file with Rust tests, because sqllogictest can't show field metadata. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`batches_with_display_names` rebuilt every batch with `RecordBatch::try_new`, which rejects a batch without columns. `DataFrame::show`, `to_string`, `show_limit` and datafusion-cli table output call it even with the option off, so a query such as `SELECT FROM t` failed to print. Return the batch unchanged when no column has a label, which also skips a needless rebuild, and keep the row count when relabeling. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Aggregate labels came from `UdafHumanDisplayBuilder`. Rendering its FILTER and ORDER BY clauses in SQL syntax for labels also changed physical EXPLAIN text with the option off and hid literal types such as `Float32(1)`. It also broke `reverse_expr`: it still looks for `ORDER BY [`, so a reversed `first_value` showed as `last_value` with the old sort direction. Restore the builder and render aggregate labels in the labeling step, as window labels already are. Ordered-set aggregates now get labels such as `percentile_cont(0.5) WITHIN GROUP (ORDER BY c)`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`remove_unnecessary_projections` gave up on any projection whose output metadata differs from what its expressions derive. With `column_labels` on, every labeled query has such a projection on top, so struct field access stopped being pushed into the scan and filters and joins lost their embedded projections. Split such a projection instead: compute the expressions in a projection with derived metadata, push it down as usual, and re-apply the metadata in a column-only projection on top. Keep the plan as is when the projection only passes columns through, or when the computation can't move. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Some labels read as a different or ambiguous expression: `NOT (a > 1 AND b > 2)` showed as `NOT a > 1 AND b > 2`, `-(a + b)` as `(- a + b)`, `a IN (1, 2)` as `a IN 1, 2`, and string literals had no quotes, so `SELECT 1, '1'` gave two `1` headers. LIKE and SIMILAR TO also printed their escape character as `CHAR '\'`. Group binary operands of NOT and negation, wrap IN lists in parentheses, quote string literals, and print `ESCAPE`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
haochunchang
force-pushed
the
feat/25903-pretty-column-names
branch
from
October 10, 2026 13:02
fbb8554 to
eb93ba8
Compare
haochunchang
marked this pull request as ready for review
October 10, 2026 13:03
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Which issue does this PR close?
Rationale for this change
An unaliased column in a SQL result is named with the planner's internal lookup key, so headers include type wrappers, table qualifiers, and default window frames:
t.a + 1t.a + Int64(1)a + 1(a + b) * ct.a + t.b * t.c(a + b) * cavg(c) FILTER (WHERE c > 5)avg(t.c) FILTER (WHERE t.c > Int64(5))avg(c) FILTER (WHERE c > 5)row_number() OVER (ORDER BY a)row_number() ORDER BY [t.a ASC NULLS LAST] RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROWrow_number() OVER (ORDER BY a)The name is also the key that the planner uses to find columns, so it can't simply be shortened (#2027, #5174, #10274). This PR implements the metadata direction proposed in this comment on #25903: column names don't change, and each output column of a top-level query gets a readable label in its Arrow field metadata.
DataFrame::show()anddatafusion-clitable output display the label as the header.Because no name changes, code that reads result columns by name keeps working, and two columns can share a label.
What changes are included in this PR?
datafusion.sql_parser.column_labels, defaultfalse. When it'sfalse, planning reads one config value and does nothing else.datafusion_common::metadata:datafusion.label(COLUMN_LABEL_KEY): the label, such asa + 1.datafusion.label_of(COLUMN_LABEL_OF_KEY): the column name that the label applies to. Field metadata stays on a field after an alias, so afterwith_column_renamed("sum(t.b)", "total")the field still carriessum(b). A label is shown only while the field name equalslabel_of.display_name,schema_with_display_names, andbatches_with_display_namesreturn the name to show for a field. The batch helper shares the arrays without copying them, returns unlabeled batches unchanged, and keeps the row count of zero-column batches.datafusion/sql/src/column_labels.rs), called from theStatement::Querybranch ofsql_statement_to_plan:Projection,Aggregate,Window,SubqueryAlias,Filter,Sort,Limit,Distinct::All,Union, andONjoins), and renders it withExpr::human_display.name(args)with theirDISTINCT, null treatment,FILTER, andORDER BYparts. Ordered-set aggregates useWITHIN GROUP (ORDER BY ...), such aspercentile_cont(0.5) WITHIN GROUP (ORDER BY c).name(args) OVER (...)and leaves out default parts: the default frame (same tie check as the planner), constant sort keys,ASC, and the defaultNULLSordering.Projectionto its own name with the label metadata. It adds a projection only when the top node is something else, such asSortorLimit.SELECT;USING/NATURALjoins,UNNEST,VALUES, recursive CTEs, subquery expressions,DISTINCT ON);DataFrame::to_string(),show(), andshow_limit(), and thedatafusion-clitable format, show labels.datafusion-cliCSV, TSV, JSON, and NDJSON output keep column names, because programs parse them and JSON keys must be unique.Expr::human_displayfixes (inSqlDisplay): renders scalar functions (coalesce(NULL, 1)), adds parentheses by operator precedence ((a + b) * c,a - (b - c)), printsCUBEinstead ofROLLUPforGroupingSet::Cube, keeps binary operands ofNOTand negation grouped (NOT (a > 1 AND b > 2),(- (a + b))), wrapsINlists in parentheses, quotes string literals ('foo','it''s'), and printsESCAPEinstead ofCHARforLIKEandSIMILAR TO.datafusion/physical-plan/src/projection.rs):remove_unnecessary_projectionsused to stop at any projection whose output metadata differs from what its expressions derive. Every labeled query has one on top, so with the option on, struct field access wasn't pushed into the scan and filters and joins lost their embedded projections. Such a projection is now split: the expressions are computed in a projection with derived metadata, which is pushed down as usual, and a column-only projection on top re-applies the metadata. The plan is kept as is when the projection only passes columns through or when the computation can't move. This applies to any projection that sets field metadata, not only labels.Sort::human_displayandSortListDisplayrender sort keys in SQL syntax for labels.UdafHumanDisplayBuilderis unchanged, so physicalEXPLAINtext for aggregates stays the same.output-field-name-semantic.md.Not affected
CREATE VIEW,CREATE TABLE AS,INSERT ... SELECT,COPY,DESCRIBE, and the DataFrame API add no labels. Column names never change, so references such asORDER BY "t.a + Int64(1)"and client lookups such asdf["sum(t.a)"]keep working.Decisions where the proposal was silent
count(*) OVER (...): the planner plans it ascount(1)and aliases it back tocount(*) .... The label keepscount(*).What is the testing strategy for this PR?
datafusion/core/tests/sql/column_labels.rslistsname => labelfor each output column of one query per row in the proposal's tables, and runs each query. It also covers:WITHIN GROUP);ORDER BYby name and by position,HAVING,DISTINCT,DISTINCT ON,UNNEST,USINGjoins, set operations, recursive CTEs, subqueries, and windows ordered by a unique key;collect(), withORDER BY ... LIMITstill planned asTopK;DataFrame::to_string()headers,select_columnsby name, andwith_column_renameddropping the stale label.datafusion-cli/src/print_format.rs:print_column_labelschecks table output (with and without rows) and CSV output.test_print_batches_zero_column_batch_with_rowschecks that a zero-column batch prints in table format.datafusion_common::metadataunit tests fordisplay_nameandbatches_with_display_names, including zero-column batches.information_schema.slt. With the option forced on locally,projection_pushdown.sltandparquet_nested_schema_pruning.sltplans match the option-off plans apart from one column-onlyProjectionExecon top.datafusion/exprunit tests and doctests for precedence, scalar functions,CUBE,NOTand negation grouping,INlists, string literals,ESCAPE, andSort::human_display.datafusion/physical-plan/src/projection.rsunit tests: a projection that sets metadata is split and pushed below a filter; an identity projection that sets metadata stays; a projection whose input can't take it stays.sqllogictest doesn't show field metadata, so the label cases are Rust tests instead of
.sltfiles.Benchmarks
datafusion/core/benches/sql_planner.rsat the current head, option off vs on, same build, run back to back on an M-series laptop. The option was set throughDATAFUSION_SQL_PARSER_COLUMN_LABELS, read by a local, uncommitted change to the bench's session setup (SessionConfig::from_env()). Clickbench cases were left out locally.logical_select_all_from_1000physical_select_all_from_1000physical_select_aggregates_from_200logical_wide_aggregate_100_exprsphysical_plan_tpch_allphysical_plan_tpcds_allSELECT *gets no labels, so it plans the same with the option on or off. The smallselect_alland TPC-H/TPC-DS changes, in both directions, are run-to-run variance on this laptop.The cost on wide aggregates comes from the labeling aliases: about 10 µs per column for physical planning and 2 µs per column for logical planning. With the option off,
optimize_projectionsremoves the projection above theAggregatebecause it only passes columns through. With aliases that carry metadata, the projection has to stay, through the optimizer and as aProjectionExecthat reuses the input arrays.Are there any user-facing changes?
datafusion.sql_parser.column_labels, defaultfalse. With the default, nothing changes.MemTable::try_new(schema, batches), can fail becauseSchema::containscompares field metadata. Usingdf.schema()avoids this.ProjectionExecthat carries the label metadata.Expr::human_displayoutput changes forNOT, negation,INlists, string literals, andESCAPE, as listed above.Sort::human_displayandSortListDisplayindatafusion_expr::expr;COLUMN_LABEL_KEY,COLUMN_LABEL_OF_KEY,column_label_metadata,display_name,schema_with_display_names, andbatches_with_display_namesindatafusion_common::metadata.Follow-up: turn the option on in
datafusion-cli, where #2027 first reported unreadable headers.🤖 Generated with Claude Code