Skip to content

[pull] main from apache:main - #281

Merged
pull[bot] merged 2 commits into
buraksenn:mainfrom
apache:main
Jun 16, 2026
Merged

pull[bot] merged 2 commits into
buraksenn:mainfrom
apache:main

Conversation

@pull

@pull pull Bot commented Jun 16, 2026 •

Copy link
Copy Markdown

See Commits and Changes for more details.


Created by pull[bot] (v2.0.0-alpha.4)

Can you help keep this open source service alive? 💖 Please sponsor : )

2010YOUY01 and others added 2 commits June 15, 2026 21:00
…ream` (#22953)

## Which issue does this PR close?

<!--
We generally require a GitHub issue to be filed for all bug fixes and
enhancements and this helps us generate change logs for our releases.
You can link an issue to this PR using the GitHub syntax. For example
`Closes #123` indicates that this PR will close issue #123.
-->

Part of #22710

## Rationale for this change

<!--
Why are you proposing this change? If this is already explained clearly
in the issue then this section is not needed.
Explaining clearly why changes are proposed helps reviewers understand
your changes and offer better suggestions for fixes.
-->
The goal is after we have fully migrated from the old `row_hash.rs`, the
existing UTs should be kept. Specifically, all tests that include
`GroupedHashAggregateStream`

There are 3 previous PRs for the migration have been merged, some
existing UTs are applicable to them, this PR migrated those tests to the
new implementation.

The test migration includes:
1. copy and paste test case
2. Change `GroupedHashAggregateStream` to `PartialHashAggregateStream`
(or other stream in new impl)
3. Left a comment on the migrated test case, so in the final delete move
it's more clear which tests have already been moved.

This PR moved 2 applicable UTs, and updated the comments for all the
tests moved previously.

(Just some random thoughts, in general I don't think it's a good idea to
write tests against low-level utilities like
`GroupedHashAggregateStream`, all tests should better be at SQL level,
or at least at `ExecutionPlan` level, so their test goal are more likely
to survive refactors)

## What changes are included in this PR?

<!--
There is no need to duplicate the description in the issue here but it
is sometimes worth providing a summary of the individual changes in this
PR.
-->

## Are these changes tested?

<!--
We typically require tests for all PRs in order to:
1. Prevent the code from being accidentally broken by subsequent changes
5. Serve as another way to document the expected behavior of the code

If tests are not included in your PR, please explain why (for example,
are they covered by existing tests)?
-->

## Are there any user-facing changes?

<!--
If there are user-facing changes then we may require documentation to be
updated before approving the PR.
-->

<!--
If there are any breaking changes to public APIs, please add the `api
change` label.
-->
…2816)

## Which issue does this PR close?

- Closes #22775.

## Rationale for this change

the `opt_filter` on `GroupsAccumulator::merge_batch` is a dead
parameter. Aggregate `FILTER` clauses only apply to raw input rows in
the update phase (`update_batch`). `merge_batch` combines already
pre-aggregated states, so there is no per-row filtering to do —
`opt_filter` is meaningless there.

The code confirms this:
- The only production caller (`row_hash.rs`) always passed `None`.
- Existing implementations already ignored it — e.g. `correlation.rs`
asserted `opt_filter.is_none()`, and Spark `avg` used `_opt_filter`.

## What changes are included in this PR?

- Removed `opt_filter` from `merge_batch` in the trait and all
implementations (built-in aggregates, `physical-expr-common`,
`functions-aggregate-common`, Spark, and FFI).
- Updated the trait docs to say `merge_batch` has no `opt_filter`
because filtering happens in the update phase.
- Changed the group zero-init path in `row_hash.rs` to always use
`update_batch` with an all-false filter instead of branching to
`merge_batch`. `update_batch` always takes raw argument types (what
`aggregate_arguments` provides), and since every row is filtered out the
data never matters — this is simpler and more correct.
- Updated all call sites and tests.

## Are these changes tested?

Yes. Existing aggregate tests cover this and were updated to the new
signature. The `first_last` tests were adjusted (with comments) to match
the merge behavior without a filter, and the FFI and Spark tests were
updated too.

## Are there any user-facing changes?

Yes — this is a breaking change to the public `GroupsAccumulator` trait:
`opt_filter` is removed from `merge_batch`. Custom implementations and
direct callers must update their signatures.
@pull pull Bot locked and limited conversation to collaborators Jun 16, 2026
@pull pull Bot added the ⤵️ pull label Jun 16, 2026
@pull
pull Bot merged commit c14379b into buraksenn:main Jun 16, 2026
20 of 21 checks passed
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants