Add Integration test to increase the coverage on Read by format #1439

yalimu-g · 2025-11-19T19:17:35Z

Add function testWindowFunctionPartitionBy in ReadByFormatIntegrationTestBase to test Window partitionBy function.

…to cover the test cases of window partition by.

gemini-code-assist · 2025-11-19T19:17:49Z

Summary of Changes

Hello @yalimu-g, I'm Gemini Code Assist¹! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a new integration test to enhance the coverage of the "Read by format" feature in the Spark BigQuery connector. The added test specifically validates the correct functioning of Spark's Window.partitionBy operation when applied to data read from BigQuery, ensuring robust data processing capabilities.

Highlights

New Integration Test: A new integration test, testWindowFunctionPartitionBy, has been added to ReadByFormatIntegrationTestBase to specifically test the Window.partitionBy function.
Increased Test Coverage: This new test increases the overall test coverage for the "Read by format" functionality within the Spark BigQuery connector, particularly for window functions.
Public Dataset Usage: The test utilizes the bigquery-public-data.google_analytics_sample.ga_sessions_20170801 dataset to validate the window function behavior when reading data from BigQuery.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature	Command	Description
Code Review	`/gemini review`	Performs a code review for the current pull request in its current state.
Pull Request Summary	`/gemini summary`	Provides a summary of the current pull request in its current state.
Comment	@gemini-code-assist	Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help	`/gemini help`	Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

gemini-code-assist

Code Review

This pull request adds an integration test for window functions when reading from BigQuery. The test logic is sound, but I've found a few issues. There's a redundant import, the window function uses a non-deterministic ordering which makes the test weak, and more critically, the assertions on the resulting DataFrame's column count and indexing are incorrect, which would cause the test to fail. I've provided suggestions to address these points.

gemini-code-assist · 2025-11-19T19:19:20Z

...c/test/java/com/google/cloud/spark/bigquery/integration/ReadByFormatIntegrationTestBase.java

+    assertThat(dataset.columns().length).isEqualTo(17);
+    assertThat(
+            Arrays.stream(dataset.schema().fields())
+                .filter(field -> field.name().equals("row_num"))
+                .count())
+        .isEqualTo(1);
+    assertThat(dataset.head().get(16)).isEqualTo(1);


The assertions about the DataFrame's column count and indexing are incorrect. The public table bigquery-public-data.google_analytics_sample.ga_sessions_20170801 contains 15 columns. After adding the row_num column, the DataFrame will have 16 columns in total, not 17. Consequently, the assertion assertThat(dataset.head().get(16)) will fail with an ArrayIndexOutOfBoundsException since valid indices are 0-15. The new row_num column will be at index 15.

Suggested change

assertThat(dataset.columns().length).isEqualTo(17);

assertThat(

Arrays.stream(dataset.schema().fields())

.filter(field -> field.name().equals("row_num"))

.count())

.isEqualTo(1);

assertThat(dataset.head().get(16)).isEqualTo(1);

assertThat(dataset.columns().length).isEqualTo(16);

assertThat(

Arrays.stream(dataset.schema().fields())

.filter(field -> field.name().equals("row_num"))

.count())

.isEqualTo(1);

assertThat(dataset.head().get(15)).isEqualTo(1);

gemini-code-assist · 2025-11-19T19:19:20Z

...c/test/java/com/google/cloud/spark/bigquery/integration/ReadByFormatIntegrationTestBase.java

 import org.apache.spark.sql.Row;
+import org.apache.spark.sql.expressions.Window;
+import org.apache.spark.sql.expressions.WindowSpec;
+import org.apache.spark.sql.types.*;


This wildcard import org.apache.spark.sql.types.* appears to be redundant. The file already specifically imports the required classes from this package (such as DataType, DataTypes, StructField, and StructType), and the new test case does not introduce any new dependencies from this package. To improve clarity and avoid unnecessary imports, this line should be removed.

gemini-code-assist · 2025-11-19T19:19:20Z

...c/test/java/com/google/cloud/spark/bigquery/integration/ReadByFormatIntegrationTestBase.java

+        Window.partitionBy(
+                concat(
+                    col("fullVisitorId"), col("visitStartTime"), col("trafficSource.referralPath")))
+            .orderBy(lit("window_ordering"));


Using lit("window_ordering") for ordering within the window function makes the ordering non-deterministic because all rows within a partition will have the same value for the ordering column. This means row_number() will assign numbers arbitrarily to rows within each partition. While the test might currently pass, it is not robust and could be flaky. It also doesn't properly test the ordering capability of the window function. Please consider ordering by a meaningful column to ensure deterministic behavior, for example visitStartTime.

Suggested change

.orderBy(lit("window_ordering"));

.orderBy(col("visitStartTime"));

yalimu-g · 2025-11-19T20:44:39Z

/gcbrun

yalimu-g · 2025-12-02T19:10:16Z

Close this with the new PR #1447 which can reproduce the customer issue with the dataset they used.

yalimu-g added 6 commits November 16, 2025 23:23

Add testWindowFunctionPartitionBy in ReadByFormatIntegrationTestBase …

4531e3c

…to cover the test cases of window partition by.

Merge branch 'GoogleCloudDataproc:master' into issue1435

0b9a73a

Format file

173e0f5

Use dataset with struct type

cb37807

Un-comment the tests

7dff794

Use a different dataset for partition.

8905419

yalimu-g self-assigned this Nov 19, 2025

yalimu-g requested a review from davidrabinowitz November 19, 2025 19:17

gemini-code-assist bot reviewed Nov 19, 2025

View reviewed changes

yalimu-g added 2 commits November 19, 2025 11:32

Remove unused imports

bb86566

Format file

1db870d

davidrabinowitz mentioned this pull request Nov 19, 2025

Using row_number().over(Window) leads to Arrow exception in 4.0 preview (0.43.1-preview) #1435

Open

yalimu-g closed this Dec 2, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Add Integration test to increase the coverage on Read by format #1439

Add Integration test to increase the coverage on Read by format #1439

yalimu-g commented Nov 19, 2025

Uh oh!

gemini-code-assist bot commented Nov 19, 2025

Uh oh!

gemini-code-assist bot left a comment

Uh oh!

gemini-code-assist bot Nov 19, 2025

Uh oh!

gemini-code-assist bot Nov 19, 2025

Uh oh!

gemini-code-assist bot Nov 19, 2025

Uh oh!

yalimu-g commented Nov 19, 2025

Uh oh!

yalimu-g commented Dec 2, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant

	.orderBy(lit("window_ordering"));
	.orderBy(col("visitStartTime"));

Add Integration test to increase the coverage on Read by format #1439

Add Integration test to increase the coverage on Read by format #1439

Conversation

yalimu-g commented Nov 19, 2025

Uh oh!

gemini-code-assist bot commented Nov 19, 2025

Summary of Changes

Highlights

Footnotes

Uh oh!

gemini-code-assist bot left a comment

Choose a reason for hiding this comment

Code Review

Uh oh!

gemini-code-assist bot Nov 19, 2025

Choose a reason for hiding this comment

Uh oh!

gemini-code-assist bot Nov 19, 2025

Choose a reason for hiding this comment

Uh oh!

gemini-code-assist bot Nov 19, 2025

Choose a reason for hiding this comment

Uh oh!

yalimu-g commented Nov 19, 2025

Uh oh!

yalimu-g commented Dec 2, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

1 participant