Skip to content

Navigation Menu

Sign in
Sign up

Slow count on parquet #1013

Unanswered
ashangit asked this question in Q&A
Discussion options

Hi,

I'm running a job doing union with count on a parquet tables.
Running this count with blaze takes 14 min while it takes 21s without blaze

Here is the config for blaze:

 spark.blaze.enable: "true"
 spark.sql.extensions: org.apache.iceberg.spark.extensions.IcebergSparkSessionExtensions,org.apache.spark.sql.blaze.BlazeSparkSessionExtension
 spark.shuffle.manager: org.apache.spark.sql.execution.blaze.shuffle.BlazeShuffleManager
 spark.memory.offHeap.enabled: "false"
 spark.executor.memoryOverhead: 27g

From the SQL plan I can see that we spend most of the time in Native.io_time total.
Also the Input bytes for this stage is around 130GB with blaze and only 25MB without balze.

Here the SQL plan with blaze
Screenshot 2025年06月06日 at 12 22 30

And without
Screenshot 2025年06月06日 at 12 25 59

It looks like we are scanning the whole parquet file to count the number of rows and not relying on metadata.
Am I understanding things correctly or am I missing some config on the job? Also it looks quite slow to get >100MB in >40s (from spark jobs the same kind of action getting >100MB for parquet file is around 20s)

You must be logged in to vote

Replies: 0 comments

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
1 participant

AltStyle によって変換されたページ (->オリジナル) /