perf: sort in memory without concat but instead use modified arrow sort [WIP] - #25108
perf: sort in memory without concat but instead use modified arrow sort [WIP]#25108rluvaton wants to merge 3 commits into
Conversation
|
run benchmarks |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (bbc14e7) to a5c809f (merge-base) diff Run configurationrun benchmark clickbench_partitionedResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (bbc14e7) to a5c809f (merge-base) diff Run configurationrun benchmark tpcdsResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (bbc14e7) to a5c809f (merge-base) diff Run configurationrun benchmark tpchResults will be posted here when complete File an issue against this benchmark runner |
|
Benchmark for this request failed before finishing (Kubernetes reason: Benchmarks requested: Runner log (last 40 lines)Kubernetes messageFile an issue against this benchmark runner |
|
Benchmark for this request failed before finishing (Kubernetes reason: Benchmarks requested: Runner log (last 40 lines)Kubernetes messageFile an issue against this benchmark runner |
|
Benchmark for this request failed before finishing (Kubernetes reason: Benchmarks requested: Runner log (last 40 lines)Kubernetes messageFile an issue against this benchmark runner |
|
run benchmarks |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark clickbench_partitionedResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpcdsResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpchResults will be posted here when complete File an issue against this benchmark runner |
|
run benchmarks tpch tpch10 tpcds |
|
run benchmark sort_tpch10 |
|
run benchmark sort_tpch |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpch
env:
prefer_hash_join: "false"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpch10
env:
prefer_hash_join: "false"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpcds
env:
prefer_hash_join: "false"Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpchCPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpcdsCPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark sort_tpchResults will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark running (GKE) | trigger CPU Details (lscpu)Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark sort_tpch10Results will be posted here when complete File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark clickbench_partitionedCPU Details (lscpu)Details
Resource Usageclickbench_partitioned — base (merge-base)
clickbench_partitioned — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpch
env:
prefer_hash_join: "false"CPU Details (lscpu)Details
Resource Usagetpch — base (merge-base)
tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpch10
env:
prefer_hash_join: "false"CPU Details (lscpu)Details
Resource Usagetpch10 — base (merge-base)
tpch10 — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark tpcds
env:
prefer_hash_join: "false"CPU Details (lscpu)Details
Resource Usagetpcds — base (merge-base)
tpcds — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark sort_tpchCPU Details (lscpu)Details
Resource Usagesort_tpch — base (merge-base)
sort_tpch — branch
File an issue against this benchmark runner |
|
🤖 Benchmark completed (GKE) | trigger Instance: Comparing check-if-sorting-without-concat-is-better (0dab556) to a5c809f (merge-base) diff Run configurationrun benchmark sort_tpch10CPU Details (lscpu)Details
Resource Usagesort_tpch10 — base (merge-base)
sort_tpch10 — branch
File an issue against this benchmark runner |
|
ok, this improve perf, but need to be in arrow |
I do not plan to merge it like this, since it duplicating a lot of code from arrow, but I'm doing it as a POC
Which issue does this PR close?
N/A
Rationale for this change
I wanna see if there is a perf benefit of avoiding concat and sort/sort_to_indecies + take and replacing it with modified sort to indices + interleave
What changes are included in this PR?
copied arrow sort and modified it to work on slice of input array and make it return
Vec<(array_index, index_in_array)>forlexort_to_indicesor return Vec forsortWhat is the testing strategy for this PR?
Unit tests and existing tests
Are there any user-facing changes?
No
I want to run benchmarks the same benchmarks as:
to compare