Skip to content

CUB Deviceless Scratch Size - #10998

Draft
Jacobfaib wants to merge 4 commits into
NVIDIA:mainfrom
Jacobfaib:jacobf/2026-08-25/deviceless-radix-sort-scratch
Draft

CUB Deviceless Scratch Size#10998
Jacobfaib wants to merge 4 commits into
NVIDIA:mainfrom
Jacobfaib:jacobf/2026-08-25/deviceless-radix-sort-scratch

Conversation

@Jacobfaib

Copy link
Copy Markdown
Contributor

Description

closes #9814

Checklist

  • New or existing tests cover these changes.
  • The documentation is up to date with these changes.

@copy-pr-bot

copy-pr-bot Bot commented Aug 25, 2026

Copy link
Copy Markdown
Contributor

Auto-sync is disabled for draft pull requests in this repository. Workflows must be run manually.

Contributors can view more details about this message here.

@cccl-authenticator-app cccl-authenticator-app Bot moved this from Todo to In Progress in CCCL Aug 25, 2026
@Jacobfaib Jacobfaib changed the title wip CUB Deviceless Scratch Size Aug 25, 2026
@Jacobfaib

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

This comment has been minimized.

@Jacobfaib

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

This comment has been minimized.

@Jacobfaib

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

This comment has been minimized.

@Jacobfaib

Copy link
Copy Markdown
Contributor Author

/ok to test

@github-actions

Copy link
Copy Markdown
Contributor

🔬 CUB benchmark SASS comparison

⚠️ The SASS changed for 3 of 84 CUB benchmark target(s). A benchmark run may be necessary

How to request a benchmark run
Request a CUB benchmark run for this PR:

1. Replace the `benchmarks:` block of ci/bench.yaml with exactly this:

benchmarks:
  filters:
    cub:
      - '^cub\.bench\.transform\.applications\.P1\.pytorch\.base$'
      - '^cub\.bench\.transform\.pytorch\.base$'
      - '^cub\.bench\.transform\.tile\.pytorch\.base$'
  gpus:
    - "h100"   # pick the GPUs that this change can affect

2. Commit with `[bench-only]` at the end of the commit summary, so that
   the unrelated CI jobs are skipped. Then push.

ci/bench.yaml must match ci/bench.template.yaml before the PR can merge.
Reset it once the measurement is done.
Run Value
Baseline fe278dc32e376737d046d27c70f6ee858ec1e89b
Tested HEAD
Architectures 75-real;80-real;90-real;100-real;110-real;120-real;120-virtual
Targets with a SASS change
Target Architectures with a SASS change
cub.bench.transform.applications.P1.pytorch.base sm_120, sm_75, sm_80, sm_100, sm_110, sm_90
cub.bench.transform.pytorch.base sm_120, sm_75, sm_80, sm_100, sm_110, sm_90
cub.bench.transform.tile.pytorch.base sm_120, sm_75, sm_80, sm_100, sm_110, sm_90

‼️ Summary of Differences ‼️

Showing 3/3 summaries.

cub.bench.transform.applications.P1.pytorch.base - sm_120

Showing 40/326 diff lines, 72 changes. - ⬇️ Full diff

--- base/cub.bench.transform.applications.P1.pytorch.base.sm_120
+++ test/cub.bench.transform.applications.P1.pytorch.base.sm_120
@@ -113796,7 +113796,7 @@
 @!P0 FMUL R0, RZ, R21 ;
 @!P0 MOV R18, RZ ;
 @!P0 BRA <+0x360> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R17, 0x1000000, RZ ;
 HFMA2 R23, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -113878,7 +113878,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R20, R20, 0x1000000, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -113963,7 +113963,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R21, 0x1000000, RZ ;
 MOV.64 R4, RZ ;
 MOV R25, RZ ;
@@ -114049,7 +114049,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R21, 0x1000000, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -114135,7 +114135,7 @@
 @!P0 BRA <+0x3b0> ;
cub.bench.transform.pytorch.base - sm_120

Showing 40/164 diff lines, 36 changes. - ⬇️ Full diff

--- base/cub.bench.transform.pytorch.base.sm_120
+++ test/cub.bench.transform.pytorch.base.sm_120
@@ -130068,7 +130068,7 @@
 @!P0 FMUL R0, RZ, R21 ;
 @!P0 MOV R18, RZ ;
 @!P0 BRA <+0x360> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R17, 0x1000000, RZ ;
 HFMA2 R23, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -130150,7 +130150,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R20, R20, 0x1000000, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -130235,7 +130235,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R21, 0x1000000, RZ ;
 MOV.64 R4, RZ ;
 MOV R25, RZ ;
@@ -130321,7 +130321,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R21, 0x1000000, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -130407,7 +130407,7 @@
 @!P0 BRA <+0x3b0> ;
cub.bench.transform.tile.pytorch.base - sm_120

Showing 40/263 diff lines, 58 changes. - ⬇️ Full diff

--- base/cub.bench.transform.tile.pytorch.base.sm_120
+++ test/cub.bench.transform.tile.pytorch.base.sm_120
@@ -119286,7 +119286,7 @@
 @!P0 FMUL R0, RZ, R21 ;
 @!P0 MOV R18, RZ ;
 @!P0 BRA <+0x360> ;
-LDC.64 R16, c[0x4][0x288] ;
+LDC.64 R16, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R21, 0x100, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -119368,7 +119368,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R23, 0x100, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -119453,7 +119453,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R23, 0x100, RZ ;
 MOV.64 R4, RZ ;
 MOV R25, RZ ;
@@ -119539,7 +119539,7 @@
 @!P0 BRA <+0x3b0> ;
 FSETP.NEU.AND P0, PT, |R23|, +INF , PT ;
 @!P0 BRA <+0x370> ;
-LDC.64 R6, c[0x4][0x288] ;
+LDC.64 R6, c[0x4][0x298] ;
 IMAD.SHL.U32 R0, R23, 0x100, RZ ;
 HFMA2 R25, -RZ, RZ, 0, 0 ;
 MOV.64 R4, RZ ;
@@ -119625,7 +119625,7 @@
 @!P0 BRA <+0x3b0> ;

@github-actions

Copy link
Copy Markdown
Contributor

😬 CI Workflow Results

🟥 Finished in 2h 25m: Pass: 98%/331 | Total: 10d 02h | Max: 1h 27m | Hits: 59%/821060

See results here.

@Jacobfaib Jacobfaib linked an issue Aug 27, 2026 that may be closed by this pull request
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

Expose cuda::execution::runs_on [FEA]: CUB Sort Deviceless Scratch Size

1 participant