[ExecuTorch][WebGPU] Add Gemma 4 plain export and artifact contract - #21673
[ExecuTorch][WebGPU] Add Gemma 4 plain export and artifact contract#21673JCNTH wants to merge 1 commit into
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21673
Note: Links to docs will display an error until the docs builds have been completed. ✅ You can merge normally! (3 Unrelated Failures)As of commit a0ca711 with merge base ceca90f ( FLAKY - The following jobs failed but were likely due to flakiness present on trunk:
BROKEN TRUNK - The following job failed but were present on the merge base:👉 Rebase onto the `viable/strict` branch to avoid these failures
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
This PR needs a
|
|
Hi @JCNTH! Thank you for your pull request. We require contributors to sign our Contributor License Agreement, and yours needs attention. You currently have a record in our system, but the CLA is no longer valid, and will need to be resubmitted. ProcessIn order for us to review and merge your suggested changes, please sign at https://code.facebook.com/cla. If you are contributing on behalf of someone else (eg your employer), the individual CLA may not be sufficient and your employer may need to sign the corporate CLA. Once the CLA is signed, our tooling will perform checks and validations. Afterwards, the pull request will be tagged with If you have received this in error or have any questions, please contact us at cla@meta.com. Thanks! |
Stack from ghstack (oldest at bottom):
Add an independently buildable plain Gemma 4 export, manifest, and browser adapter. The export path keeps XNNPACK defaults intact while applying scoped WebGPU rewrites, BSHD KV-cache semantics, independent 512-token input and 8960-token context bounds, compact token output, and an ordered three-PTD contract. Artifact manifests remain source-only and distinguish historical behavior evidence from final-source rebuilds.
Key changes:
export_gemma4.py/webgpu_partitioner.pyapplies edge rewrites before graph-identity-preserving partitioning.webgpu_artifact_manifest.pybinds acquisition, architecture, method/delegate census, and ordered PTE/PTD roles.gemma4_plain_wasm.cppexposes the three-argument load ABI, compact int64 output, reset reload, and owned context lifecycle.createWebGPULlama/webgpu_llamawith valid--no-entryand explicit browser exports.No PTE/PTD/WASM binary or dashboard dependency is committed.
Co-authored-with: Claude Code.
Differential Revision: D115234082