[js/webgpu] Optimize transpose #21964

qjia7 · 2024-09-03T06:20:34Z

Description

Fix bugs in previous implementation and add more situations to go the optimized path.

Below situations will go to the optimized path.

2d inputs or squeezed 2d inputs
channels last or channels first transpose. For example, channel last transpose: [1, 256, 512, 512] -> [1, 512, 512, 256]
For this case, the transpose becomes [256, 512x512] -> [512x512, 256]

Motivation and Context

For SD Turbo demo, the total transpose time becomes 39.98ms from 122.09ms. And the correspnding percents becomes 3.89% from 11.05% in this demo.

This PR will also help #21618, the total transpose time in that demo becomes 17.32 ms from 70.25 ms on my iGPUs.

Fix bugs in previous implementation and add more situations to go to the optimized path. 1. 2d inputs or squeezed 2d inputs 2. channel last or channel first transpose. For example, channel last transpose: [1, 256, 512, 512] -> [1, 512, 512, 256] For this case, the transpose becomes [256, 512x512] -> [512x512, 256] For SD Turbo demo, the total transpose time becomes 39.98ms from 122.09ms. And the correspnding percents becomes 3.89% from 11.05% in this demo.

qjia7 · 2024-09-03T06:41:36Z

@guschmue @fs-eire @satyajandhyala Please take a look, thanks.

fs-eire · 2024-09-03T08:54:01Z

/azp run Windows ARM64 QNN CI Pipeline,Windows x64 QNN CI Pipeline,Windows CPU CI Pipeline,Windows GPU CUDA CI Pipeline,Windows GPU DML CI Pipeline,Windows GPU Doc Gen CI Pipeline,Windows GPU TensorRT CI Pipeline,ONNX Runtime Web CI Pipeline,Linux CPU CI Pipeline,Linux CPU Minimal Build E2E CI Pipeline

fs-eire · 2024-09-03T08:54:04Z

/azp run Linux GPU CI Pipeline,Linux GPU TensorRT CI Pipeline,Linux OpenVINO CI Pipeline,Linux QNN CI Pipeline,MacOS CI Pipeline,orttraining-amd-gpu-ci-pipeline,orttraining-linux-ci-pipeline,orttraining-linux-gpu-ci-pipeline,orttraining-ortmodule-distributed,onnxruntime-binary-size-checks-ci-pipeline

fs-eire · 2024-09-03T08:54:06Z

/azp run Big Models,Linux Android Emulator QNN CI Pipeline,Android CI Pipeline,iOS CI Pipeline,ONNX Runtime React Native CI Pipeline

azure-pipelines · 2024-09-03T08:54:17Z

Azure Pipelines successfully started running 1 pipeline(s).

azure-pipelines · 2024-09-03T08:54:17Z

Azure Pipelines successfully started running 1 pipeline(s).

azure-pipelines · 2024-09-03T08:54:19Z

Azure Pipelines successfully started running 1 pipeline(s).

satyajandhyala · 2024-09-03T17:33:47Z

Can we add testcases that exercise the newly added code if not already exists?

qjia7 · 2024-09-04T05:30:41Z

Can we add testcases that exercise the newly added code if not already exists?

Done. Please take another look, thanks.

fs-eire · 2024-09-04T05:49:33Z

/azp run Windows ARM64 QNN CI Pipeline,Windows x64 QNN CI Pipeline,Windows CPU CI Pipeline,Windows GPU CUDA CI Pipeline,Windows GPU DML CI Pipeline,Windows GPU Doc Gen CI Pipeline,Windows GPU TensorRT CI Pipeline,ONNX Runtime Web CI Pipeline,Linux CPU CI Pipeline,Linux CPU Minimal Build E2E CI Pipeline

fs-eire · 2024-09-04T05:49:35Z

/azp run Linux GPU CI Pipeline,Linux GPU TensorRT CI Pipeline,Linux OpenVINO CI Pipeline,Linux QNN CI Pipeline,MacOS CI Pipeline,orttraining-amd-gpu-ci-pipeline,orttraining-linux-ci-pipeline,orttraining-linux-gpu-ci-pipeline,orttraining-ortmodule-distributed,onnxruntime-binary-size-checks-ci-pipeline

fs-eire · 2024-09-04T05:49:37Z

/azp run Big Models,Linux Android Emulator QNN CI Pipeline,Android CI Pipeline,iOS CI Pipeline,ONNX Runtime React Native CI Pipeline

azure-pipelines · 2024-09-04T05:49:47Z

Azure Pipelines successfully started running 1 pipeline(s).

azure-pipelines · 2024-09-04T05:49:48Z

Azure Pipelines successfully started running 1 pipeline(s).

azure-pipelines · 2024-09-04T05:49:50Z

Azure Pipelines successfully started running 1 pipeline(s).

guschmue added the ep:WebGPU ort-web webgpu provider label Sep 3, 2024

guschmue previously approved these changes Sep 3, 2024

View reviewed changes

add tests

7030e9b

qjia7 dismissed guschmue’s stale review via 7030e9b September 4, 2024 05:29

satyajandhyala approved these changes Sep 4, 2024

View reviewed changes

fs-eire approved these changes Sep 4, 2024

View reviewed changes

fs-eire merged commit a80bfed into microsoft:main Sep 4, 2024
50 of 53 checks passed

qjia7 mentioned this pull request Sep 5, 2024

[js/webgpu] Optimize InstanceNormalization #21995

Merged

qjia7 deleted the opt_transpose branch November 18, 2024 07:10

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[js/webgpu] Optimize transpose #21964

[js/webgpu] Optimize transpose #21964

qjia7 commented Sep 3, 2024 •

edited

Loading

qjia7 commented Sep 3, 2024

fs-eire commented Sep 3, 2024

fs-eire commented Sep 3, 2024

fs-eire commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

satyajandhyala commented Sep 3, 2024

qjia7 commented Sep 4, 2024

fs-eire commented Sep 4, 2024

fs-eire commented Sep 4, 2024

fs-eire commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

[js/webgpu] Optimize transpose #21964

[js/webgpu] Optimize transpose #21964

Conversation

qjia7 commented Sep 3, 2024 • edited Loading

Description

Motivation and Context

qjia7 commented Sep 3, 2024

fs-eire commented Sep 3, 2024

fs-eire commented Sep 3, 2024

fs-eire commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

azure-pipelines bot commented Sep 3, 2024

satyajandhyala commented Sep 3, 2024

qjia7 commented Sep 4, 2024

fs-eire commented Sep 4, 2024

fs-eire commented Sep 4, 2024

fs-eire commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

azure-pipelines bot commented Sep 4, 2024

qjia7 commented Sep 3, 2024 •

edited

Loading