Zheyuan Chen and GitHub
13d36cf891
ggml-webgpu: enable FLASH_ATTN_EXT on browser without subgroup matrix ( #22199 )
...
* ggml-webgpu: add tile flash attention fallback
* ggml-webgpu: add new fields and discard usage of mnk for tile version
* ggml-webgpu: modify the vec path to discard the mnk parameter
* ggml-webgpu: enable flash attention vec and tile version for broswer
* ggml-webgpu: stagging KV for flash attention tile version
* formatting
* turn on subgroup uniformity check
* remove Q_TILE as it is always 1 for vec path
* make row_max and exp_sum to local register
* make different bindings with same underlying buffer to have the same usage flags
* move path selection into the shader library and have the host consume a single flash-attn decision object.
* turn off skip_validation and address buffer overlapping when nwg==1
* formatting
* merge binding when kv overlap
2026-04-24 10:39:09 -07:00
a1cfb64530
ggml-webgpu: add vectorized flash attention ( #20709 )
...
* naive vectorized version
* add vectorized flash attention
* update vec version
* remove unused path and shader
* remove unused helper functions
* add comments
* remove pad path
* ggml-webgpu: fix flash-attn vec nwg=1 path and tighten vec specialization
* change back to vec4
* enable multi split
* enable vec path when:
- Q->ne[1] < 20
- Q->ne[0] % 32 == 0
- V->ne[0] % 4 == 0
- K->type == f16
* update flast_attn_vec_split.wgsl to reduce redundant workgroup barrier usage and use select
* enable vec path for q4 and q8
* flash-attn vec nwg=1 fast path (skip tmp/reduce staging)
* use packed f16 K loads in flash-attn vec split
* use packed f16 K loads in flash-attn vec split on host side
* tune flash-attn vec f16 VEC_NE by head dim
* cleanup
* cleanup
* keep host side clean
* cleanup host side
* change back to original host wait/submit behavior
* formatting
* reverted param-buffer pool r ecfactor
* add helper functions
* ggml-webgpu: move flash-attn vec pipeline caching back into shader lib
* ggml-webgpu: remove duplicate functions
* ggml-webgpu: reserve flash-attn vec scratch in dst buffer allocation
* ggml-webgpu: revert unrelated change
* ggml-webgpu: revert deleted comment
* disable uniformity check
* remove unnecessary change
* Update ggml/src/ggml-webgpu/wgsl-shaders/flash_attn_vec_split.wgsl
* Update ggml/src/ggml-webgpu/ggml-webgpu.cpp
---------
Co-authored-by: Reese Levine <reeselevine1@gmail.com >
2026-04-02 10:40:42 -07:00
bd90fc74c3
ggml-webgpu: improve flastAttention performance by software pipelining ( #19151 )
...
* webgpu : pipeline flash_attn Q/K loads in WGSL
* ggml-webgpu: unroll Q*K accumlation inner loop
* ggml-webgpu: vectorization
* ggml-webgpu: unrolling
* ggml-webgpu: remove redundant unrolling
* ggml-webgpu: restore the config
* ggml-webgpu: remove redundant comments
* ggml-webgpu: formatting
* ggml-webgpu: formatting and remove vectorization
* ggml-webgpu: remove unnecessary constants
* ggml-webgpu: change QKV buffer to read_write to pass validation
* ggml-webgpu: add explanation for the additional bracket around Q K accumulate
* Indentation and for -> if for tail
* Kick off CI on wgsl only commits
---------
Co-authored-by: Reese Levine <reeselevine1@gmail.com >
2026-01-29 14:05:30 -08:00