This website requires JavaScript.
Explore
Help
Sign In
nikola
/
llama_cpp
Watch
1
Star
0
Fork
0
Code
Issues
Pull Requests
Actions
28
Packages
Projects
Releases
Wiki
Activity
Files
b14e3fb90ca8c760f4254ddc9aa7845ebbdb2edf
llama_cpp
/
ggml
/
src
/
ggml-webgpu
T
History
Masashi Yoshimura
and
GitHub
6e9007ae61
ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (
#24530
)
...
* Improve prefill speeds for i-quants * Fix #if defined() usage in preprocessor guards.
2026-06-14 18:15:30 -07:00
..
wgsl-shaders
ggml-webgpu: improve i-quants mul_mat performance and speed up prefill (
#24530
)
2026-06-14 18:15:30 -07:00
CMakeLists.txt
ggml-webgpu: FlashAttention refactor + standardize quantization support (
#23834
)
2026-06-04 08:05:04 +03:00
ggml-webgpu-shader-lib.hpp
ggml-webgpu: Add clang-format job (
#24308
)
2026-06-08 20:54:24 -07:00
ggml-webgpu.cpp
Remove padding and multiple D2D copies for MTP (
#24086
)
2026-06-10 23:21:16 +05:30
pre_wgsl.hpp
ggml-webgpu: FlashAttention refactor + standardize quantization support (
#23834
)
2026-06-04 08:05:04 +03:00