Johannes Gäßler and GitHub
a8fd165fec
CUDA: lower-case PCI bus id, standardize for ggml ( #22820 )
2026-05-08 10:09:38 +02:00
Johannes Gäßler and GitHub
9dbb372610
Github: update issue templates ( #22594 )
2026-05-02 07:56:13 +02:00
Johannes Gäßler and GitHub
e82aaf2587
CUDA: fix tile FA kernel on Pascal ( #22541 )
2026-04-30 13:04:50 +02:00
Johannes Gäßler and GitHub
739393beeb
TP: fix delayed AllReduce + zero-sized slices ( #22489 )
2026-04-29 08:55:07 +02:00
Johannes Gäßler and GitHub
7ec36aa861
Github: set meta backend code owner ( #22388 )
2026-04-26 13:34:13 +02:00
Johannes Gäßler and GitHub
9725a313be
CUDA: reduce MMQ stream-k overhead ( #22298 )
...
* CUDA: reduce MMQ stream-k overhead
* use 32 bit integers for kbc
2026-04-25 14:15:03 +02:00
Johannes Gäßler and GitHub
fb19f94c71
TP: fix 0-sized tensor slices, AllReduce fallback ( #21808 )
...
* TP: fix 0-sized tensor slices, AllReduce fallback
* fix layer structure <-> GPU count aliasing
* add missing std::fill
* fix CUDA device set, max ggml ctx size
2026-04-20 18:09:39 +02:00
Johannes Gäßler and GitHub
4eac5b4509
CUDA: refactor mma data loading for AMD ( #22051 )
...
* CUDA: refactor mma data loading for AMD
* fix CDNA MMQ occupancy
* fix CDNA3 mma
* fix RDNA3 compile
2026-04-19 18:26:59 +02:00
Johannes Gäßler and GitHub
fd1c0ec3f0
llama: fit ctx size for CPU only ( #21568 )
2026-04-18 08:16:04 +02:00
Johannes Gäßler and GitHub
a6206958d2
CUDA: require explicit opt-in for P2P access ( #21910 )
2026-04-15 16:01:46 +02:00
Johannes Gäßler and GitHub
014dca49d6
CUDA: manage NCCL communicators in context ( #21891 )
...
* CUDA: manage NCCL communicators in context
* add check that all backends are CUDA
* remove unused vector, limit init to > 1 GPUs
* fix warnings
* fix cuda device, cache allreduce
2026-04-15 15:58:40 +02:00
Johannes Gäßler and GitHub
ff5ef82786
CUDA: skip compilation of superfluous FA kernels ( #21768 )
2026-04-11 18:52:11 +02:00
Johannes Gäßler and GitHub
865ff06b2f
TP: fix Qwen 3 Next data split ( #21732 )
2026-04-11 09:23:42 +02:00
Johannes Gäßler and GitHub
0893f50f2d
common: mark --split-mode tensor as experimental ( #21684 )
2026-04-10 12:27:27 +02:00
d6f3030047
ggml: backend-agnostic tensor parallelism (experimental) ( #19378 )
...
* ggml: backend-agnostic tensor parallelism
* support for GPT-OSS, Qwen 3 MoE
* partial Vulkan fix
* add support for 4/8 GPUs
* unconditional peer access
* re-use buffers + ggml contexts
* fix output pattern
* NCCL support
* GGML: HIP: add RCCL support
* Remove shfl and AllReduce from backend interface
* move allocation workaround out of ggml-alloc.c
* 2d tensor set/get support
* Fix the seg fault without NCCL
* Apply suggestion from JohannesGaessler
* support for tensor dims % n_devs != 0
* fix view_offs scaling
* arbitrary num. of GPUs/tensor split
* fix compilation
* better granularity estimate
* Support device-specific host buffer types if all underlying backends expose the same type. This allows using pinned memory instead of pageable memory for CUDA.
Fix compilation errors.
* partial Qwen 3 Next support
* Fix qwen3 30b (#8 )
* Fix crash with Qwen-30B-A3B Q4_0
Qwen-30B-A3B Q4_0 has an intermediate dimension of 768. Using a granularity of 256 forces an uneven split between GPUs, which is not supported by the current implementation.
* Decide block size based on tensor quantization type
* Fix crashes due to KV cache serialization (#9 )
KV cache serialization requires non-zero offsets on the tensor. Add support in the meta backend to set/get a tensor with a non-zero offset.
* metal : fix build (#7 )
* static memory allocations, fix usage count
* fix tensor granularity
* more even memory distribution
* use BF16 for allreduce
* rebase fixup
* better error message for unsupported architectures
* Fix device mismatch during scatter of allReduce. (#11 )
There is a mismatch between the dst buffer device and the backend device, causing the use of sync copies
* Enable the previous allreduce implementation. It is better in both perf and stability (#12 )
* delay AllReduce for Moe for less I/O
* build : clean-up compile warnings
* backend : move most of the meta backend API to ggml-backend-impl.h
* cont : hide unused public API in the implementation
* llama : use llama_device + remove ggml_backend_dev_is_meta()
* ggml-backend : remove unused alloc include
* minor : remove regex include
* ggml : introduce ggml-ext.h for staging new APIs
* rebase fixup
* fix tests
* llama : more robust logic for determining Meta devices (#16 )
* llama : more robust logic for determining Meta devices
* cont : fix devs size check
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* cont : fix log type
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
* disable roundtrip for meta backend
* fix arch selection
* Qwen 3.5 support
* fix Gemma 4 MoE
* fix OpenVino, SYCL
* fix test-llama-archs for CPU-only builds
* Fix Qwen 3.5 MoE
* disable meta backend tests for WebGPU
* tests : filter CPU-based devices from the Meta backend tests (#17 )
* meta : formatting, naming, indentation (#18 )
* formatting : llama-model.cpp
* formatting : ggml-ext.h
* formatting : ggml-backend-meta.cpp
* meta : add TODO
* add documentation
* better error messages
* fix GPT-OSS
---------
Co-authored-by: Carl Philipp Klemm <carl@uvos.xyz >
Co-authored-by: Gaurav Garg <gaugarg@nvidia.com >
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
2026-04-09 16:42:19 +02:00
Johannes Gäßler and GitHub
a8ec0df461
llama: remove per-arch tensor name lists ( #21531 )
2026-04-07 15:02:03 +02:00
Johannes Gäßler and GitHub
86221cf6da
CUDA: fix FA kernel selection logic ( #21271 )
2026-04-01 22:28:19 +03:00
36dafba5c4
llama: fix llama-model-saver ( #20503 )
...
* llama : add fd-based model loading via llama_model_load_from_fd
* llama : address review feedback for fd-based model loading
* llama : use FILE pointer instead of fd in public API
* llama : use FILE pointer consistently, address review feedback
* fixup
* fix tensor names
* fix llama-model-saver
* roundtrip tests
* fixup
* refactor tests
* fix prints
* fix model saving
* fix CI, disable Chameleon
* print seed
---------
Co-authored-by: Siddhesh2377 <siddheshsonar2377@gmail.com >
2026-03-25 12:53:16 +02:00
Johannes Gäßler and GitHub
bd3f1d9d65
CUDA: fix BF16 FA compilation ( #20865 )
2026-03-22 17:53:33 +01:00
Johannes Gäßler and GitHub
ae40cd27c8
CUDA: limit number of FA stream-k CUDA blocks ( #20586 )
2026-03-15 18:30:47 +01:00
Johannes Gäßler and GitHub
a976ff081b
llama: end-to-end tests ( #19802 )
...
* tests: add end-to-end tests per model architecture
* fixup for rebase
* fix use-after-free in llama-model-loader.cpp
* fix CI
* fix WebGPU
* fix CI
* disable CI for macOS-latest-cmake-arm64
* use expert_weights_scale only if != 0.0f
* comments
2026-03-08 12:30:21 +01:00
Johannes Gäßler and GitHub
2850bc6a13
ggml-cpu: fix data race for debug asserts ( #20148 )
2026-03-06 09:12:49 +01:00
Johannes Gäßler and GitHub
7f5ee54968
ggml: fix ggml_is_contiguous_n for ne == 1 ( #20092 )
2026-03-04 12:04:31 +01:00
Johannes Gäßler and GitHub
c78e682245
CUDA: fix kernel selection logic for tile FA ( #19686 )
...
* CUDA: fix kernel selection logic for tile FA
* add comment
2026-02-19 12:42:58 +01:00
Johannes Gäßler and GitHub
ada90bf2ba
docs: ban AI for issues and discussions [no CI] ( #19512 )
2026-02-11 12:49:40 +01:00
Johannes Gäßler and GitHub
59377a6c87
ggml-backend: fix async set/get fallback sync ( #19179 )
2026-02-02 10:00:05 +01:00
Johannes Gäßler and GitHub
a5bb8ba4c5
CUDA: tune GLM 4.7 Flash FA kernel selection logic ( #19097 )
2026-01-27 14:28:56 +01:00
Johannes Gäßler and GitHub
b0311c16d2
CUDA: fix padding of GQA to power of 2 in FA ( #19115 )
2026-01-26 23:24:58 +01:00
Johannes Gäßler and GitHub
0c21677e43
CUDA: faster FA for GQA > 1 but not power of 2 ( #19092 )
2026-01-25 21:19:47 +01:00
Johannes Gäßler and GitHub
e9fd8dcab4
llama-fit-params: keep explicit --ctx-size 0 ( #19070 )
2026-01-24 22:13:08 +01:00
Johannes Gäßler and GitHub
4e5b83b226
GGUF: check that tensor size is representable ( #19072 )
2026-01-24 21:57:51 +01:00
Johannes Gäßler and GitHub
8f91ca54ec
CUDA: re-use MLA K data for V in MMA FA ( #19057 )
2026-01-24 10:09:36 +01:00
Johannes Gäßler and GitHub
e2baf02162
CUDA: fix alignment check for FA ( #19023 )
2026-01-22 20:39:25 +01:00
Johannes Gäßler and GitHub
5c662d21a3
CUDA: fix allignment on register spill for FA ( #18815 )
2026-01-15 15:14:50 +01:00
Johannes Gäßler and GitHub
c1e79e610f
doc: ban AI-generated PR descriptions [no ci] ( #18765 )
2026-01-13 13:43:12 +01:00
Johannes Gäßler and GitHub
d2ff4e23ac
HIP: adjust RDNA3.5 MMQ kernel selction logic ( #18666 )
2026-01-10 17:19:01 +01:00
Johannes Gäßler and GitHub
64848deb18
llama-fit-params: free memory target per device ( #18679 )
2026-01-08 10:07:58 +01:00
Johannes Gäßler and GitHub
68b4d516c3
llama-params-fit: fix last devices with low VRAM ( #18494 )
2026-01-06 20:02:30 +01:00
Johannes Gäßler and GitHub
df17a4c94f
CUDA: fix FA FP16 accumulator overflow for Granite ( #18614 )
2026-01-05 19:51:13 +01:00
Johannes Gäßler and GitHub
0f2e42ca1d
CUDA: only allocate FA tmp buffer if needed ( #18564 )
2026-01-03 13:55:53 +01:00
Johannes Gäßler and GitHub
ecc343de63
CUDA: fix KQ max calculation ( #18487 )
2025-12-31 09:37:00 +01:00
Johannes Gäßler and GitHub
0bd1212a43
CUDA: fix replacment of bad archs in CMake ( #18457 )
2025-12-29 17:58:20 +01:00
Johannes Gäßler and GitHub
e70e640db3
CUDA: Blackwell features for non-native builds ( #18436 )
2025-12-29 09:35:42 +01:00
Johannes Gäßler and GitHub
f8d561eb87
llama-fit-params: fix step size for last device ( #18415 )
2025-12-28 10:52:09 +01:00
e59efe6a78
github: update issue templates [no ci] ( #18410 )
...
* github: update issue templates [no ci]
* Apply suggestions from code review
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2025-12-28 10:50:56 +01:00
Johannes Gäßler and GitHub
a4bf35889e
llama-fit-params: fix overflow check ( #18354 )
2025-12-27 20:20:45 +01:00
Johannes Gäßler and GitHub
026d2ad472
llama: fix magic number of 999 for GPU layers ( #18266 )
...
* llama: fix magic number of 999 for GPU layers
* use strings for -ngl, -ngld
* enacapsulate n_gpu_layers, split_mode
2025-12-27 20:18:35 +01:00
Johannes Gäßler and GitHub
a52dc60ba3
llama_fit_params: return enum for fail vs. error ( #18374 )
2025-12-27 09:59:19 +01:00
Johannes Gäßler and GitHub
9045c9afe5
llama-fit-params: fix Gemma 3 calculation ( #18372 )
2025-12-27 09:56:04 +01:00
Johannes Gäßler and GitHub
147a521636
tool/ex/tests: consistently free ctx, then model ( #18168 )
2025-12-22 11:00:37 +01:00
Johannes Gäßler and GitHub
0e1ccf15c7
llama: fix RPC for -fit on ( #18233 )
2025-12-21 19:33:08 +01:00
Johannes Gäßler and GitHub
57c1e05643
llama: offload output layer to GPU first ( #18148 )
2025-12-18 08:12:18 +01:00
Johannes Gäßler and GitHub
8dcc3662a2
llama-fit-params: fix memory print ( #18136 )
2025-12-17 21:10:03 +01:00
Johannes Gäßler and GitHub
a2c199e479
common: clarify instructions for bug reports ( #18134 )
2025-12-17 18:44:13 +01:00
Johannes Gäßler and GitHub
6f1f6a961a
Github: ask for -v logs for params_fit [no ci] ( #18128 )
2025-12-17 13:46:48 +01:00
Johannes Gäßler and GitHub
d0794e89d9
llama-fit-params: force disable mlock ( #18103 )
2025-12-17 00:50:12 +01:00
Johannes Gäßler and GitHub
9dcac6cf9f
llama-fit-params: lower ctx size for multi GPU ( #18101 )
2025-12-17 00:49:34 +01:00
Johannes Gäßler and GitHub
0e49a7b8b4
llama-fit-params: fix underflow for dense models ( #18095 )
2025-12-17 00:47:37 +01:00
Johannes Gäßler and GitHub
4164596c76
llama-fit-params: QoL impr. for prints/errors ( #18089 )
2025-12-17 00:03:19 +01:00
Johannes Gäßler and GitHub
ec98e20021
llama: fix early stop in params_fit if ctx is set ( #18070 )
2025-12-16 14:24:00 +01:00
Johannes Gäßler and GitHub
b1f3a6e5db
llama: automatically set parameters not set by the user in such a way that maximizes GPU utilization ( #16653 )
...
* llama: automatically fit args to free memory
llama-fit-params tool
* fix CI
* hints for bug reports, ensure no reallocation
* fix segfault with Vulkan
* add llama-fit-params to CI
* fix CI
* fix CI
* fix CI
* minor adjustments
* fix assignment of 1 dense layer
* fix logger not being reset on model load failure
* remove --n-gpu-layer hint on model load failure
* fix llama-fit-params verbosity
* fix edge case
* fix typo [no ci]
2025-12-15 09:24:59 +01:00
Johannes Gäßler and GitHub
482211438d
CUDA: fix overflow in MMA kernel without stream-k ( #17939 )
2025-12-12 17:43:58 +01:00
Johannes Gäßler and GitHub
17f7f4baad
CUDA: fix unpadded strides in MMA FA kernel ( #17891 )
2025-12-10 12:39:56 +01:00
Johannes Gäßler and GitHub
48f47565a7
docs: clarify that CPU support should be first ( #17886 )
2025-12-09 20:10:36 +01:00
Johannes Gäßler and GitHub
0cdce38a97
CUDA: fix FP16 overflow in tile FA kernel ( #17875 )
2025-12-09 09:34:02 +01:00
Johannes Gäßler and GitHub
f334b79494
HIP: fix RDNA3 FP16/BF16 matrix multiplication ( #17817 )
2025-12-06 13:45:36 +01:00
Johannes Gäßler and GitHub
6016d0bd41
HIP : fix RDNA4 build ( #17792 )
2025-12-05 13:47:52 +01:00
Johannes Gäßler and GitHub
e95d0bc8fd
CUDA: fix FA VKQ accumulator overflow ( #17746 )
2025-12-05 09:18:10 +01:00
2e1c9cd814
CUDA: generalized (mma) FA, add Volta support ( #17505 )
...
* CUDA: generalized (mma) FA, add Volta support
* use struct for MMA FA kernel config
---------
Co-authored-by: Aman Gupta <aman>
2025-12-03 16:57:05 +01:00
Johannes Gäßler and GitHub
73955f7d2a
CUDA: no FP16 arithmetic for vector FA kernel ( #17558 )
2025-11-28 10:29:09 +01:00
Johannes Gäßler and GitHub
5d6838b74f
CUDA: static assert to prevent misuse of memcpy_1 ( #17198 )
2025-11-12 23:13:55 +01:00
Johannes Gäßler and GitHub
e14e842e87
CUDA: fix MMQ stream-k fixup ne1 indices ( #17089 )
2025-11-08 08:26:18 +01:00
6515610506
CUDA: fix should_use_mmvf for ne11 == 1 ( #17085 )
...
* CUDA: fix should_use_mmvf for ne11 == 1
* Apply suggestion from @am17an
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
---------
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
2025-11-07 20:53:14 +01:00
Johannes Gäßler and GitHub
aa374175c3
CUDA: fix crash on uneven context without FA ( #16988 )
2025-11-06 14:05:47 +01:00
Johannes Gäßler and GitHub
22c8c3c6ad
docs: explain CUDA 11 compilation [no ci] ( #16824 )
2025-11-06 08:14:35 +01:00
31c511a968
CUDA: Volta tensor core support for MMF ( #16843 )
...
* CUDA: Volta tensor core support for MMF
* more generic checks for hardware support
* Update ggml/src/ggml-cuda/mmf.cuh
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
---------
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
2025-10-31 15:57:19 +01:00
Johannes Gäßler and GitHub
7a0e900e36
llama: consistent ctx <-> buf order for KV cache ( #16746 )
2025-10-28 11:23:54 +01:00
Johannes Gäßler and GitHub
80d28f104c
HIP: fix AMDGPU_TARGETS, update documentation ( #16803 )
2025-10-27 21:39:49 +01:00
Johannes Gäßler and GitHub
945501f5ea
llama: fix leaked buffers for mmap + split files ( #16765 )
2025-10-27 09:17:31 +01:00
Johannes Gäßler and GitHub
0bf47a1dbb
server: add memory breakdown print ( #16740 )
2025-10-23 21:30:17 +02:00
Johannes Gäßler and GitHub
51d1a8c997
CUDA: better error for FA kernel with 0 occupancy ( #16643 )
2025-10-21 15:27:53 +02:00
Johannes Gäßler and GitHub
ee09828cb0
HIP: fix GPU_TARGETS ( #16642 )
2025-10-18 14:47:32 +02:00
Johannes Gäßler and GitHub
66b0dbcb2d
llama-model: fix insonsistent ctxs <-> bufs order ( #16581 )
2025-10-17 17:41:09 +02:00
Johannes Gäßler and GitHub
9c7185dd28
CUDA: enable FA for FP32 KV cache ( #16546 )
2025-10-14 14:22:47 +02:00
Johannes Gäßler and GitHub
7049736b2d
CUDA: fix numerical issues in tile FA kernel ( #16540 )
2025-10-13 17:29:45 +03:00
Johannes Gäßler and GitHub
11f0af5504
CUDA: faster tile FA, add oob checks, more HSs ( #16492 )
2025-10-11 20:54:32 +02:00
Johannes Gäßler and GitHub
75a3a6c2cd
CUDA: refactor and deduplicate vector FA kernels ( #16208 )
...
* CUDA: refactor and deduplicate vector FA kernels
2025-09-27 18:45:07 +02:00
Johannes Gäßler and GitHub
4cdd0bb453
docs: fix typo [no ci] ( #16244 )
2025-09-25 12:12:27 +03:00
Johannes Gäßler and GitHub
e789095502
llama: print memory breakdown on exit ( #15860 )
...
* llama: print memory breakdown on exit
2025-09-24 16:53:48 +02:00
Johannes Gäßler and GitHub
368560a1e3
CUDA: fix compilation on CC 6.0 ( #16091 )
2025-09-18 19:28:32 +02:00
Johannes Gäßler and GitHub
c959b676be
CUDA: fix FA occupancy, optimize tile kernel ( #15982 )
2025-09-17 15:32:42 +02:00
Johannes Gäßler and GitHub
0e6ff0046f
CUDA: larger SRAM reads for tile FA, AMD FP16 dot ( #15927 )
...
* CUDA: larger SRAM reads for tile FA, AMD FP16 dot
* fix logic for availability of v_dot2_f32_f16
2025-09-11 21:19:58 +02:00
Johannes Gäßler and GitHub
17bc5a815f
HIP: use v_dot2_f32_f16 instruction for FA ( #15884 )
2025-09-09 14:04:43 +02:00
Johannes Gäßler and GitHub
550cf726e1
CUDA: fix GET_ROWS for large tensors ( #15882 )
2025-09-09 08:11:01 +02:00
Johannes Gäßler and GitHub
79bc429262
CUDA: faster tile FA (Pascal/AMD), headsize 256 ( #15769 )
2025-09-07 00:26:28 +02:00
Johannes Gäßler and GitHub
01806e7771
ggml-cpu: document use of "free" memory [no ci] ( #15834 )
2025-09-06 13:28:44 +02:00
Johannes Gäßler and GitHub
5143fa895e
CUDA: fastdiv, launch bounds for mmvq + q8_1 quant ( #15802 )
...
* CUDA: fastdiv, launch bounds for mmvq + q8_1 quant
2025-09-05 16:07:02 +02:00
Johannes Gäßler and GitHub
c466abe158
llama: -fa 1/0/-1 aliases for -fa on/off/auto ( #15746 )
2025-09-02 18:17:26 +02:00
Johannes Gäßler and GitHub
5d804a4938
ggml-backend: raise GGML_MAX_SPLIT_INPUTS ( #15722 )
2025-09-01 16:14:55 -07:00
Johannes Gäßler and GitHub
e81b8e4b7f
llama: use FA + max. GPU layers by default ( #15434 )
...
* llama: use max. GPU layers by default, auto -fa
* ggml-backend: abort instead of segfault
2025-08-30 16:32:10 +02:00