Georgi Gerganov and GitHub
8ea8fee966
gitignore : add .pi + personal SYSTEM.md ( #22316 )
...
* gitignore : add .pi + personal SYSTEM.md
* cont : fix requirements heading in PR template
* cont : shorten line
2026-04-25 09:20:45 +03:00
Georgi Gerganov and GitHub
15fa3c493b
metal : print GPU description ( #22318 )
2026-04-24 13:56:03 +03:00
Georgi Gerganov and GitHub
e583f3b4f5
ggml : minor coding style ( #22308 )
2026-04-24 11:02:00 +03:00
Georgi Gerganov and GitHub
017f090442
jinja : remove unused header ( #22310 )
2026-04-24 11:01:46 +03:00
Georgi Gerganov and GitHub
ffdd983fb8
server : fix swa-full logic ( #22288 )
2026-04-24 10:17:37 +03:00
Georgi Gerganov and GitHub
8635e221c8
metal : fix event synchronization ( #22260 )
2026-04-23 08:22:49 +03:00
Georgi Gerganov and GitHub
930e0210d1
gitignore: add AGENTS.local.md ( #22246 )
...
* gitignore: add AGENTS.local
Assisted-by: llama.cpp:local pi
Signed-off-by: Georgi Gerganov <ggerganov@gmail.com >
* gitignore: rename AGENTS.local to AGENTS.local.md
Assisted-by: llama.cpp:local pi
Signed-off-by: Georgi Gerganov <ggerganov@gmail.com >
---------
Signed-off-by: Georgi Gerganov <ggerganov@gmail.com >
2026-04-23 08:22:24 +03:00
Georgi Gerganov and GitHub
96c1db26c4
ggml-base: use MATH_LIBRARY variable instead of hardcoded 'm' ( #22239 )
...
Fixes #22237 — the find_library(MATH_LIBRARY m) result was being
discarded and the target linked against the literal 'm' string.
This prevents users from overriding the math library (e.g. for AMD AOCL)
via CMake variables. Now the discovered MATH_LIBRARY is used directly.
2026-04-23 08:22:08 +03:00
Georgi Gerganov and GitHub
bcb5eeb645
speculative-simple : add checkpoint support ( #22227 )
...
* speculative-simple : add checkpoint support
* cont : fix build
2026-04-22 15:44:45 +03:00
Georgi Gerganov and GitHub
84652b80cf
arg : add --spec-default ( #22223 )
2026-04-21 19:52:02 +03:00
Georgi Gerganov and GitHub
7fc1c4ef78
metal : workaround macOS GPU interactivity watchdog ( #22216 )
2026-04-21 17:24:55 +03:00
Georgi Gerganov and GitHub
cd03ec7642
llama-ext : fix exports ( #22202 )
2026-04-21 11:04:46 +03:00
Georgi Gerganov
4889afba5f
sync : ggml
2026-04-21 11:04:21 +03:00
Georgi Gerganov
041fe83d74
ggml : bump version to 0.10.0 (ggml/1463)
2026-04-21 11:04:21 +03:00
Georgi Gerganov and GitHub
cfe9838d26
fit-params : refactor + add option to output estimated memory per device ( #22171 )
...
* fit-params : add option to output estimated memory per device
* cont : minor
* cont : refactor
* cont : move fit params implementation to libcommon
* cont : header
* cont : headers
* cont : codeowners
2026-04-21 09:54:36 +03:00
Georgi Gerganov and GitHub
cf8b0dbda9
server : remove /api endpoints ( #22165 )
...
* server : remove /api endpoints
* cont : remove /api/tags
2026-04-20 20:41:19 +03:00
Georgi Gerganov and GitHub
de71b5f81c
server : refactor "use checkpoint" logic ( #22114 )
2026-04-20 08:42:37 +03:00
Georgi Gerganov and GitHub
6990e2f1f7
libs : rename libcommon -> libllama-common ( #21936 )
...
* cmake : allow libcommon to be shared
* cmake : rename libcommon to libllama-common
* cont : set -fPIC for httplib
* cont : export all symbols
* cont : fix build_info exports
* libs : add libllama-common-base
* log : add common_log_get_verbosity_thold()
2026-04-17 11:11:46 +03:00
Georgi Gerganov and GitHub
c0de6eda72
metal : fix FA support logic ( #21898 )
2026-04-14 17:32:29 +03:00
Georgi Gerganov and GitHub
f4b5bf2f32
ci : re-enable mac workflows ( #21894 )
...
* ci : re-enable mac workflows
* vulkan : fix compile warning
2026-04-14 15:58:09 +03:00
Georgi Gerganov and GitHub
5e9c635463
metal : add missing mm-id specializations for q1_0 ( #21662 )
2026-04-09 10:54:00 +03:00
4a05e0c566
webui : send both backend_sampling == false/true ( #18781 )
...
* webui : send both backend_sampling == false/true
* feat: Parameter sync
---------
Co-authored-by: Aleksander Grygier <aleksander.grygier@gmail.com >
2026-04-08 16:35:52 +02:00
Georgi Gerganov and GitHub
5764d7c6a6
gemma : perform per-layer projections in the first layer ( #21612 )
...
* gemma : reduce graph splits by keeping per-layer ops in the input layer
* gemma : put the per-layer proj in the first layer
* cont : move the projection before the layer loop
2026-04-08 16:06:30 +03:00
Georgi Gerganov and GitHub
ae65fbdf33
tests : remove obsolete .mjs script ( #21615 )
2026-04-08 13:20:46 +03:00
Georgi Gerganov and GitHub
4eb19514dd
kv-cache : support attention rotation for heterogeneous iSWA ( #21513 )
...
* kv-cache : support attention rotation for heterogeneous iSWA
* cont : remove assert
2026-04-07 20:31:28 +03:00
Georgi Gerganov and GitHub
e8f5082697
server : fix restore for checkpoints with pos_min == 0 ( #21510 )
2026-04-07 15:29:17 +03:00
Georgi Gerganov and GitHub
22fc79134e
ggml : deprecate GGML_OP_ADD1 ( #21363 )
...
* ggml : deprecate GGML_OP_ADD1
* cont : remove tests
* cont : re-enable vulkan check
2026-04-07 15:28:27 +03:00
Georgi Gerganov and GitHub
400ac8e194
convert : set "add bos" == True for Gemma 4 ( #21500 )
...
* convert : set "add bos" == True for Gemma 4
* cont : handle old GGUFs
2026-04-06 13:52:07 +03:00
Georgi Gerganov and GitHub
57ace0d612
chat : avoid including json in chat.h ( #21306 )
2026-04-03 09:07:59 +03:00
Georgi Gerganov and GitHub
39b27f0da0
(revert) kv-cache : do not quantize SWA KV cache ( #21332 )
...
This reverts commit 17193cce34 .
2026-04-03 09:07:01 +03:00
Georgi Gerganov and GitHub
17193cce34
kv-cache : do not quantize SWA KV cache ( #21277 )
2026-04-02 11:54:05 +03:00
Georgi Gerganov
dae2bf41c9
sync : ggml
2026-04-02 10:39:00 +03:00
Georgi Gerganov
bc07d55922
ggml : bump version to 0.9.11 (ggml/1456)
2026-04-02 10:39:00 +03:00
Georgi Gerganov and GitHub
744c0c7310
llama : rotate activations for better quantization ( #21038 )
...
* llama : rotate activations for better quantization
* cont : rotate V more + refactor
* cont : rotate caches separately + support non-power-of-2 head sizes
* cont : simplify
* cont : add reference for V rotation
* cont : refactor
* cont : support context shift
* cont : consolidate
* cont : dedup + allow different types for the rotation matrix
* cont : add env variable to disable rotation
* cont : simplify attn rot kv cache logic + rename env
* cont : pre-compute the Hadamard matrices
2026-04-01 16:58:01 +03:00
Georgi Gerganov
6422036fcb
sync : ggml
2026-04-01 16:03:17 +03:00
Georgi Gerganov
296bc0538b
ggml : bump version to 0.9.10 (ggml/1454)
2026-04-01 16:03:17 +03:00
Georgi Gerganov and GitHub
d43375ff7f
ggml : fix RWKV ops thread assignment ( #21226 )
2026-04-01 11:10:25 +03:00
Georgi Gerganov
9281dd135d
sync : ggml
2026-03-31 14:00:41 +03:00
Georgi Gerganov
0be6c7c9ce
ggml : bump version to 0.9.9 (ggml/1449)
2026-03-31 14:00:41 +03:00
Georgi Gerganov and GitHub
edfb440a2f
server : fix processing of multiple back-to-back mtmd chunks ( #21107 )
2026-03-28 16:27:36 +02:00
Georgi Gerganov and GitHub
3fab96cd04
ci : disable self-hosted mac jobs ( #20985 )
2026-03-25 14:46:40 +02:00
Georgi Gerganov and GitHub
9f102a1407
models : move the token embedding norms to the first layer ( #20943 )
...
* models : move the token embedding norms to the first layer
* cont : fix LLM_TENSOR_CONV1D + fix il indexing
2026-03-24 17:00:30 +02:00
Georgi Gerganov and GitHub
342d6125bc
metal : add FA instantiations for HSK=512, HSV=512 ( #20902 )
2026-03-24 10:03:09 +02:00
Georgi Gerganov and GitHub
f93c09e267
memory : fix seq_id bounds in llama_memory_recurrent::state_read_meta() ( #20887 )
2026-03-23 14:08:46 +02:00
Georgi Gerganov and GitHub
e32d243849
ai : update gh permissions ( #20895 )
2026-03-23 13:21:41 +02:00
Georgi Gerganov and GitHub
4cb7e0bd61
ai : limit runtime of the agent ( #20816 )
2026-03-20 20:31:25 +02:00
Georgi Gerganov and GitHub
b31b30f31d
ai : do not run bash commands in the prompt ( #20810 )
2026-03-20 19:06:33 +02:00
Georgi Gerganov and GitHub
ab9d4c3678
server : improve mtmd ctx checkpoints ( #20726 )
...
* server : improve mtmd ctx checkpoints
* server : fix off-by-one in pos_min_thold
2026-03-20 11:13:12 +02:00
Georgi Gerganov and GitHub
464fd0e71f
ai : update find-related action ( #20790 )
...
* ai : update "related issues" prompt
* cont
* cont
* cont
2026-03-20 10:28:14 +02:00
Georgi Gerganov and GitHub
6c72646a61
ci : improve action for duplicate issue ( #20772 )
...
* ci : show thinking traces of the agent
* cont : increase thinking
* cont : remove agent files
* cont : move the model selection to the provider
2026-03-19 21:11:53 +02:00
Georgi Gerganov and GitHub
900efd531d
ci : clarify gh command for viewing issues ( #20766 )
2026-03-19 18:43:54 +02:00
f071ce67c9
ci : add action for finding duplicate issues ( #20756 )
...
* ci : add action for finding duplicates issues
* cont : gen info
* cont : formatting
* cont : fix
* cont : instructions
* cont : bump checkout action
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-03-19 16:17:37 +02:00
Georgi Gerganov
4efd326e71
sync : ggml
2026-03-18 15:17:28 +02:00
Georgi Gerganov
b08f7322ee
ggml : bump version to 0.9.8 (ggml/1442)
2026-03-18 15:17:28 +02:00
Georgi Gerganov
79187f2fb8
ggml : restore ggml_type_sizef() to aboid major version bump (ggml/1441)
2026-03-18 15:17:28 +02:00
Georgi Gerganov and GitHub
8cc2d81264
server : fix ctx checkpoint invalidation ( #20671 )
2026-03-17 15:21:14 +02:00
Georgi Gerganov and GitHub
45172df4d6
ci : disable AMX jobs ( #20654 )
...
[no ci]
2026-03-16 22:38:59 +02:00
Georgi Gerganov and GitHub
9b342d0a9f
benches : add Nemotron 3 Nano on DGX Spark ( #20652 )
...
[no ci]
2026-03-16 21:50:43 +02:00
Georgi Gerganov
f47a246a08
sync : ggml
2026-03-16 17:22:06 +02:00
Georgi Gerganov
c0ccbd1f86
ggml : try fix arm build (whisper/0)
2026-03-16 17:22:06 +02:00
Georgi Gerganov and GitHub
88915cb55c
server : fix wait in test_cancel_requests() test ( #20601 )
...
* server : fix wait in test_cancel_requests() test
* codeowners : add team for server tests
2026-03-15 20:54:37 +02:00
9cd4ebcfb1
ci : split build.yml + server.yml ( #20546 )
...
* ci : split build.yml
* cont : split server.yml
* cont : reduce paths
* cont : split build-android.yml + update paths
* ci : make msys workflows manual (#20588 )
* ci : make cross-build workflows manual (#20585 )
* cont : fix release paths
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-03-15 15:11:17 +02:00
Georgi Gerganov and GitHub
b30a5fdf37
metal : add FA specialization for HSK = 320, HSV = 256 ( #20549 )
2026-03-14 23:15:47 +02:00
Georgi Gerganov and GitHub
b4768955c4
ci : move self-hosted workflows to separate files ( #20540 )
2026-03-14 23:15:35 +02:00
Georgi Gerganov and GitHub
9f774e45ee
ci : reduce webgpu tests timeout to 900s ( #20538 )
...
[no ci]
2026-03-14 17:08:26 +02:00
e30f1fdf74
graph : remove redundant GDN state transposes ( #20443 )
...
* ggml : transpose fused GDN state access for coalesced memory reads (#20436 )
The fused Gated Delta Net kernel accessed the [S_v, S_v] state matrix
column-wise on row-major storage, causing strided reads (stride S_v =
128 floats = 512 bytes) that waste GPU cache bandwidth. This produced a
39% regression on Qwen3.5-9B (Metal, M4 Max) compared to the unfused
path.
Transpose the state indexing so threads read contiguously:
- Metal: s_ptr[is*S_v] -> s_ptr[is] (stride 1 vs S_v)
- CUDA: curr_state[i*S_v+col] -> curr_state[col*S_v+i] (coalesced)
- CPU: restructured loops for row-wise transposed access
Also add --fused-gdn [on|off|auto] CLI flag (mirrors --flash-attn) so
users can control fused GDN independently of auto-detection.
All GATED_DELTA_NET backend-ops tests pass.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com >
* ggml : use SIMD dot products in CPU GDN kernel, couple AR/chunked fused flags
- Replace scalar inner loops with ggml_vec_dot_f32 for SIMD-optimized
dot products in the CPU fused GDN kernel (delta and attention output)
- Couple fused_gdn_ar and fused_gdn_ch flags in auto-detection: if one
path lacks device support, disable both to prevent state layout mismatch
between transposed (fused) and non-transposed (unfused) formats
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com >
* llama : rever fgdn argument changes
* graph : remove GDN state transposes
* vulkan : adapt
* cuda : remove obsolete smem code
---------
Co-authored-by: Paul Flynn <paul@arkavo.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Oliver Simons <osimons@nvidia.com >
2026-03-13 22:12:54 +02:00
Georgi Gerganov and GitHub
73c9eb8ced
metal : fix l2 norm scale ( #20493 )
2026-03-13 11:43:20 +02:00
Georgi Gerganov and GitHub
57819b8d4b
llama : disable graph reuse with pipeline parallelism ( #20463 )
2026-03-12 21:04:13 +02:00
Georgi Gerganov and GitHub
e4cff0956b
metal : avoid divisions in bin kernel ( #20426 )
...
* metal : avoid modulus in bin kernel when not broadcasting
* metal : fix capture_started flag
2026-03-12 09:42:40 +02:00
d28961d81e
llama : enable chunked fused GDN path ( #20340 )
...
* llama : enable chunked fused GDN path
* models : avoid Q and K repeats when using fused GDA
* cont : fix comment
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
* cont : fix the fix
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
* cont : fix
* metal : add GDN kernel (#20361 )
* metal : add Metal backend for GGML_OP_GATED_DELTA_NET
Add a fused Metal kernel for the gated delta net recurrence op
(#19504 ), enabling GPU-accelerated inference for DeltaNet-based
models (Qwen3.5, etc.) on Apple Silicon.
Supports both GDA (scalar gate) and KDA (per-row gate) modes
with head_size 64 and 128. Unsupported configurations (head_size
32, non-contiguous tensors) gracefully fall back to CPU.
Performance: Qwen3.5-0.8B Q4_K_M on M4 Max
tg128: 170 -> 213 t/s (+25%)
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com >
* metal : validate contiguity of all input tensors in supports_op
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com >
* metal : add algorithm equivalence comment for GDA decay path
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com >
* cont : unslop + optimize
* cont : clean-up
---------
Co-authored-by: Paul Flynn <paul@arkavo.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
* CUDA: AR gated delta net improvements (#20391 )
* Add FastDiv to gated_delta_net_cuda
* Shard columns across warps
This reduces register pressure (avoids spill for S_v = 128) and gives
the warp-scheduler more CTAs to schedule (thus hiding data-access
latencies).
* Remove unneded include in gated_delta_net.cu
* Improve comments
* Apply code-formating
* Make sharding HIP-compatible
1. Use ggml_cuda_get_physical_warp_size() to determine warp size flexibly
2. Add test with partial warp to test sum reduction on CUDA
* Remove fastdiv_s64, as we can treat neqk1 and rq3 as uint32_t
* Rename variables
* Enable GDN also for prefill, move TODO for chunked_GDN
* Actually remove the TODO from 206890897546bd16602c3b79394fd5ea09ef199f
* Get warp size at runtime
warp_size is not known at compile time in hip host code.
* Don't expose ggml_cuda_get_physical_warp_size on host
---------
Co-authored-by: uvos <devnull@uvos.xyz >
* llama : refactor llm_build_delta_net_base API
---------
Co-authored-by: Aman Gupta <amangupta052@gmail.com >
Co-authored-by: Paul Flynn <paul@arkavo.com >
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com >
Co-authored-by: Oliver Simons <osimons@nvidia.com >
Co-authored-by: uvos <devnull@uvos.xyz >
2026-03-11 22:46:40 +02:00
Georgi Gerganov and GitHub
3ca19b0e9f
benches : add nemotron super ( #20420 )
2026-03-11 21:39:40 +02:00
Georgi Gerganov and GitHub
76ea1c1c46
metal : fix capture_compute counter logic ( #20410 )
2026-03-11 18:38:22 +02:00
Georgi Gerganov and GitHub
b541241104
metal : fix q5_k mul_mv register spill ( #20399 )
2026-03-11 16:25:27 +02:00
Georgi Gerganov and GitHub
c363256839
metal : add env var to trigger graph capture ( #20398 )
2026-03-11 16:25:10 +02:00
Georgi Gerganov and GitHub
90b2731894
ggml : bump RPC version ( #20330 )
2026-03-10 21:36:57 +02:00
1274fbee9e
models : fix assert in mamba2 (cont) ( #20335 )
...
* models : fix assert in mamba2 (cont)
* cont : add n_group mod
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-03-10 15:00:08 +02:00
Georgi Gerganov and GitHub
a7b3dee7a5
server : make 2 checkpoints near the end of the prompt ( #20288 )
...
* server : make 2 checkpoints near the end of the prompt
* cont : adjust checkpoints
2026-03-10 14:28:23 +02:00
Georgi Gerganov and GitHub
96cfc4992c
server : fix checkpoints n_tokens calculation ( #20287 )
2026-03-09 16:47:06 +02:00
Georgi Gerganov and GitHub
ed0007aa32
metal : add upscale ( #20284 )
2026-03-09 16:45:11 +02:00
Georgi Gerganov and GitHub
344ee2a38a
server : warn swa-full is not supported for non-SWA models ( #20291 )
2026-03-09 16:44:25 +02:00
Georgi Gerganov and GitHub
d6e1556499
server : fix off-by-1 in server_tokens::size_up_to_pos() ( #20279 )
...
* server : fix off-by-1 in server_tokens::size_up_to_pos()
* cont : fix typo [no ci]
2026-03-09 16:43:38 +02:00
Georgi Gerganov and GitHub
43e1cbd6c1
models : fix assert in mamba2 graph ( #20270 )
2026-03-09 13:15:15 +02:00
Georgi Gerganov and GitHub
107d599952
server : add kill switch when server is stuck ( #20277 )
2026-03-09 10:33:12 +02:00
Georgi Gerganov and GitHub
d417bc43dd
server : do not create checkpoints right after mtmd chunks ( #20232 )
2026-03-08 22:16:46 +02:00
Georgi Gerganov and GitHub
17a4258946
kv-cache : fix M-RoPE checkpoints ( #20132 )
2026-03-06 08:46:51 +02:00
Georgi Gerganov and GitHub
37964f44f9
mtmd : fix padding of n_tokens ( #19930 )
2026-02-26 18:39:49 +02:00
Georgi Gerganov and GitHub
01cd448b8c
server : fix ctx checkpoint restore logic ( #19924 )
2026-02-26 18:20:16 +02:00
Georgi Gerganov and GitHub
99bd67c9b2
kv-cache : fix can_shift() check to take into account M-RoPE ( #19928 )
2026-02-26 18:08:54 +02:00
Georgi Gerganov and GitHub
1ca3d1de15
gguf : avoid too many file size calls ( #19919 )
2026-02-26 12:46:32 +02:00
Georgi Gerganov and GitHub
f20469d919
server : enable multi-modal prompt caching ( #19877 )
2026-02-25 15:15:42 +02:00
d7d826b3c1
server : support multi-modal context checkpoints ( #19849 )
...
* Modify llama-memory-hybrid-iswa.cpp
* Modify llama-memory-recurrent.cpp
* Modify server-common.cpp
* Modify server-common.h
* Modify server-context.cpp
* Modify server-task.h
* Added comment to llama-memory-hybrid-iswa.cpp
* Remove comment from server-context.cpp
* Stylistic fix server-context.cpp
* Fix an issue when seqrm isn't called in server-context.cpp
* cont : alternative impl
* cont : cleanup
* cont : n_tokens -> int64_t
---------
Co-authored-by: timkhronos <timkhronos@gmail.com >
2026-02-25 15:14:27 +02:00
Georgi Gerganov and GitHub
244641955f
models : fix graph splits ( #19866 )
2026-02-25 00:01:13 +02:00
418dea39ce
ggml/gguf : prevent integer overflows ( #19856 )
...
* gguf : prevent integer overflow for ggml_context mem size
* ggml : fix int overflows in ggml_new_object()
* gguf : prevent string exhaustion
* gguf : prevent array elements exhaustion
* ggml : fix negative tensor type oob
* py : assert that alignment is non-zero power of 2
* ggml : check int overflow in ggml_new_tensor_impl and ggml_new_object
* gguf-py : error on duplicate keys when reading
* py : restore tensor_fields
* enforce proper alignment in add_custom_alignment
* gguf : better name
* gguf : fix ctx size for no_alloc == true
* gguf : minor print fix
* ggml : print values when overflow
* ggml : remove deprecated ggml_type_sizef()
* ggml : relax ggml_type asserts to debug-only
* gguf : add mem_size overflow test
* gguf : add file size check for arrays
* ggml : relax asseerts for ggml_get_type_traits()
* flake8 fix
---------
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com >
2026-02-24 20:17:11 +02:00
Georgi Gerganov and GitHub
da348c9dfb
models : fix qwen3.5 beta/gate shapes ( #19730 )
...
* models : fix qwen3.5 beta/gate shapes
* cont : avoid extra reshapes
2026-02-19 15:19:53 +02:00
Georgi Gerganov and GitHub
27326bfce1
models : dedup qwen35 graphs ( #19660 )
...
* models : dedup qwen35 graphs
* cont : add missing sigmoid
2026-02-19 08:17:49 +02:00
ad8207af77
cuda : enable CUDA graphs for MMID 1 <= BS <= 4 ( #19645 )
...
* cuda : enable CUDA graphs for MMID BS <= 4
* cont : add stream capture check
Co-authored-by: Oliver Simons <osimons@nvidia.com >
* cont : add MMVQ_MMID_MAX_BATCH_SIZE
---------
Co-authored-by: Oliver Simons <osimons@nvidia.com >
2026-02-17 12:31:49 +02:00
Georgi Gerganov and GitHub
cc45f2ada6
models : deduplicate delta-net graphs for Qwen family ( #19597 )
...
* models : add llm_build_delta_net_base
* cont : keep qwen35 and qwen35moe graphs intact
* cont : add comments
2026-02-16 14:35:04 +02:00
Georgi Gerganov and GitHub
d5dfc33027
graph : fix KQ mask, lora, cvec reuse checks ( #19644 )
...
* graph : fix KQ mask reuse condition
* cont : dedup KQ mask build and can_reuse
* cont : fix build
* graph : fix adapter check for reuse
2026-02-16 09:21:11 +02:00
Georgi Gerganov
ff4affb4c1
sync : ggml
2026-02-15 22:24:29 +02:00
Georgi Gerganov
55d58599c8
ggml : bump version to 0.9.7 (ggml/1425)
2026-02-15 22:24:29 +02:00