Georgi Gerganov and GitHub
254098a279
common : refactor common_sampler + grammar logic changes ( #17937 )
...
* common : refactor common_sampler + grammar logic changes
* tests : increase max_tokens to get needed response
* batched : fix uninitialized samplers
2025-12-14 10:11:13 +02:00
Georgi Gerganov and GitHub
77ad8542bd
model-conversion : cast logits to float32 ( #18009 )
2025-12-14 08:58:13 +02:00
Georgi Gerganov and GitHub
609a2d0268
models : fix YaRN regression + consolidate logic ( #18006 )
...
* models : fix YaRN regression + consolidate logic
* cont : fix the fix
* cont : remove header
* cont : add header
2025-12-14 08:34:56 +02:00
Georgi Gerganov
a63cbafbbc
ggml : arm repack fix build
2025-12-14 08:33:51 +02:00
Georgi Gerganov
0e59224990
sync : ggml
2025-12-14 08:33:51 +02:00
Georgi Gerganov
71fdcf0616
ggml : arm repack fix build (whisper/0)
2025-12-14 08:33:51 +02:00
Georgi Gerganov and GitHub
3c6391e748
speculative-simple : free batch on exit ( #17985 )
2025-12-13 09:48:34 +02:00
Georgi Gerganov and GitHub
7bed317f53
models : fix the attn_factor for mistral3 graphs + improve consistency ( #17945 )
...
* models : fix the attn_factor for mistral3 graphs
* cont : rework attn_factor correction logic
* cont : make deepseek2 consistent
* cont : add TODO
* cont : special-case DSv2
* cont : revert Mistral 3 Large changes
* cont : fix DS2 to use the original attn_factor
* cont : minor comments
2025-12-12 17:12:40 +02:00
Georgi Gerganov and GitHub
c6f6e4f96a
ggml-alloc : fix reuse-parent logic for misaligned sizes ( #17884 )
2025-12-11 14:30:10 +02:00
Georgi Gerganov and GitHub
d9f8f60618
batch : fix sequence id ownership ( #17915 )
...
* batch : fix sequence id ownage
* cont : reduce allocations
2025-12-11 14:29:47 +02:00
Georgi Gerganov and GitHub
4dff236a52
ggml : remove GGML_KQ_MASK_PAD constant ( #17910 )
...
* ggml : remove GGML_KQ_MASK_PAD constant
* cont : remove comment
2025-12-10 20:53:16 +02:00
Georgi Gerganov and GitHub
6b82eb7883
metal : print node names for debugging ( #17882 )
2025-12-09 15:25:49 +02:00
Georgi Gerganov and GitHub
2bc96931d2
server : make cache_reuse configurable per request ( #17858 )
2025-12-08 12:43:12 +02:00
Georgi Gerganov and GitHub
8e5f4987b1
contrib : stale PRs ( #17803 )
2025-12-06 09:34:18 +02:00
Georgi Gerganov and GitHub
8ce774a102
metal : fix build( #17799 )
...
* metal : fix build
* tests : fix context destruction
2025-12-06 09:33:59 +02:00
Georgi Gerganov and GitHub
8160b38a5f
rpc : fix alloc size logic ( #17116 )
...
* rpc : fix alloc size logic
* rpc : bump version
2025-12-05 19:39:04 +02:00
Georgi Gerganov and GitHub
c41bde6fbd
metal : add residency sets keep-alive heartbeat ( #17766 )
...
* examples : add idle
* metal : attach residency sets to queue
* idle : add link
* idle : adjust intervals
* metal : add residency sets keep-alive heartbeat
* cont : adjust default keep-alive time
2025-12-05 19:38:54 +02:00
Georgi Gerganov and GitHub
0d1324856f
metal : use params per pipeline instance ( #17739 )
2025-12-04 10:34:11 +02:00
Georgi Gerganov and GitHub
a67ef0f47f
llama : fix sanity checks during quantization ( #17721 )
2025-12-04 10:33:42 +02:00
Georgi Gerganov and GitHub
190c4838bd
chat : reserve memory in compute_diffs and improve naming ( #17729 )
2025-12-03 17:22:10 +02:00
Georgi Gerganov and GitHub
3d94e967a1
metal : fix data race in pipeline library ( #17731 )
2025-12-03 14:03:40 +02:00
Georgi Gerganov and GitHub
649495c9d9
metal : add FA head size 48 ( #17619 )
2025-12-01 12:49:53 +02:00
Georgi Gerganov and GitHub
90c72a614a
ggml : extend the GGML_SCHED_NO_REALLOC debug logic of the scheduler ( #17617 )
2025-12-01 12:49:33 +02:00
Georgi Gerganov and GitHub
c386114922
arch : add description about LLM_TENSOR_INFOS ( #17550 )
2025-11-27 16:34:13 +02:00
Georgi Gerganov and GitHub
6783b11fb0
models : fix LFM2 tensors ( #17548 )
2025-11-27 16:04:29 +02:00
Georgi Gerganov and GitHub
583cb83416
ggml : add ggml_top_k ( #17365 )
...
* ggml : add ggml_top_k
* cont : add ggml_argsort_top_k
* metal : add top_k support
* ggml : cleanup
* tests : add virtual err() function for test_case
* ggml : add comments
2025-11-25 15:31:43 +02:00
Georgi Gerganov
2d50b9d8cb
sync : ggml
2025-11-24 15:26:31 +02:00
Georgi Gerganov
2286a360ff
sync : ggml
2025-11-20 14:10:44 +02:00
Georgi Gerganov and GitHub
196f5083ef
common : more accurate sampling timing ( #17382 )
...
* common : more accurate sampling timing
* eval-callback : minor fixes
* cont : add time_meas impl
* cont : fix log msg [no ci]
* cont : fix multiple definitions of time_meas
* llama-cli : exclude chat template init from time measurement
* cont : print percentage of unaccounted time
* cont : do not reset timings
2025-11-20 13:40:10 +02:00
Georgi Gerganov and GitHub
f40a2e5f11
gitignore : be more specific about ignored stuff ( #17354 )
2025-11-18 16:44:53 +02:00
Georgi Gerganov and GitHub
7aaeedc098
metal : support I32 -> I32 copy ( #17317 )
2025-11-17 11:52:00 +02:00
Georgi Gerganov and GitHub
3347e6d904
metal : faster argsort ( #17315 )
...
* metal : faster argsort
* cont : keep data in registers
2025-11-17 11:51:48 +02:00
Georgi Gerganov and GitHub
1a139644a8
metal : add cumsum ( #17305 )
2025-11-17 11:51:13 +02:00
Georgi Gerganov and GitHub
416e7c7f47
metal : remove obosolete asserts ( #17295 )
2025-11-16 09:50:26 +02:00
Georgi Gerganov and GitHub
5b2093becc
server : handle context overflow during decode ( #17267 )
...
* server : handle context overflow during decode
* server : minor refactor
2025-11-16 09:23:37 +02:00
Georgi Gerganov and GitHub
d396b43748
server : fix "can batch with" bug ( #17263 )
2025-11-14 14:03:45 +02:00
Georgi Gerganov and GitHub
45c6ef7307
metal : support argsort for ne00 > 1024 ( #17247 )
...
* metal : refactor argsort
* cont : sort chunks
* cont : merge sorted buckets
* cont : cleanup
2025-11-14 09:36:06 +02:00
Georgi Gerganov and GitHub
2606b0adab
metal : make the FA extra sizes consistent ( #17143 )
2025-11-14 09:13:34 +02:00
Georgi Gerganov and GitHub
2776db6c81
Revert "ggml-cpu: handle 3d tensors in repack mat_mul ( #17030 )" ( #17233 )
...
This reverts commit 1c398dc9ec .
2025-11-13 12:59:37 +02:00
Georgi Gerganov and GitHub
374fe09cdd
ggml : use std::sort in ggml_argsort CPU implementation ( #17211 )
...
* ggml : use std::sort in ggml_argsort CPU implementation
* cont : add missing header
2025-11-12 20:43:38 +02:00
Georgi Gerganov and GitHub
13730c183b
metal : cap threadgroups size of set_rows ( #17146 )
2025-11-10 21:33:35 +02:00
Georgi Gerganov and GitHub
c27efd2bd1
metal : enable tensor API for A19 ( #17087 )
2025-11-10 15:38:42 +02:00
Georgi Gerganov and GitHub
f914544b16
batched-bench : add "separate text gen" mode ( #17103 )
2025-11-10 12:59:29 +02:00
Georgi Gerganov and GitHub
9898b57cbe
editorconfig : ignore benches/ ( #17140 )
...
[no ci]
2025-11-10 12:17:19 +02:00
Georgi Gerganov and GitHub
15274c0c50
benches : add eval results ( #17139 )
...
[no ci]
2025-11-10 10:44:10 +02:00
Georgi Gerganov and GitHub
b8595b16e6
mtmd : fix embedding size for image input ( #17123 )
2025-11-09 18:31:02 +02:00
Georgi Gerganov and GitHub
cb1adf8851
server : handle failures to restore host cache ( #17078 )
...
* server : handle failures to restore host cache
* server : add tests for the prompt cache
2025-11-09 14:27:05 +02:00
Georgi Gerganov and GitHub
ef1d826997
benches : add folder with benchmarks ( #16931 )
...
* benches : add folder with benchmarks
* benches : update dgx-spark bench
2025-11-09 12:53:29 +02:00
Georgi Gerganov and GitHub
0750a59903
metal : retain src and dst buffers during async ops ( #17101 )
2025-11-09 08:28:51 +02:00
Georgi Gerganov and GitHub
7956bb4d7f
bench : cache the llama_context state at computed depth ( #16944 )
...
* bench : cache llama_context state at depth
* cont : handle failures to restore the old state
* cont : print information when the state is being reused
2025-11-07 21:23:11 +02:00
Georgi Gerganov and GitHub
16bcc1259d
kv-cache : pad the cache size to 256 for performance ( #17046 )
...
* kv-cache : pad the size of the small SWA cache for performance
* context : pad the total context to 256
* cont : future-proof the swa pad
* server : adjust test params to new logic
2025-11-07 20:03:25 +02:00
Georgi Gerganov and GitHub
8c0d6bb455
server : print the samplers chain for each request ( #17070 )
2025-11-07 12:24:47 +02:00
Georgi Gerganov and GitHub
5b180c3d60
metal : initial Metal4 tensor API support ( #16634 )
...
* metal : rework mat-mat multiplication
* metal : initial Metal4 support
* cont
* metal : detect tensor support
* cont : better ifdefs
* metal : support tensors in mul_mm_id
* metal : add env for disabling tensor API
* tests : restore
* metal : remove unused constants
* metal : fix check for bfloat tensor support
* cont : handle API incompatibilities
* cont : handle even more incompatibilities
* metal : use tensor API only on M5 and later
2025-11-06 14:45:10 +02:00
Georgi Gerganov and GitHub
b7f9010d24
server : disable checkpoints with mtmd ( #17045 )
2025-11-06 12:09:29 +02:00
Georgi Gerganov and GitHub
13b339bcd9
server : do not default to multiple slots with speculative decoding ( #17017 )
...
* server : do not default to multiple slots with speculative decoding
* cont : fix
2025-11-05 14:32:55 +02:00
Georgi Gerganov
cdabeb2c27
sync : ggml
2025-11-05 10:41:51 +02:00
852ce5180a
ggml : fix conv2d_dw SVE path (ggml/1380)
...
* Fix test-conv2d-dw failure on ARM SVE by using runtime vector length
The ggml_compute_forward_conv_2d_dw_cwhn function was using a hardcoded GGML_F32_EPR (8) for SIMD vectorization, but on ARM SVE the actual vector length varies by hardware. This caused incorrect computation when processing CWHN layout tensors on ARM machines.
Fix by using svcntw() to get the runtime SVE vector length instead of the compile-time constant.
Co-authored-by: ggerganov <1991296+ggerganov@users.noreply.github.com >
* ci : reduce sam score threshold
* ci : update bbox checks for sam test
---------
Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com >
Co-authored-by: ggerganov <1991296+ggerganov@users.noreply.github.com >
2025-11-05 10:41:51 +02:00
Georgi Gerganov and GitHub
66d8eccd42
server : do context shift only while generating ( #17000 )
2025-11-04 19:21:36 +02:00
Georgi Gerganov and GitHub
afd353246d
readme : update hot topics ( #17002 )
2025-11-04 17:21:31 +02:00
Georgi Gerganov and GitHub
48bd26501b
server : add props.model_alias ( #16943 )
...
* server : add props.model_alias
* webui : npm run format
2025-11-03 14:38:23 +01:00
2f966b8ed8
clip : use FA ( #16837 )
...
* clip : use FA
* cont : add warning about unsupported ops
* implement "auto" mode for clip flash attn
* clip : print more detailed op support info during warmup
* cont : remove obsolete comment [no ci]
* improve debugging message
* trailing space
* metal : remove stray return
---------
Co-authored-by: Xuan Son Nguyen <son@huggingface.co >
2025-11-02 21:21:48 +01:00
Georgi Gerganov and GitHub
cd5e3b5754
server : support unified cache across slots ( #16736 )
...
* server : support unified context across slots
* cont : fix speculative decoding initialization
* context : fix n_ctx_per_seq computation
* server : purge slots one by one
* tests : add unified cache server tests
* llama : update per-seq context computation
* test-thread-safety : handle tiny training context of the input model
* server : fix server_tokens clear()
* server : use 4 slots + unified KV by default
* llama : add note about context size queries
* cont : update todos [no ci]
* context : do not cap the size of the context
* tests : adjust parameters to be CI friendlier
* context : add warning
2025-11-02 18:14:04 +02:00
Georgi Gerganov and GitHub
7fd205a8e8
scripts : add script to bench models ( #16894 )
2025-11-02 00:15:31 +02:00
Georgi Gerganov
6d39015a74
sync : ggml
2025-10-31 16:26:28 +02:00
Georgi Gerganov and GitHub
8da3c0e200
batch : fix consistency checks for the input positions ( #16890 )
2025-10-31 13:50:33 +02:00
Georgi Gerganov and GitHub
c22473b580
server : don't print user inputs to console ( #16871 )
2025-10-31 10:54:19 +02:00
b52edd2558
server : remove n_past ( #16818 )
...
* server : remove n_past
* server : replace slot.n_prompt_tokens() with slot.task->n_tokens()
* server : fixes + clean-up
* cont : fix context shift
* server : add server_tokens::pos_next()
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
* server : fix pos_next() usage
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
---------
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
2025-10-30 18:42:57 +02:00
Georgi Gerganov and GitHub
85a7d8677b
memory : remove KV cache size padding ( #16812 )
...
* memory : remove KV cache size padding
* cont : restore padding for n_kv tensor shape
* server : use slot context size instead of training context size
* server : simplify context limit logic
2025-10-28 20:19:44 +02:00
Georgi Gerganov and GitHub
a8ca18b4b8
llama-bench : clarify benchmarked parts of the computation ( #16823 )
2025-10-28 19:41:43 +02:00
Georgi Gerganov and GitHub
17304cbcc1
server : fix img token logs ( #16595 )
2025-10-15 16:53:12 +03:00
Georgi Gerganov and GitHub
554fd578a5
server : fix mtmd checkpoints ( #16591 )
2025-10-15 11:51:27 +02:00
Georgi Gerganov and GitHub
fa882fd2b1
metal : avoid using Metal's gpuAddress property ( #16576 )
...
* metal : avoid using Metal's gpuAddress property
* metal : fix rope kernels buffer check
2025-10-14 20:33:05 +03:00
Georgi Gerganov and GitHub
bc07349a7f
server : dynamic token limit for prompt cache ( #16560 )
...
* server : dynamic token limit for prompt cache
* cont : print estimated token limit
2025-10-14 08:48:50 +03:00
Georgi Gerganov and GitHub
e60f241eac
metal : FA support F32 K and V and head size = 32 ( #16531 )
...
* metal : FA support F32 K and V and head size = 32
* graph : remove obsolete comment [no ci]
2025-10-13 23:07:57 +03:00
Georgi Gerganov and GitHub
e38b7c6e9e
graph : support cacheless embeddings with FA and iSWA ( #16528 )
...
* graph : support cacheless embeddings with FA and iSWA
* cont : deduplicate mask creation
* cont : fix name
2025-10-13 22:42:37 +03:00
Georgi Gerganov and GitHub
c515fc5771
ggml : fix scalar path for computing norm ( #16558 )
2025-10-13 11:22:27 +03:00
Georgi Gerganov and GitHub
4b2dae383d
common : update presets ( #16504 )
...
* presets : add --embd-gemma-default and remove old embedding presets
* presets : add gpt-oss presets
* presets : add vision presets
* cont : remove reasoning overrides [no ci]
* cont : fix batch size for embedding gemma [no ci]
2025-10-12 09:29:13 +03:00
Georgi Gerganov and GitHub
a3cb04744f
metal : fix mul-mm condition + fix mul-mv permuted kernels ( #16494 )
2025-10-11 16:54:10 +03:00
Georgi Gerganov and GitHub
e60f01d941
server : fix division by zero when reporting stats ( #16501 )
2025-10-10 22:15:05 +03:00
Georgi Gerganov and GitHub
81086cd6a3
vocab : mark EOT token for Granite models ( #16499 )
...
* vocab : mark EOT token for Granite models
* sampling : fallback to EOS when EOT is not found
2025-10-10 17:17:31 +03:00
d00cbea63c
server : host-memory prompt caching ( #16391 )
...
* minor : code style
* server : fix prompt similarity calculation
* server : initial host-memory prompt caching
* cont
* server : refactor
* cont
* cont : make the server task of the slot const
* cont : minor [no ci]
* server : cache prompts and checkpoints only for completion tasks
* server : improve prompt caching logic
* cont : fix check for number of cached prompts [no ci]
* server : improve caching logic, add -cram CLI arg
* server : print prompt mismatch info
* cont : better naming [no ci]
* server : improve prompt cache loading logic
* server : add option to debug the slot contents (#16482 )
* server : add option to debug the slot contents
* Update tools/server/server.cpp
---------
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
* server : add option to disable prompt cache
---------
Co-authored-by: Xuan-Son Nguyen <son@huggingface.co >
2025-10-09 18:54:51 +03:00
Georgi Gerganov and GitHub
b2c08c9ec4
metal : mark FA blocks ( #16372 )
...
* metal : better unroll in the FA kernels
* metal : index FA blocks
* tests : restore [no ci]
* metal : prevent division by zero in FA kernels
* metal : fix -INF detection logic
2025-10-08 10:57:53 +03:00
Georgi Gerganov and GitHub
7fdd16b432
server : improve context checkpoint logic ( #16440 )
2025-10-08 10:57:29 +03:00
Georgi Gerganov and GitHub
df1b612e29
server : add /v1/health endpoint ( #16461 )
...
* server : add /v1/health endpoint
* cont : update readme
2025-10-07 15:57:14 +03:00
Georgi Gerganov and GitHub
ef4c5b87ea
presets : fix pooling param for embedding models ( #16455 )
2025-10-07 10:32:32 +03:00
Georgi Gerganov and GitHub
0123ff38f5
memory : use sequential equal splits for recurrent modules ( #16442 )
2025-10-07 08:24:17 +03:00
Georgi Gerganov and GitHub
0a319bb75e
metal : add support for non-padded FA KV ( #16148 )
...
* metal : pad K, V and Mask when needed
* cont : simplify
* cuda : add TODO about KV padding requirement
* metal : add comments
* metal : remove mask padding requirement
2025-10-07 08:23:30 +03:00
1d6092fc72
tests : add -INF blocks to the KQ mask in the FA tests ( #16380 )
...
* tests : add -INF blocks to the KQ mask in the FA tests
* cont : bump -INF block size to 64
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
* ggml : prevent division by zero in FA CPU op
---------
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
2025-10-07 08:22:35 +03:00
Georgi Gerganov and GitHub
8ae32dc9ec
metal : various optimizations + refactoring ( #16446 )
...
* metal : ssm_scan minor opts
* metal : get_rows optimize
* metal : cpy optimize
* metal : ssm_conv opt
* metal : ssm_scan simplify
* metal : ssm_Scan opt
2025-10-07 08:21:40 +03:00
Georgi Gerganov and GitHub
a23b9bdbd3
ggml : fix unaligned access in AMX code ( #16315 )
2025-10-06 16:05:27 +03:00
Georgi Gerganov and GitHub
606a73f531
metal : fix loop bound in ggml_mem_ranges ( #16412 )
2025-10-03 19:18:56 +03:00
Georgi Gerganov and GitHub
bbd32bc038
ci : fix clean-up of old logs ( #16381 )
2025-10-02 10:35:43 +03:00
Georgi Gerganov
075c01567b
ggml : bump version to 0.9.4 (ggml/1363)
2025-09-30 13:53:55 +03:00
Georgi Gerganov and GitHub
35fb82497e
metal : dynamic simdgroups for MV kernels ( #16340 )
...
* metal : dynamic simdgroups for MV kernels
* cont : minor
2025-09-30 11:03:23 +03:00
Georgi Gerganov and GitHub
d72f5f7ba2
ci : add AMD runners and workflows ( #16249 )
...
* ci : add AMD runners and workflows
* ci : move AMD jobs to separate workflow
* cont : fix paths
2025-09-29 17:51:48 +03:00
Georgi Gerganov
2ddd3f2356
sync : ggml
2025-09-29 17:43:58 +03:00
Georgi Gerganov
4d3d455d3c
sync : whisper.cpp (ggml/1359)
...
* ggml : Fix MKL detection by quoting BLAS_INCLUDE_DIRS (whisper/3426)
* sync : whisper.cpp
2025-09-29 17:43:58 +03:00
Georgi Gerganov
b6dff20e2f
ggml : prepare for development of 0.9.2-dev
2025-09-29 17:43:58 +03:00
Georgi Gerganov
2db78c75e4
ggml : bump version to 0.9.1
2025-09-29 17:43:58 +03:00
Georgi Gerganov and GitHub
a4a0aa5ea2
ggml : fix dependencies for ggml_set_rows ( #16318 )
2025-09-29 08:41:28 +03:00