Georgi Gerganov and GitHub
b2c08c9ec4
metal : mark FA blocks ( #16372 )
...
* metal : better unroll in the FA kernels
* metal : index FA blocks
* tests : restore [no ci]
* metal : prevent division by zero in FA kernels
* metal : fix -INF detection logic
2025-10-08 10:57:53 +03:00
Georgi Gerganov and GitHub
7fdd16b432
server : improve context checkpoint logic ( #16440 )
2025-10-08 10:57:29 +03:00
Georgi Gerganov and GitHub
df1b612e29
server : add /v1/health endpoint ( #16461 )
...
* server : add /v1/health endpoint
* cont : update readme
2025-10-07 15:57:14 +03:00
Georgi Gerganov and GitHub
ef4c5b87ea
presets : fix pooling param for embedding models ( #16455 )
2025-10-07 10:32:32 +03:00
Georgi Gerganov and GitHub
0123ff38f5
memory : use sequential equal splits for recurrent modules ( #16442 )
2025-10-07 08:24:17 +03:00
Georgi Gerganov and GitHub
0a319bb75e
metal : add support for non-padded FA KV ( #16148 )
...
* metal : pad K, V and Mask when needed
* cont : simplify
* cuda : add TODO about KV padding requirement
* metal : add comments
* metal : remove mask padding requirement
2025-10-07 08:23:30 +03:00
1d6092fc72
tests : add -INF blocks to the KQ mask in the FA tests ( #16380 )
...
* tests : add -INF blocks to the KQ mask in the FA tests
* cont : bump -INF block size to 64
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
* ggml : prevent division by zero in FA CPU op
---------
Co-authored-by: Jeff Bolz <jbolz@nvidia.com >
2025-10-07 08:22:35 +03:00
Georgi Gerganov and GitHub
8ae32dc9ec
metal : various optimizations + refactoring ( #16446 )
...
* metal : ssm_scan minor opts
* metal : get_rows optimize
* metal : cpy optimize
* metal : ssm_conv opt
* metal : ssm_scan simplify
* metal : ssm_Scan opt
2025-10-07 08:21:40 +03:00
Georgi Gerganov and GitHub
a23b9bdbd3
ggml : fix unaligned access in AMX code ( #16315 )
2025-10-06 16:05:27 +03:00
Georgi Gerganov and GitHub
606a73f531
metal : fix loop bound in ggml_mem_ranges ( #16412 )
2025-10-03 19:18:56 +03:00
Georgi Gerganov and GitHub
bbd32bc038
ci : fix clean-up of old logs ( #16381 )
2025-10-02 10:35:43 +03:00
Georgi Gerganov
075c01567b
ggml : bump version to 0.9.4 (ggml/1363)
2025-09-30 13:53:55 +03:00
Georgi Gerganov and GitHub
35fb82497e
metal : dynamic simdgroups for MV kernels ( #16340 )
...
* metal : dynamic simdgroups for MV kernels
* cont : minor
2025-09-30 11:03:23 +03:00
Georgi Gerganov and GitHub
d72f5f7ba2
ci : add AMD runners and workflows ( #16249 )
...
* ci : add AMD runners and workflows
* ci : move AMD jobs to separate workflow
* cont : fix paths
2025-09-29 17:51:48 +03:00
Georgi Gerganov
2ddd3f2356
sync : ggml
2025-09-29 17:43:58 +03:00
Georgi Gerganov
4d3d455d3c
sync : whisper.cpp (ggml/1359)
...
* ggml : Fix MKL detection by quoting BLAS_INCLUDE_DIRS (whisper/3426)
* sync : whisper.cpp
2025-09-29 17:43:58 +03:00
Georgi Gerganov
b6dff20e2f
ggml : prepare for development of 0.9.2-dev
2025-09-29 17:43:58 +03:00
Georgi Gerganov
2db78c75e4
ggml : bump version to 0.9.1
2025-09-29 17:43:58 +03:00
Georgi Gerganov and GitHub
a4a0aa5ea2
ggml : fix dependencies for ggml_set_rows ( #16318 )
2025-09-29 08:41:28 +03:00
Georgi Gerganov and GitHub
6a2c6145a0
metal : extend mat-mat multiplication support ( #16225 )
...
* metal : support mul_mm with src1->type == GGML_TYPE_F16
* metal : support mul_mm_id with src1->type == GGML_TYPE_F16
[no ci]
* metal : mul_mm support ne00 % 32 != 0
* metal : support mul_mm_id with ne00 % 32 != 0
* cont : remove unnecessary unrolls
* cont : simplify data loading
* metal : optimize mul_mm when output bounds checks are not needed
2025-09-28 09:34:44 +03:00
Georgi Gerganov and GitHub
3b53634fe3
metal : fuse non-sequential nodes ( #16102 )
...
* metal : fuse non-sequential nodes
* cont : add comment
* cont : simplify bounds checks
2025-09-28 09:34:05 +03:00
Georgi Gerganov and GitHub
54dbc37053
metal : report OOM errors ( #16274 )
2025-09-26 14:14:28 +03:00
Georgi Gerganov and GitHub
dfcd53f7ec
metal : fuse NORM + MUL + ADD, support non-multiples of 4 ( #16220 )
...
* metal : fuse NORM + MUL + ADD
* metal : support norms of non-multiple of 4
* cont : fix comment [no ci]
2025-09-25 11:30:16 +03:00
Georgi Gerganov and GitHub
4ea00794b8
metal : relax reorder conditions ( #16216 )
2025-09-25 11:29:42 +03:00
Georgi Gerganov and GitHub
02a6a82ae7
metal : restore im2col perf ( #16219 )
2025-09-25 11:29:08 +03:00
Georgi Gerganov and GitHub
f505bd83ca
ci : disable AMD workflows + update NVIDIA workflows ( #16200 )
...
* ci : disable AMD workflows + update NVIDIA workflows
* cont : fixes
* cont : update nvidia vulkan workflows
2025-09-23 20:41:40 +03:00
Georgi Gerganov and GitHub
0889589dbe
ci : enable Vulkan workflow on Mac ( #16194 )
2025-09-23 13:44:25 +03:00
432cf4304c
codeowners : update + cleanup ( #16174 )
...
---------
Co-authored-by: slaren <slarengh@gmail.com >
2025-09-22 18:20:21 +03:00
Georgi Gerganov and GitHub
4f324a556c
ggml : extend ggml_can_fuse to work with non-sequential nodes ( #16123 )
...
* ggml : extend ggml_can_fuse to work with non-sequential nodes in the graph
* cont : fix wrong bounds check condition
* cont : remove unnecessary overload
2025-09-22 11:12:37 +03:00
Georgi Gerganov and GitHub
a71ae3ba7a
ggml : add ggml_op_is_empty ( #16122 )
...
* ggml : add ggml_op_is_empty
* ggml : move to ggml-impl.h
2025-09-22 11:12:09 +03:00
Georgi Gerganov and GitHub
5c6106a696
contrib : update roles ( #16113 )
...
* contrib : update roles
* contrib : merge PR sections + add link to CI instructions
Updated pull request guidelines for contributors and collaborators, and clarified merging practices for maintainers.
2025-09-22 10:58:02 +03:00
Georgi Gerganov and GitHub
ec65fb52f0
ci : remove vulkaninfo calls ( #16169 )
2025-09-22 10:16:05 +03:00
Georgi Gerganov and GitHub
1d660d2fae
ci : use smaller model ( #16168 )
...
* ci : switch from gemma to qwen3 0.6b
* ci : use smaller model for some tests
2025-09-22 09:11:39 +03:00
Georgi Gerganov and GitHub
4d0a7cbc61
ci : adjust params for less runtime ( #16167 )
...
* ci : adjust params for less runtime
* ci : gate BF16 on some hardware
* ci : move extra tests to Arm runner
2025-09-22 08:31:40 +03:00
Georgi Gerganov and GitHub
da30ab5f86
ci : add label for the RISC-V runner ( #16150 )
2025-09-21 19:00:27 +03:00
Georgi Gerganov and GitHub
28baac9c9f
ci : migrate ggml ci to self-hosted runners ( #16116 )
...
* ci : migrate ggml ci to a self-hosted runners
* ci : add T4 runner
* ci : add instructions for adding self-hosted runners
* ci : disable test-backend-ops from debug builds due to slowness
* ci : add AMD V710 runner (vulkan)
* cont : add ROCM workflow
* ci : switch to qwen3 0.6b model
* cont : fix the context size
2025-09-21 16:50:45 +03:00
Georgi Gerganov
7f766929ca
sync : ggml
2025-09-20 13:02:14 +03:00
Georgi Gerganov and GitHub
703f9e32c4
metal : use function constants for mul_mv_ext kernels ( #16074 )
...
* metal : use function constants for mul_mv_ext kernels
ggml-ci
* metal : remove NW template argument
ggml-ci
* metal : adjust constants
ggml-ci
2025-09-18 16:28:41 +03:00
Georgi Gerganov and GitHub
e58174cecb
llama : bump max seq limit from 64 to 256 ( #15916 )
...
ggml-ci
2025-09-18 12:47:56 +03:00
Georgi Gerganov and GitHub
b213fce89b
metal : improve F32, F16 and BF16 mat-vec multiplication ( #16057 )
...
* metal : improve F32, F16 and BF16 mat-vec multiplication
ggml-ci
* metal : make the NSG a function constant in mul_mv kernels
ggml-ci
2025-09-18 12:33:45 +03:00
Georgi Gerganov and GitHub
f2f28380ea
metal : handle nil cv during pipeline creation ( #16065 )
...
ggml-ci
2025-09-18 10:03:24 +03:00
Georgi Gerganov and GitHub
0320ac5264
metal : refactor + optimize v2 ( #15995 )
...
* metal : improve naming
* metal : refactor device
ggml-ci
* cont : props
ggml-ci
* metal : apply ggml_mem_ranges_t
ggml-ci
* metal : remove GGML_METAL_USE_BF16
ggml-ci
* metal : refactor device buffer
ggml-ci
* cont : fix naming
* metal : sync before destroying the backend
ggml-ci
* metal : refactor context
ggml-ci
* metal : migrate ggml-metal.m to ggml-metal.cpp
ggml-ci
* metal : adjust ops API
ggml-ci
* metal : use C++ to store piplienes
ggml-ci
* metal : migrate ops to separate functions
ggml-ci
* metal : add ggml_metal_library_t
ggml-ci
* metal : improve naming
ggml-ci
* metal : cleanp
ggml-ci
* metal : add support for GGML_OP_LOG
ggml-ci
* metal : fix error handling
ggml-ci
2025-09-17 20:38:12 +03:00
Georgi Gerganov and GitHub
9dcd200d57
metal : remove memory pools ( #15966 )
...
* metal : remove mem pool usage
ggml-ci
* metal : remove mem pool implementation
ggml-ci
* metal : take into account the actual allocated memory of the tensor
ggml-ci
* cont : use ggml_backend_buft_get_alloc_size
ggml-ci
* cont : improve, comments
ggml-ci
* cont : add functions for the extra tensor sizes
* metal : add comments
ggml-ci
* metal : implement .get_alloc_size for the rest of the buffer types
ggml-ci
* metal : remove ggml_metal_heap
ggml-ci
2025-09-14 22:02:32 +03:00
Georgi Gerganov and GitHub
a14bd35014
metal : fix kernel requirements ( #15983 )
...
* metal : fix kernel requirements
ggml-ci
* cont : fix supports_op
* cont : fix supports_op for ARGMAX
2025-09-14 15:33:22 +03:00
Georgi Gerganov and GitHub
55758b00ca
metal : refactor kernel loading ( #15964 )
...
* metal : refactor bin kernels loading
ggml-ci
* metal : refactor rms kernel loading
ggml-ci
* ci : try to add memory leaks check
ggml-ci
* ci : try to enable memory leak detection for Mac
* cont : seems to be working
2025-09-13 16:24:22 +03:00
Georgi Gerganov and GitHub
f161463a54
metal : allow ops to run concurrently ( #15929 )
...
* metal : run graphs ops concurrently
ggml-ci
* cont : add flags for debugging and disabling concurrency
ggml-ci
* cont : refactor and handle fusing
ggml-ci
* cont : simplify - no need to use GPU address
ggml-ci
* cont : prepare mem ranges for reuse + add ggml-metal-common.cpp
ggml-ci
* cont : avoid redundant keywords in cpp [no ci]
* metal : reorder graph for better concurrency
ggml-ci
* metal : fix race on mem pool buffers
ggml-ci
* cont : add env GGML_METAL_GRAPH_OPTIMIZE_DISABLE
ggml-ci
* cont : refactor, optimize, add comments
ggml-ci
* cont : refactor ggml-metal.m
ggml-ci
* minor : update logs [no ci]
2025-09-13 13:54:28 +03:00
Georgi Gerganov and GitHub
84d7b2fca1
metal : fix memory leaks ( #15962 )
...
ggml-ci
2025-09-13 12:45:04 +03:00
Georgi Gerganov and GitHub
f088b6a84f
server : adjust prompt similarity thold + add logs ( #15913 )
...
ggml-ci
2025-09-12 17:02:55 +03:00
Georgi Gerganov and GitHub
0f0a3c2851
metal : make the backend async ( #15906 )
...
* metal : make the backend async
ggml-ci
* cont : add comments, extend op offload, clean up
ggml-ci
* metal : fix batch size for MUL_MAT_ID
* metal : remove deprecated ggml_backend_metal_buffer_from_ptr
* metal : create only metal buffers, no wrapping of host memory
ggml-ci
* metal : restore .alloc_buffer for buffer_from_ptr_type
ggml-ci
* metal : remove broken implementation of GGML_OP_SET
ggml-ci
* metal : clean-up loose ends, ready for tests
ggml-ci
* metal : support both private and shared buffers
ggml-ci
* metal : enable private buffers + add global device queue
* metal : disable host buffer to prevent races
ggml-ci
* metal : avoid extra copy during set_tensor
ggml-ci
* metal : use separate buffer types for shread and private Metal buffers
ggml-ci
* metal : simplify synchronization logic
ggml-ci
* metal : fix build
ggml-ci
* metal : do not implement cpy_tensor
ggml-ci
* metal : separate implementations for shared and private buffers
ggml-ci
2025-09-10 17:52:35 +03:00
c252ce67c4
contrib : add notes about merging PRs ( #15881 )
...
* contrib : add notes about merging PRs
* Update CONTRIBUTING.md
Co-authored-by: Diego Devesa <slarengh@gmail.com >
* Update CONTRIBUTING.md
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
---------
Co-authored-by: Diego Devesa <slarengh@gmail.com >
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2025-09-09 08:42:10 +03:00
Georgi Gerganov and GitHub
b0d52998b9
cuda : fix supports_op condition for get_rows when number of blocks is too large ( #15868 )
...
* cuda : fix supports_op condition for get_rows when src1->ne2 > 1
ggml-ci
* ggml : add comment about ggml_get_rows
ggml-ci
* cuda : add FIXME [no ci]
* cuda : update support condition
ggml-ci
2025-09-08 13:56:51 +03:00
Georgi Gerganov and GitHub
f28d4f4ac9
metal : refactor + optimize ( #15857 )
...
* metal : refactor
ggml-ci
* cont : refactor FA-vec kernel
* cont : print metal library load time
* minor : warn to debug + bettern kernel names
ggml-ci
* metal : optimize mul_mv q8_0
ggml-ci
* metal : simplify FA pipeline creation functions
ggml-ci
* metal : improve naming consistency
* metal : safer function constants offsets
ggml-ci
* metal : comments
ggml-ci
2025-09-08 13:34:56 +03:00
Georgi Gerganov and GitHub
a885dcff11
batched-bench : fix llama_synchronize usage during prompt processing ( #15835 )
...
ggml-ci
2025-09-08 10:27:07 +03:00
Georgi Gerganov and GitHub
663027fd54
context : fix n_outputs during reserve ( #15858 )
...
ggml-ci
2025-09-08 10:26:36 +03:00
Georgi Gerganov and GitHub
cf0e3ba150
model : avoid ggml_cont_3d for fused QKV weights ( #15662 )
...
* model : avoid ggml_cont_3d for fused QKV weights
ggml-ci
* kv-cache : make cpy_k and cpy_v implementation more readable
ggml-ci
* cont : add comments
ggml-ci
* cont : minor fix [no ci]
* cont : one more fix
* cont : clarity
ggml-ci
* kv-cache : require contiguous heads of k_cur and v_cur
ggml-ci
2025-09-08 10:25:33 +03:00
Georgi Gerganov and GitHub
c610b6c11b
kv-cache : fix SWA checks + disable cacheless iSWA ( #15811 )
...
ggml-ci
2025-09-05 10:39:22 +03:00
Georgi Gerganov and GitHub
cdedb70a99
sampling : optimize dist sampler ( #15704 )
...
ggml-ci
2025-09-03 18:16:26 +03:00
e92d53b29e
sampling : optimize samplers by reusing bucket sort ( #15665 )
...
* sampling : optimize sorting using bucket sort in more places
ggml-ci
* sampling : do not sort in dist sampler
ggml-ci
* sampling : avoid heap allocations for sort buffers
ggml-ci
* common : add option to sort sampling candidates by probability
ggml-ci
* sampling : revert the change for preserving sort buffers
* sampling : use std::copy instead of memcpy
* sampling : clarify purpose of partial sort helpers
ggml-ci
* cont : remove wrong comment [no ci]
* common : update comment
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
---------
Co-authored-by: Johannes Gäßler <johannesg@5d6.de >
2025-08-31 20:41:02 +03:00
Georgi Gerganov and GitHub
0d161f021a
server : enable /slots by default and make it secure ( #15630 )
...
* server : enable /slots by default and make it secure
ggml-ci
* server : fix tests to pass `--no-slots` when necessary
* server : extend /props with info about enabled endpoints
2025-08-31 20:11:58 +03:00
Georgi Gerganov and GitHub
4efd5a8316
metal : fix checks for available FA kernels ( #15700 )
...
* metal : fix checks for available FA kernels
ggml-ci
* cont : fix comment [no ci]
2025-08-31 19:43:30 +03:00
Georgi Gerganov and GitHub
c8d0d14e77
kv-cache : fix find_slot to not search for continuous slot ( #15638 )
...
ggml-ci
2025-08-28 17:09:05 +03:00
Georgi Gerganov and GitHub
8a4280ce43
kv-cache : remove LLAMA_SET_ROWS checks ( #15505 )
...
ggml-ci
2025-08-28 12:27:02 +03:00
Georgi Gerganov and GitHub
da54f9f1a2
presets : add qwen3-30B-a3b FIM ( #15616 )
2025-08-27 15:48:07 +03:00
Georgi Gerganov and GitHub
1bded5a3b3
kv-cache : better estimate of n_kv for multi-sequence batches ( #15610 )
...
ggml-ci
2025-08-27 13:55:12 +03:00
Georgi Gerganov and GitHub
0373486dbc
graph : fix assert in memory-less build_attn ( #15590 )
...
ggml-ci
2025-08-26 17:45:17 +03:00
Georgi Gerganov and GitHub
b3964c1e89
metal : optimize FA vec for large sequences and BS <= 8 ( #15566 )
...
* metal : optmize FA vec for large heads and sequences
* metal : adjust small-batch mul mv kernels
ggml-ci
* batched-bench : fix total speed computation
ggml-ci
* cont : add comments
ggml-ci
2025-08-26 14:22:14 +03:00
Georgi Gerganov and GitHub
85cc1ae998
context : print graph stats for memory-less contexts ( #15586 )
...
ggml-ci
2025-08-26 12:47:00 +03:00
Georgi Gerganov and GitHub
1d8d83deaa
metal : improve MUL_MAT_ID ( #15541 )
...
* metal : mul_mm_id remove hdst
* metal : remove mul_mm_id hsrc1
* metal : mul_mm_id simplify + add test
* metal : opt mul_mm_id map0
* metal : optimize mul_mm_id id gathering
* metal : mul/div opt
* metal : optimize mul_mm_id_map0
ggml-ci
2025-08-26 12:46:15 +03:00
Georgi Gerganov and GitHub
6b64f74b55
batched-bench : fix unified KV cache handling + pp timing ( #15562 )
...
* batched-bench : fix unified KV cache handling + pp timing
* cont : run dummy token only with split KV cache
2025-08-25 13:56:43 +03:00
Georgi Gerganov and GitHub
b0ba31f525
metal : add FA kernels for HS=40 ( #15559 )
...
ggml-ci
2025-08-25 10:14:48 +03:00
Georgi Gerganov and GitHub
b730706a49
kv-cache : support layer reuse ( #15504 )
...
* kv-cache : support layer reuse
ggml-ci
* cont : update comments [no ci]
2025-08-24 13:07:07 +03:00
Georgi Gerganov and GitHub
9ebebef62f
llama : remove KV cache defragmentation logic ( #15473 )
...
ggml-ci
2025-08-22 12:22:13 +03:00
Georgi Gerganov and GitHub
cd36b5e5c7
llama : remove deprecated llama_kv_self API ( #15472 )
...
ggml-ci
2025-08-21 19:13:45 +03:00
Georgi Gerganov and GitHub
3f196be84b
graph : remove build_attn_with_sinks overload ( #15469 )
...
ggml-ci
2025-08-21 18:44:45 +03:00
Georgi Gerganov and GitHub
715a6db02c
kv-cache : drop the "unified" prefix ( #15467 )
...
* kv-cache : drop the "unified" prefix
ggml-ci
* cont : fix comment [no ci]
2025-08-21 17:00:33 +03:00
Georgi Gerganov and GitHub
30649cab65
ci : continue file download with wget ( #15471 )
...
ggml-ci
2025-08-21 13:42:55 +03:00
Georgi Gerganov and GitHub
2f37014073
lookahead : add sample command to readme ( #15447 )
...
* lookahead : add sample command to readme
* cont : build-agnostic command
2025-08-20 13:30:46 +03:00
Georgi Gerganov and GitHub
9ef6b0b835
model : add gpt-oss type strings ( #15424 )
2025-08-19 19:58:28 +03:00
Georgi Gerganov and GitHub
d2fcd91cf9
server : disable context shift by default ( #15416 )
...
* server : disable context shift by default
ggml-ci
* server : make scopr of test parameters local
2025-08-19 16:46:37 +03:00
Georgi Gerganov and GitHub
9d262f4bad
server : remove swa_full warning ( #15399 )
2025-08-19 08:45:26 +03:00
Georgi Gerganov and GitHub
f0d3c7405c
batched-bench : use rand tokens ( #15398 )
2025-08-19 08:45:12 +03:00
Georgi Gerganov
6d7f1117e3
codeowners : remove mmv.*
2025-08-18 22:06:44 +03:00
Georgi Gerganov
60212f1ead
sync : ggml
2025-08-18 22:06:44 +03:00
Georgi Gerganov
f0c541d315
scripts : update sync scripts
2025-08-18 22:06:44 +03:00
Georgi Gerganov and GitHub
3007baf201
readme : update hot topics ( #15397 )
2025-08-18 18:11:44 +03:00
Georgi Gerganov and GitHub
5edf1592fd
vulkan : fix out-of-bounds access in argmax kernel ( #15342 )
...
ggml-ci
2025-08-15 16:16:36 +02:00
Georgi Gerganov and GitHub
db3010bd23
vulkan : fix compile warnings on macos ( #15340 )
...
ggml-ci
2025-08-15 15:28:28 +02:00
Georgi Gerganov and GitHub
df36bce667
eval-callback : stop on first NaN ( #15320 )
...
* eval-callback : stop on first NaN
* cont : log error
2025-08-14 22:10:51 +03:00
Georgi Gerganov and GitHub
1a01899b61
readme : update hot topics ( #15315 )
2025-08-14 17:16:03 +03:00
Georgi Gerganov and GitHub
d32e03f449
server : add SWA checkpoints ( #15293 )
...
* server : add SWA checkpoints
ggml-ci
* cont : server clean-up
* server : handle state restore fails
* llama : add extended llama_state_seq_ API
* server : do not make checkpoints if --swa-full
ggml-ci
* llama : remove flags value for NONE
* server : configure number of SWA checkpoints with CLI arg
ggml-ci
* args : fix scope of new argument
2025-08-14 14:59:50 +03:00
Georgi Gerganov
3973163bff
sync : ggml
...
ggml-ci
2025-08-14 14:59:27 +03:00
Georgi Gerganov
8b2483730f
tests : remove unused includes (ggml/0)
2025-08-14 14:59:27 +03:00
Georgi Gerganov and GitHub
00f35d509e
ggml : repack block_iq4_nlx8 ( #14904 )
...
ggml-ci
2025-08-13 11:09:39 +03:00
Georgi Gerganov and GitHub
228f724d9c
kv-cache : fix seq_rm with seq_id == -1 ( #15226 )
...
* kv-cache : fix seq_rm with seq_id == -1
ggml-ci
* cont : iterate over streams
ggml-ci
2025-08-11 13:58:24 +03:00
fd1234cb46
llama : add gpt-oss ( #15091 )
...
* oai moe
* compat with new checkpoint
* add attn sink impl
* add rope scaling yarn
* logits match with latest transformers code
* wip chat template
* rm trailing space
* use ggml_scale_bias
* rm redundant is_swa_all
* convert interleaved gate_up
* graph : fix activation function to match reference (#7 )
* vocab : handle o200k_harmony special tokens
* ggml : add attention sinks support (#1 )
* llama : add attn sinks
* ggml : add attn sinks
* cuda : add attn sinks
* vulkan : add support for sinks in softmax
remove unnecessary return
* ggml : add fused swiglu_oai op (#11 )
* ggml : add fused swiglu_oai op
* Update ggml/src/ggml-cpu/ops.cpp
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
* update CUDA impl
* cont : metal impl
* add vulkan impl
* test-backend-ops : more test cases, clean up
* llama : remove unfused impl
* remove extra lines
---------
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com >
---------
Co-authored-by: slaren <slarengh@gmail.com >
* repack mxfp4 upon conversion
* clean up a bit
* enable thinking
* add quick hack to render only some special tokens
* fix bf16 conversion
* remove vocab hack
* webui ok
* support chat parsing for gpt-oss
* fix webui
* direct mapping mxfp4, FINALLY
* force using mxfp4
* properly use lazy tensor
* ggml : add mxfp4
ggml : use e8m0 conversion instead of powf
Co-authored-by: Diego Devesa <slarengh@gmail.com >
change kvalues_mxfp4 table to match e2m1 (#6 )
metal : remove quantization for now (not used)
cuda : fix disabled CUDA graphs due to ffn moe bias
vulkan : add support for mxfp4
cont : add cm2 dequant
* ggml : add ggml_add_id (#13 )
* ggml : add ggml_add_id
* add cuda impl
* llama : add weight support check for add_id
* perf opt
* add vulkan impl
* rename cuda files
* add metal impl
* allow in-place ggml_add_id
* llama : keep biases on CPU with --cpu-moe
* llama : fix compile error
ggml-ci
* cuda : add fallback for __nv_cvt_e8m0_to_bf16raw
ggml-ci
* cleanup
ggml-ci
* sycl : fix supports_op for MXFP4
ggml-ci
* fix Unknown reasoning format
* ggml-cpu : fix AVX build
ggml-ci
* fix hip build
ggml-ci
* cuda : add mxfp4 dequantization support for cuBLAS
ggml-ci
* ggml-cpu : fix mxfp4 fallback definitions for some architectures
ggml-ci
* cuda : fix version required for __nv_cvt_e8m0_to_bf16raw
---------
Co-authored-by: Xuan Son Nguyen <son@huggingface.co >
Co-authored-by: slaren <slarengh@gmail.com >
2025-08-05 22:10:36 +03:00
Georgi Gerganov and GitHub
be42642581
readme : update hot topics ( #15097 )
2025-08-05 20:19:33 +03:00
Georgi Gerganov and GitHub
a4569c41fd
llama : enable LLAMA_SET_ROWS=1 by default ( #14959 )
...
ggml-ci
2025-08-02 17:14:21 +03:00
Georgi Gerganov and GitHub
15e92fd337
cuda, sycl : fix batched gemm when ne02 == 1 && ne03 > 1 ( #15038 )
...
* cuda, sycl : fix batched gemm when ne02 == 1 && ne03 > 1
ggml-ci
* cont : fix cont types
ggml-ci
* cont : adopt variable names and comment from the other branch
2025-08-02 17:13:05 +03:00
Georgi Gerganov and GitHub
ba42794c9e
graph : fix equal_seq() check ( #14986 )
...
ggml-ci
2025-08-01 06:38:12 +03:00
Georgi Gerganov
e32a4ec60e
sync : ggml
...
ggml-ci
2025-07-30 17:33:11 +03:00