Files
llama_cpp/src
PascalandGitHub f0156d1401 kv-cache: follow the source cache size when sharing cells (#24267)
A fitted target context can end up smaller than the draft default, the
oversized assistant views then overflow the shared K/V tensors and trip
the ggml_view_4d size assert during graph reserve.
2026-06-07 18:33:00 +03:00
..
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-06-07 20:50:54 +08:00
2026-04-03 10:33:03 +02:00