  
|
3c7450cee1
|
ggml-cpu: extend RVV quantization vec dot to higher VLENs (#22754)
* ggml-cpu: add rvv 512b,1024b impls for iq4_xs
* ggml-cpu: refactor; add rvv 512b, 1024b impls for q6_K, i-quants
* ggml-cpu: refactor; add 512 and 1024 implementations of tq3_s, iq3_xxs, iq2_s, iq2_xs, iq2_xxs
improve iq2_xs impl for rvv 256
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
---------
Co-authored-by: taimur-10x <taimur.ahmad@10xengineers.ai>
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
|
2026-06-04 08:03:40 +03:00 |
|
  
|
1e796eb41f
|
ggml-cpu: add 128-bit RVV implementation for Quantization Vector Dot (#20633)
* ggml-cpu: add 128-bit impls for i-quants, ternary quants
* ggml-cpu: add 128-bit impls for iq2_xs, iq3_s, iq3_xxs, tq2_0
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
* ggml-cpu: refactor; add rvv checks
---------
Co-authored-by: taimur-10x <taimur.ahmad@10xengineers.ai>
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
|
2026-04-16 11:15:15 +03:00 |
|
   
|
fbaa95bc29
|
ggml-cpu: add RVV vec dot kernels for quantization types (#18859)
* ggml-cpu: add rvv quantize_row_q8_K kernel
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
* ggml-cpu: add rvv vec_dot for iq4_nl, mxfp4, iq2_xxs
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
* ggml-cpu: add rvv vec_dot for iq4_xs, refactor
* ggml-cpu: remove ifunc for rvv vec dot
* ggml-cpu: add vec_dot for iq2_xs, iq3_xxs
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
* ggml-cpu: refactor quants.c
---------
Co-authored-by: taimur-10x <taimur.ahmad@10xengineers.ai>
Co-authored-by: Rehan Qasim <rehan.qasim@10xengineers.ai>
Co-authored-by: Rehan Qasim <rehanbhatti0317@gmail.com>
|
2026-03-13 17:36:04 +02:00 |
|