 Jiahao LiandGitHub
|
f56e1baec3
|
metal : alibi for arbitrary number of heads (#3426)
|
2023-10-03 19:55:21 +03:00 |
|
 Jiahao LiandGitHub
|
35195689cd
|
2x faster (rms) norm cuda kernels (3.7% e2e improvement) (#2985)
* 2x faster (rms) norm cuda kernels
* Fix code style
|
2023-09-04 08:53:30 +02:00 |
|
 Jiahao LiandGitHub
|
800c9635b4
|
Fix CUDA softmax by subtracting max value before exp (#2665)
|
2023-08-22 20:27:06 +02:00 |
|
 Jiahao LiandGitHub
|
875086bdb9
|
ggml : relax contiguous constraints in activation function (#2371)
|
2023-07-25 15:58:32 +03:00 |
|
 Jiahao LiandGitHub
|
83a00ce69b
|
metal : support bcast add & dup & cont op (#2323)
|
2023-07-23 14:00:37 +03:00 |
|
 Jiahao LiandGitHub
|
7568d1a2b2
|
Support dup & cont ops on CUDA (#2242)
|
2023-07-17 20:39:29 +03:00 |
|
 
|
206e01de11
|
cuda : support broadcast add & mul (#2192)
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
|
2023-07-14 21:38:24 +03:00 |
|