Skip to content

Navigation Menu

Sign in
Appearance settings

Search code, repositories, users, issues, pull requests...

Provide feedback

We read every piece of feedback, and take your input very seriously.

Saved searches

Use saved searches to filter your results more quickly

Appearance settings

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

ui: rendering performance follow-uo server/ui
#26097 opened Jul 24, 2026 by allozaur Contributor Loading…
docs: use ROCM_PATH instead of HIP_PATH in linux HIP build command (#26060) documentation Improvements or additions to documentation
#26096 opened Jul 24, 2026 by amd-hasnasir Loading…
Add LlamaNet to the list of tools in the README. documentation Improvements or additions to documentation
#26095 opened Jul 24, 2026 by unixguru2k Loading…
opencl: fix fused RMS norm mul view offset ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend
#26085 opened Jul 24, 2026 by happyyzy Contributor Loading…
metal: fix memory leak if model is freed without any GPU operations Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26082 opened Jul 24, 2026 by nikwen Contributor Loading…
llama: add default load-mode auto, which avoids mmap on iGPUs AMD ZenDNN Issues related to the AMD ZenDNN backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) Ascend NPU issues specific to Ascend NPUs CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning Hexagon IBM zDNN issues specific to IBM zDNN Accelerator OpenCL Issues specific to the OpenCL backend OpenVINO SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language Vulkan Issues specific to the Vulkan backend WebGPU
#26081 opened Jul 24, 2026 by 0cc4m Contributor Loading…
CUDA: runtime GGML_CUDA_MMVQ_MAX to tune the mvq->MMQ decode crossover CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26079 opened Jul 24, 2026 by praneshgo Contributor Draft
kleidiai: Update KleidiAI Documentation documentation Improvements or additions to documentation ggml changes relating to the ggml tensor library for machine learning
#26078 opened Jul 24, 2026 by JonathanC-ARM Draft
kleidiai: Rework KleidiAI Build System/Integration ggml changes relating to the ggml tensor library for machine learning
#26077 opened Jul 24, 2026 by JonathanC-ARM Draft
kleidiai: Add runtime feature detection mechanism for aarch64/kleidiai ggml changes relating to the ggml tensor library for machine learning
#26076 opened Jul 24, 2026 by JonathanC-ARM Loading…
ggml : handle graph buffer reservation failure ggml changes relating to the ggml tensor library for machine learning
#26070 opened Jul 24, 2026 by FaiChou Loading…
ggml-cpu: Enable tiled gemm for BF16 and FP16 ggml changes relating to the ggml tensor library for machine learning
#26068 opened Jul 24, 2026 by shalinib-ibm Contributor Loading…
vendor : update cpp-httplib to 0.51.0 vendor
#26067 opened Jul 24, 2026 by angt Member Loading…
server: support MCP stdio server vendor
#26062 opened Jul 24, 2026 by ngxson Collaborator Loading…
4 tasks done
metal : add support for GGML_OP_REPEAT_BACK Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning
#26057 opened Jul 24, 2026 by wuisabel-gif Loading…
1 task done
tests: synchronize save-load-state generation testing Everything test related
#26056 opened Jul 24, 2026 by helanfxz Contributor Loading…
CUDA: Optimize prefil via fuse of w_s scale in epilogue MMQ for nvfp4 checkpoints CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26048 opened Jul 23, 2026 by kmorennv Draft
ggml: fix backend split scheduler race condition ggml changes relating to the ggml tensor library for machine learning
#26040 opened Jul 23, 2026 by 0cc4m Contributor Loading…
server: add optional repetition detection documentation Improvements or additions to documentation server testing Everything test related
#26039 opened Jul 23, 2026 by seryogakovalyov Contributor Loading…
Feature: Add p-less sampling testing Everything test related
#26035 opened Jul 23, 2026 by BII-wushuang Loading…
metal: implement soft max backward operation in Metal backend Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning
#26033 opened Jul 23, 2026 by kunwar-vikrant Loading…
CUDA: support non-contiguous rows in L2_NORM CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26026 opened Jul 23, 2026 by MagicalFlames Loading…
sycl: fuse RMS_NORM + MUL ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26015 opened Jul 22, 2026 by Titaniumtown Contributor Draft
Windows unbuffered model load
#26014 opened Jul 22, 2026 by JTischbein Contributor Loading…
ProTip! Type g i on any issue or pull request to go back to the issue listing page.
Morty Proxy This is a proxified and sanitized view of the page, visit original site.