Skip to content

batched-bench : fix unified KV cache handling + pp timing#15562

Merged
ggerganov merged 2 commits into
masterfrom
gg/batched-bench-pps
Aug 25, 2025
Merged

batched-bench : fix unified KV cache handling + pp timing#15562
ggerganov merged 2 commits into
masterfrom
gg/batched-bench-pps

Conversation

@ggerganov

Copy link
Copy Markdown
Member
  • Take into account split KV cache when computing N_KV
  • Do not measure KV cache copying when -pps is enabled

@ggerganov
ggerganov merged commit 6b64f74 into master Aug 25, 2025
47 of 48 checks passed
@ggerganov
ggerganov deleted the gg/batched-bench-pps branch August 25, 2025 10:56
Minh141120 pushed a commit to janhq/llama.cpp that referenced this pull request Aug 26, 2025
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
Minh141120 pushed a commit to janhq/llama.cpp that referenced this pull request Aug 27, 2025
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
blime4 referenced this pull request in blime4/llama.cpp Feb 5, 2026
* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
Seunghhon pushed a commit to Seunghhon/llama.cpp that referenced this pull request Apr 26, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
ljubomirj pushed a commit to ljubomirj/llama.cpp that referenced this pull request May 6, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
my-other-github-account pushed a commit to my-other-github-account/llama.cpp that referenced this pull request May 15, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
phibya pushed a commit to ziee-ai/llama.cpp that referenced this pull request May 29, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
fewtarius pushed a commit to fewtarius/CachyLLama that referenced this pull request May 30, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
fukuro-kun pushed a commit to fukuro-kun/fukuro-llama-cpp-turboquant that referenced this pull request Jul 5, 2026
…5562)

* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
MrLordCat referenced this pull request in MrLordCat/llama.cpp-with-GUI Jul 16, 2026
* batched-bench : fix unified KV cache handling + pp timing

* cont : run dummy token only with split KV cache
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant