diff --git a/CLAUDE.md b/CLAUDE.md index 2e559254..671acb6c 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI. -Current llama.cpp pinned version: **b10903** +Current llama.cpp pinned version: **b10905** ## Upgrading CUDA Version @@ -502,7 +502,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi ships no UI): ```bash # needs node/npm + network for the asset build; the embed step is plain cmake -P -git clone --depth 1 --branch b10903 https://github.com/ggml-org/llama.cpp /tmp/lc +git clone --depth 1 --branch b10905 https://github.com/ggml-org/llama.cpp /tmp/lc ( cd /tmp/lc/tools/ui && npm ci && npm run build ) mkdir -p webui-generated /tmp/ui-gen cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \ @@ -542,7 +542,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend: - `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored as the repo secret **`DEPOT_TOKEN`**. -Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10903`), the +Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b10905`), the ~280 upstream object files are byte-identical every run, so a warm cache recompiles only the *changed* files. Depot's cache is **shared across all branches** (unlike GitHub's per-branch `actions/cache`), so every branch builds incrementally; a `b` version bump @@ -1536,7 +1536,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson" #### Upstream source location (in CMake build tree) -llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10903`. +llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b10905`. **GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.17.0`), used solely by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the diff --git a/README.md b/README.md index d5878636..0f09078b 100644 --- a/README.md +++ b/README.md @@ -11,7 +11,7 @@ **Build:** ![Java 8+](https://img.shields.io/badge/Java-8%2B-informational) ![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey) -[![llama.cpp b10903](https://img.shields.io/badge/llama.cpp-%23b10903-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10903) +[![llama.cpp b10905](https://img.shields.io/badge/llama.cpp-%23b10905-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b10905) [![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/) ![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162) [![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev) diff --git a/docs/history/llama-cpp-breaking-changes.md b/docs/history/llama-cpp-breaking-changes.md index 7b3e0a99..3c5fb52c 100644 --- a/docs/history/llama-cpp-breaking-changes.md +++ b/docs/history/llama-cpp-breaking-changes.md @@ -708,3 +708,5 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r | b10883–b10902 | patches + upstream verification | **All nine patches still apply, and `0001`, `0010` and `0012` are all still required.** Only one patch target moved in the range — `src/llama-model.{cpp,h}` (`0012`) — and its diff is the two additive enumerators above, nowhere near `load_tensors`. The other eight targets (`common/arg.{cpp,h}`, `common/peg-parser.cpp`, every `tools/server/*`, `tests/CMakeLists.txt`, `vendor/*`) are **byte-unchanged**, verified by diffing those paths explicitly rather than inferred from the aggregate. All three standing drop-checks were run against the pristine tag, because the fail-loud applier detects "does not apply" but never "upstream already fixed this": **`0001`** — `common_params_parse_main` appears **0 times** in `b10902:common/arg.h` and the `#ifdef _WIN32` count-guarded `argv = utf8.ptrs.data()` override is still at `common/arg.cpp:1282`; **`0010`** — `b10902:tools/server/server-context.cpp:4554` still emits `{"vocab_type", meta.model_vocab_type}` uncast, so the `common_json` enum-to-bool trap is still live; **`0012`** — `b10902:src/llama-model.cpp:1491` still normalises with a bare `splits[i] /= split_sum;` and carries no `split_sum == 0` guard. Verified for real: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp written at head `df03399b885831b2a1603b3abb0d8c156808e363` (= `b10902`) with **all nine SHA-256 lines**; full `cmake --build --config Release` clean; `ctest` **537/537**, including the seven `LlamaModelSplits.*` cases that are the only place `0012`'s two extracted functions are linked in CI, and the six `test_wire_contracts.cpp` cases whose configure-time `OAI_LAYER` reader sweep re-ran against b10902's sources (138 CLI / 57 request / 15 trainer names extracted, unchanged). `nm -D` on the fresh `libjllama.so`: **40** `Java_*` exports, **0** C++-mangled ones. `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped**, including `nativeBuildInfoMatchesPinnedVersionConstant` — the end-to-end proof that the four pin sites and the linked binary agree. Run with `clean`: `LLAMA_CPP_VERSION` is a compile-time constant javac inlines into the test class, and Maven's incremental compilation cannot see that dependency. | | b10902–b10903 | `ggml/src/ggml-vulkan/vulkan-shaders/argsort.comp` + `argsort_large.comp` — and nothing else. One commit (**#28705**, "vulkan: fix data race and OOB access in argsort(large)"), 2 files, 21 insertions, 12 deletions, **3 KiB**. | **The smallest bump in this file's history, and a genuinely empty review surface.** Upstream fixes two defects in the Vulkan bitonic argsort. The *OOB read*: the initialising write now leaves `value.y` at 0 for a padded column instead of reading `data_a[row_offset + col]` past `p.ncols` (both shaders). The *data race*: the compare-exchange body is hoisted **inside** the `ixj > col` guard, so only one thread of each pair touches the two shared-memory slots — previously both threads ran the read-modify-write and only the final store was guarded, i.e. the partner thread read `dst_row[idx_0]`/`dst_row[idx_1]` concurrently with its own unsynchronised write. No C++, no header, no build-system file, no Python pin. Every row of the API-compatibility table is **vacuously satisfied** and the three mechanical `tools/server/` contract greps have **no input** — for the second bump running. Well under the 100 KiB chunking threshold, so no chunking question arises. **Where it does land:** GLSL under `ggml/src/ggml-vulkan/vulkan-shaders/` is compiled by `glslc` at build time and embedded, so the change reaches exactly the three Vulkan classifier artifacts (`vulkan-linux-x86-64`, `vulkan-linux-aarch64`, `vulkan-windows-x86-64`). All three are **build-only** jobs on GPU-less runners, so CI proves the shaders still compile, not that the race is fixed — that needs real Vulkan hardware, which no job here has. The default JAR and every non-Vulkan classifier are bit-for-bit unaffected by this range. | | b10902–b10903 | patches + upstream verification | **All nine patches apply untouched, and the intersection with `patches/` is empty by inspection rather than by aggregate:** the two changed files are Vulkan shaders, which no patch in this repo touches. The three standing drop-checks were nevertheless run against the pristine tag rather than waved through on that basis — the fail-loud applier detects "does not apply" but never "upstream already fixed this", and a drop-check firing is a reason to **delete** a patch, which no amount of "the diff is small" substitutes for. **`0001`** — `common_params_parse_main` appears **0 times** in `b10903:common/arg.h` and the `#ifdef _WIN32` count-guarded `argv = utf8.ptrs.data()` override is still at `common/arg.cpp:1282`; **`0010`** — `b10903:tools/server/server-context.cpp:4554` still emits `{"vocab_type", meta.model_vocab_type}` uncast, so the `common_json` enum-to-bool trap is still live; **`0012`** — `b10903:src/llama-model.cpp:1491` still normalises with a bare `splits[i] /= split_sum;` and carries no `split_sum == 0` guard. All three line numbers are **identical to b10902**, which is what a byte-unchanged file looks like. Verified for real: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp written at head `481c65f091f74c5e7089dd0a3a1cc6b50cced31e` (= `b10903`) with **all nine SHA-256 lines**; full `cmake --build --config Release` clean, zero errors; `ctest` **537/537**; wire-name extraction unchanged at **138 CLI / 57 request / 15 trainer** names (the configure-time `OAI_LAYER` reader sweep re-ran against b10903's sources). `nm -D` on the fresh `libjllama.so`: **40** `Java_*` exports, **0** C++-mangled ones. `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped**, including `nativeBuildInfoMatchesPinnedVersionConstant` — run with `clean` because javac inlines `LLAMA_CPP_VERSION` into the test class and Maven's incremental compile cannot see that dependency. Full `mvn test`: **1755 run, 0 failures, 0 errors** (269 skipped — the model-gated classes, no GGUF in this sandbox). SpotBugs **0** findings; `spotless:check` clean; `javadoc:jar` BUILD SUCCESS. | +| b10903–b10905 | `ggml/src/ggml-cuda/fattn-common.cuh`, `fattn-mma-f16.cuh`, `fattn.cu` (**#28102**, Flash-Attention tuning for `gfx1201` — AMD RDNA4 via HIP), `tests/test-backend-ops.cpp` (a case for it), and `.github/workflows/server-sanitize.yml` (**#28708**, upstream keys its sanitizer cache per matrix entry). 2 commits, 5 files, 56 insertions, 12 deletions, **8 KiB**. | **Nothing on the review surface, and nothing this project compiles differently.** Zero files under `common/`, `include/`, `tools/server/`, `tools/mtmd/` or `src/`, so every row of the API-compatibility table is vacuously satisfied and the three mechanical `tools/server/` contract greps have **no input** — third bump running. Far under the 100 KiB threshold, so no chunking question. **Where it lands:** the three `ggml-cuda` files are compiled by the GPU classifier jobs that use that backend — `cuda13-linux-x86-64` / `cuda13-windows-x86-64` and the HIP ones (`rocm-linux-x86-64` / `rocm-windows-x86-64`, which is where `gfx1201` is actually relevant). All are **build-only** jobs on GPU-less runners, so CI proves they still compile, not that the tuning helps. The two remaining files reach nothing here at all: `tests/test-backend-ops.cpp` is never compiled (a FetchContent subproject sets `LLAMA_BUILD_TESTS=OFF`), and `server-sanitize.yml` is upstream's own CI. The default JAR and every CPU classifier are bit-for-bit unaffected. | +| b10903–b10905 | patches + upstream verification | **All ten patches apply untouched, and every patch-target file is byte-unchanged in the range** — verified by diffing those paths explicitly (`common/arg.{cpp,h}`, `common/peg-parser.cpp`, all of `tools/server/`, `src/llama-model.{cpp,h}`, `tests/CMakeLists.txt`, `ggml/src/ggml-cpu/arch/s390/`), not inferred from the aggregate. **This is the first bump with four standing drop-checks rather than three**, because `0013` joined the set; all four were run against the pristine tag, since the fail-loud applier detects "does not apply" but never "upstream already fixed this". **`0001`** — `common_params_parse_main` appears **0 times** in `b10905:common/arg.h`, override still at `common/arg.cpp:1282`. **`0010`** — `b10905:tools/server/server-context.cpp:4554` still emits `{"vocab_type", meta.model_vocab_type}` uncast. **`0012`** — `b10905:src/llama-model.cpp:1491` still normalises with a bare `splits[i] /= split_sum;` and no zero-sum guard. **`0013`** — `b10905`'s `ggml/src/ggml-cpu/arch/s390/repack.cpp` still leaves `vxe_dot_acc` / `vxe_splat_granule` / `vxe_fold` at file scope (lines 73/77/83) between the guarded blocks at 28–70 and 100–155, so the non-VXE s390x build still needs the patch. (Upstream `master` at the time of this bump, `a2878d30d`, carries that file byte-identical to b10905 — the defect is live there too, and was reported on the PR that introduced it.) Verified for real: `rm -rf build` then `cmake -B build -DBUILD_TESTING=ON` through the real `FetchContent` path, configure clean, stamp at head `16378d93f94012d4228c8c7683adce3f286aee5d` (= `b10905`) with **all ten** SHA-256 lines; full `cmake --build --config Release` clean; `ctest` **537/537**; extraction unchanged at **138 CLI / 57 request / 15 trainer** names. **`0013` re-verified with the real cross toolchain** — it has no runnable guard beyond the s390x CI job, so the bump routine now includes it: `s390x-linux-gnu-g++` compiles the applier's `repack.cpp` clean both with the job's own (scalar) flags and with `-mvx -mzvector -march=z15`. `nm -D`: **40** `Java_*` exports, **0** mangled. `mvn -pl llama clean test -Dtest=NativeLibraryLoadSmokeTest` **4/4, 0 skipped**; full `mvn test` **1755 run, 0 failures, 0 errors** (269 skipped — the model-gated classes). SpotBugs **0**; `spotless:check` clean (243 files); `javadoc:jar` BUILD SUCCESS. | diff --git a/llama/CMakeLists.txt b/llama/CMakeLists.txt index e6d8b382..5fc4c5dd 100644 --- a/llama/CMakeLists.txt +++ b/llama/CMakeLists.txt @@ -173,7 +173,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE) FetchContent_Declare( llama.cpp GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git - GIT_TAG b10903 + GIT_TAG b10905 PATCH_COMMAND ${CMAKE_COMMAND} -DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches -DLLAMA_SRC= diff --git a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java index e86bf71c..d17cb5d5 100644 --- a/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java +++ b/llama/src/main/java/net/ladenthin/llama/value/LlamaCppVersion.java @@ -10,13 +10,13 @@ * library was compiled against, exposed as a compile-time constant so callers can render a badge or * emit a startup log line without loading the native library. * - *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10903"}) that mirrors the + *

{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b10905"}) that mirrors the * {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is * absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a * lightweight version badge in Android or other UIs.

* *

For the authoritative value that is baked into the native binary — the build number - * plus the resolved upstream commit, e.g. {@code "b10903-"} — call + * plus the resolved upstream commit, e.g. {@code "b10905-"} — call * {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own * {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires * the native library to be loaded).

@@ -24,14 +24,14 @@ public final class LlamaCppVersion { /** - * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10903"}. + * The pinned llama.cpp release tag this library was built against, e.g. {@code "b10905"}. * *

Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the * "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the * compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the * value actually linked into the native binary.

*/ - public static final String LLAMA_CPP_VERSION = "b10903"; + public static final String LLAMA_CPP_VERSION = "b10905"; // Constants holder — not instantiable. private LlamaCppVersion() {}