Summary
Compilation of 3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c fails on any ARM CPU that has the i8mm extension when GGML_NATIVE=ON (the default used by setup_env.py). The I2_S fast path uses a variable that only exists when llamafile is enabled, but the file explicitly disables llamafile on i8mm/SVE targets. Independent of the model and of the OS (found on macOS; a Linux ARM host with i8mm, e.g. Graviton3 or Neoverse V1/N2, should hit the same line).
Environment
Error
3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.c:1495:40: error: use of undeclared identifier 'src1_cont'
1495 | const size_t src1_col_stride = src1_cont || src1->type != vec_dot_type ? i2s_row_size : nb11;
| ^~~~~~~~~
1 error generated.
Compile flags for that translation unit (from compile_commands.json): -mcpu=native+dotprod+i8mm+nosve+nosme ... -DGGML_USE_LLAMAFILE.
Root cause
ggml-cpu.c line 48:
#if defined(__ARM_FEATURE_SVE) || defined(__ARM_FEATURE_MATMUL_INT8)
#undef GGML_USE_LLAMAFILE
#endif
With i8mm enabled (-mcpu=native+...+i8mm, so __ARM_FEATURE_MATMUL_INT8 is defined), the llamafile block in ggml_compute_forward_mul_mat is compiled out. src1_cont is declared only inside that block (line 1361, #if GGML_USE_LLAMAFILE). The I2_S fast path added at lines 1491-1495 uses src1_cont unconditionally, so it only compiles when llamafile happens to be on: x86, and ARM without i8mm (e.g. Apple M1). Every ARM CPU with i8mm fails.
Suggested fix
Compute the flag locally where it is needed, e.g. in the I2_S block:
const bool src1_cont = ggml_is_contiguous(src1);
or hoist the existing declaration out of the #if GGML_USE_LLAMAFILE region.
Workarounds
-DCMAKE_C_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 -DCMAKE_CXX_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 (same mechanism ggml itself uses when a feature check fails), or
-DGGML_NATIVE=OFF -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16 (build without i8mm).
Both make the build complete. Note that I2_S inference on the resulting arm64 binary is still wrong (garbage output even with the official BitNet-b1.58-2B-4T-gguf), which is the separate, already reported #585 / #600.
Summary
Compilation of
3rdparty/llama.cpp/ggml/src/ggml-cpu/ggml-cpu.cfails on any ARM CPU that has the i8mm extension whenGGML_NATIVE=ON(the default used bysetup_env.py). The I2_S fast path uses a variable that only exists when llamafile is enabled, but the file explicitly disables llamafile on i8mm/SVE targets. Independent of the model and of the OS (found on macOS; a Linux ARM host with i8mm, e.g. Graviton3 or Neoverse V1/N2, should hit the same line).Environment
microsoft/BitNetat0b341e5(currentmain), submodule3rdparty/llama.cppat390c3077(branchrelease-bitnet-embedding-0.6b-270m)setup_env.pyconfiguration (-DBITNET_ARM_TL1=OFF), plus-DCMAKE_SHARED_LINKER_FLAGS=-Wl,-undefined,dynamic_lookupto get past the macOS link error of setup_env.py on MacBook M2 failed #611 / [Windows] Linker error "undefined symbol: quantize_i2_s" when building with -q i2_s or -q tl2 on Windows 11 #595 (otherwise the build stops earlier and this error is never reached)Error
Compile flags for that translation unit (from
compile_commands.json):-mcpu=native+dotprod+i8mm+nosve+nosme ... -DGGML_USE_LLAMAFILE.Root cause
ggml-cpu.cline 48:With i8mm enabled (
-mcpu=native+...+i8mm, so__ARM_FEATURE_MATMUL_INT8is defined), the llamafile block inggml_compute_forward_mul_matis compiled out.src1_contis declared only inside that block (line 1361,#if GGML_USE_LLAMAFILE). The I2_S fast path added at lines 1491-1495 usessrc1_contunconditionally, so it only compiles when llamafile happens to be on: x86, and ARM without i8mm (e.g. Apple M1). Every ARM CPU with i8mm fails.Suggested fix
Compute the flag locally where it is needed, e.g. in the I2_S block:
or hoist the existing declaration out of the
#if GGML_USE_LLAMAFILEregion.Workarounds
-DCMAKE_C_FLAGS=-U__ARM_FEATURE_MATMUL_INT8 -DCMAKE_CXX_FLAGS=-U__ARM_FEATURE_MATMUL_INT8(same mechanism ggml itself uses when a feature check fails), or-DGGML_NATIVE=OFF -DGGML_CPU_ARM_ARCH=armv8.2-a+dotprod+fp16(build without i8mm).Both make the build complete. Note that I2_S inference on the resulting arm64 binary is still wrong (garbage output even with the official
BitNet-b1.58-2B-4T-gguf), which is the separate, already reported #585 / #600.