Backlinks: Local AI
Created: 2026-08-10 06:55
Last edited: 2026-08-10 06:55
llama.cpp
Setup
Built from source with Vulkan support on Linux:
https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#for-linux-users
LunarG Vulkan SDK
- download latest SDK
- unpack
$ tar -xf vulkan-sdk.tar.gz
- install runtime dependencies
- `# apt install libxcb-xinput0 libxcb-xinerama0 libxcb-cursor-dev
- source
$ source <vulkan_sdk_dir>/setup-env.sh
llama.cpp
- install dependencies
# apt install libvulkan-dev glslc spirv-headers
- confirm Vuklan
$ vulkaninfo
- build
$ cmake -B build -DGGML_VULKAN=1$ cmake --build build --config Release
Get models
- curated selection of open models
- more
Use
https://github.com/ggml-org/llama.cpp/blob/master/tools/server/README.md#common-params
- check backend
$ build/bin/llama-server --list-devices
- start server
$ build/bin/llama-server --gpu-layers 'all' --model models/Ministral-3-8B-Reasoning-2512-Q8_0.gguf
- access
- open 127.0.0.1:8080