Running on Apple Silicon
- Overview
- GPU, Metal and the Neural Engine
- 1. Prerequisites
- 2. Checking
llama-server - 3. The backend and the model
- 4.
gko backend serve,status,stop - 5.
gko doctor - Troubleshooting
Overview
gko needs nothing Apple-specific. On a Mac it drives the same thing as anywhere else, an
OpenAI-compatible server, and the piece that uses the hardware is that server. This guide uses
llama.cpp's llama-server, started and stopped by gko
itself through a process runtime, with no
container and no daemon.
llama-server (Metal) ← gko backend serve / stop / status / logs
↑
gko <command> → POST /v1/chat/completions
Any other server speaking the protocol (an MLX server, for instance) is wired the same way: a
different command and arguments, not a different gko.
GPU, Metal and the Neural Engine
An Apple Silicon chip has two accelerators, and they are not interchangeable:
- the GPU, programmed through Metal. This is what
llama-serveruses (as does MLX), and what this guide sets up; - the Neural Engine (ANE), which is reachable only through Core ML.
llama-serverdoes not use it, and neither does the server this guide starts.
So, despite the name, nothing on a Mac runs on an NPU in the sense of the Intel guide. The model runs on the GPU, in the unified memory the CPU shares, which is why a model has to fit in RAM with room left over for everything else.
1. Prerequisites
-
An Apple Silicon Mac:
uname -mprintsarm64. -
The
gkobinary foraarch64-apple-darwin, from the releases. -
llama.cpp. Homebrew's build has Metal enabled:
brew install llama.cpp
-
A model in GGUF format, for instance:
hf download Qwen/Qwen2.5-1.5B-Instruct-GGUF qwen2.5-1.5b-instruct-q4_k_m.gguf --local-dir ~/models
Docker is not needed.
2. Checking llama-server
gko backend serve looks command up on PATH the way a shell does, so check it from the shell you will
run gko in:
command -v llama-server
llama-server --versionIf command -v prints nothing, gko backend serve fails in the same way (see
Troubleshooting). A llama-server installed somewhere else can be named by
its absolute path in command.
3. The backend and the model
# .gekko/backends/llamacpp.toml
id = "llamacpp"
type = "openai-compatible"
port = 8080
base_url = "http://127.0.0.1:{{ backend.port }}"
[operations.chat]
method = "POST"
path = "/v1/chat/completions"
[runtime]
type = "process"
command = "llama-server"
arguments = [
"--model", "{{ env.HOME }}/models/{{ args.model }}",
"--alias", "{{ args.model }}",
"--host", "127.0.0.1",
"--port", "{{ backend.port }}",
"--n-gpu-layers", "99",
]
startup_timeout_secs = 120# .gekko/models/qwen-local.toml
id = "qwen-local"
backend = "llamacpp"
operation = "chat"
model = "qwen2.5-1.5b-instruct-q4_k_m.gguf"
[generation]
temperature = 0.0
max_tokens = 512What each line is for:
-
portis declared once.base_urland--portboth read it as{{ backend.port }}, so they cannot drift apart. It has to be a fixed number, because a process runtime cannot hand back a port the kernel picked (port = "auto"is rejected here). -
{{ args.model }}is the model'smodelfield. It names the file to load, and--aliasmakesllama-serveranswer under that same name. -
--n-gpu-layers 99offloads every layer to the GPU. Any number at least as large as the model's layer count means "all of them". -
startup_timeout_secsis how longgko backend servewaits for the port to answer. Loading a large model from disk takes a while, and 30 seconds (the default) can be too short.
Every key is described in Starting a backend as a process.
4. gko backend serve, status, stop
Pids differ from run to run, and so do the OS error numbers: macOS reports "connection refused" as
os error 61.
serve returns only once the server answers on its port, and prints its pid:
$ gko backend serve qwen-local
1269338status lists every backend that declares a runtime:
$ gko backend status
BACKEND RUNTIME INSTANCE URL STATE
llamacpp process 1269338 http://127.0.0.1:8080 runningA second serve refuses to start a second server beside the first one (exit 3):
$ gko backend serve qwen-local
backend error: backend "llamacpp" is already served by process 1269338 — `gko backend status` to see it, `gko backend stop qwen-local` to end itstop sends SIGTERM, then SIGKILL if the server does not exit in time, and prints the backend
id. It is idempotent: stopping a backend that is not running still succeeds.
$ gko backend stop qwen-local
llamacppgko backend logs qwen-local prints what the server wrote on both streams (-f follows it). Its startup
lines are the quickest way to confirm that Metal was actually used:
gko backend logs qwen-local | grep -i -e metal -e offloadedThe record and the log live in $HOME/Library/Application Support/gekko/state/. See
What gko remembers.
5. gko doctor
doctor checks the backend's port and, because the backend declares a process runtime, that its
command can be found. Before serve, the port is expected to fail, and that gives exit 3:
$ gko doctor
✓ configuration loaded
✗ backend "llamacpp" reachable: TCP connection to "127.0.0.1:8080" failed: Connection refused (os error 111)
✓ runtime command "llama-server" available
✓ model "qwen-local"Once the server is up:
$ gko doctor
✓ configuration loaded
✓ backend "llamacpp" reachable
✓ runtime command "llama-server" available
✓ model "qwen-local"doctor connects to the port and stops there. It does not send a chat request, so a green report
does not prove that the model answers well.
Troubleshooting
llama-server is not on PATH. doctor says so:
✗ runtime command "llama-server" available: command "llama-server" not found on PATH — install it, or point the backend's [runtime].command at itand serve refuses with exit 3:
$ gko backend serve qwen-local
backend error: backend "llamacpp": command "llama-server" not found (an absolute or relative path is used as-is, a bare name is looked up on PATH)A shell that finds it while gko does not usually means a different PATH. Homebrew's prefix on
Apple Silicon is /opt/homebrew/bin, and a service manager does not load your shell profile.
The port is taken. serve checks before spawning anything:
$ gko backend serve qwen-local
backend error: backend "llamacpp": port 8080 is already in use by something else — change its "port" key, or stop what is listening on itChange port in the backend file. base_url and --port follow on their own.
The server exits during startup. This happens with a missing model file, an unsupported GGUF,
or not enough memory. serve fails with exit 3, naming the executable, its exit status and the
log file it kept. The server's own explanation is in that file: gko backend logs qwen-local.
serve times out while the model is still loading. Raise startup_timeout_secs.
The server dies when the terminal closes. It is a child of the shell gko backend serve ran in, not
a detached daemon. Run gko backend serve from something that outlives the terminal (nohup, tmux, a
launchd agent) if the server has to stay up.
gko backend stop succeeded but the port is still held. command is a wrapper script that does not
end with exec. gko signalled the wrapper, and the real server survived. End the script with
exec llama-server ….