Gekko

Running on Apple Silicon


Overview

gko needs nothing Apple-specific. On a Mac it drives the same thing as anywhere else, an OpenAI-compatible server, and the piece that uses the hardware is that server. This guide uses llama.cpp's llama-server, started and stopped by gko itself through a process runtime, with no container and no daemon.

llama-server (Metal)  ←  gko backend serve / stop / status / logs
        ↑
gko <command>  →  POST /v1/chat/completions

Any other server speaking the protocol (an MLX server, for instance) is wired the same way: a different command and arguments, not a different gko.

GPU, Metal and the Neural Engine

An Apple Silicon chip has two accelerators, and they are not interchangeable:

So, despite the name, nothing on a Mac runs on an NPU in the sense of the Intel guide. The model runs on the GPU, in the unified memory the CPU shares, which is why a model has to fit in RAM with room left over for everything else.


1. Prerequisites

Docker is not needed.

2. Checking llama-server

gko backend serve looks command up on PATH the way a shell does, so check it from the shell you will run gko in:

command -v llama-server
llama-server --version

If command -v prints nothing, gko backend serve fails in the same way (see Troubleshooting). A llama-server installed somewhere else can be named by its absolute path in command.


3. The backend and the model

# .gekko/backends/llamacpp.toml
id = "llamacpp"
type = "openai-compatible"
port = 8080
base_url = "http://127.0.0.1:{{ backend.port }}"

[operations.chat]
method = "POST"
path = "/v1/chat/completions"

[runtime]
type = "process"
command = "llama-server"
arguments = [
    "--model", "{{ env.HOME }}/models/{{ args.model }}",
    "--alias", "{{ args.model }}",
    "--host", "127.0.0.1",
    "--port", "{{ backend.port }}",
    "--n-gpu-layers", "99",
]
startup_timeout_secs = 120
# .gekko/models/qwen-local.toml
id = "qwen-local"
backend = "llamacpp"
operation = "chat"
model = "qwen2.5-1.5b-instruct-q4_k_m.gguf"

[generation]
temperature = 0.0
max_tokens = 512

What each line is for:

Every key is described in Starting a backend as a process.


4. gko backend serve, status, stop

Pids differ from run to run, and so do the OS error numbers: macOS reports "connection refused" as os error 61.

serve returns only once the server answers on its port, and prints its pid:

$ gko backend serve qwen-local
1269338

status lists every backend that declares a runtime:

$ gko backend status
BACKEND   RUNTIME  INSTANCE  URL                    STATE
llamacpp  process  1269338   http://127.0.0.1:8080  running

A second serve refuses to start a second server beside the first one (exit 3):

$ gko backend serve qwen-local
backend error: backend "llamacpp" is already served by process 1269338 — `gko backend status` to see it, `gko backend stop qwen-local` to end it

stop sends SIGTERM, then SIGKILL if the server does not exit in time, and prints the backend id. It is idempotent: stopping a backend that is not running still succeeds.

$ gko backend stop qwen-local
llamacpp

gko backend logs qwen-local prints what the server wrote on both streams (-f follows it). Its startup lines are the quickest way to confirm that Metal was actually used:

gko backend logs qwen-local | grep -i -e metal -e offloaded

The record and the log live in $HOME/Library/Application Support/gekko/state/. See What gko remembers.


5. gko doctor

doctor checks the backend's port and, because the backend declares a process runtime, that its command can be found. Before serve, the port is expected to fail, and that gives exit 3:

$ gko doctor
✓ configuration loaded
✗ backend "llamacpp" reachable: TCP connection to "127.0.0.1:8080" failed: Connection refused (os error 111)
✓ runtime command "llama-server" available
✓ model "qwen-local"

Once the server is up:

$ gko doctor
✓ configuration loaded
✓ backend "llamacpp" reachable
✓ runtime command "llama-server" available
✓ model "qwen-local"

doctor connects to the port and stops there. It does not send a chat request, so a green report does not prove that the model answers well.


Troubleshooting

llama-server is not on PATH. doctor says so:

✗ runtime command "llama-server" available: command "llama-server" not found on PATH — install it, or point the backend's [runtime].command at it

and serve refuses with exit 3:

$ gko backend serve qwen-local
backend error: backend "llamacpp": command "llama-server" not found (an absolute or relative path is used as-is, a bare name is looked up on PATH)

A shell that finds it while gko does not usually means a different PATH. Homebrew's prefix on Apple Silicon is /opt/homebrew/bin, and a service manager does not load your shell profile.

The port is taken. serve checks before spawning anything:

$ gko backend serve qwen-local
backend error: backend "llamacpp": port 8080 is already in use by something else — change its "port" key, or stop what is listening on it

Change port in the backend file. base_url and --port follow on their own.

The server exits during startup. This happens with a missing model file, an unsupported GGUF, or not enough memory. serve fails with exit 3, naming the executable, its exit status and the log file it kept. The server's own explanation is in that file: gko backend logs qwen-local.

serve times out while the model is still loading. Raise startup_timeout_secs.

The server dies when the terminal closes. It is a child of the shell gko backend serve ran in, not a detached daemon. Run gko backend serve from something that outlives the terminal (nohup, tmux, a launchd agent) if the server has to stay up.

gko backend stop succeeded but the port is still held. command is a wrapper script that does not end with exec. gko signalled the wrapper, and the real server survived. End the script with exec llama-server ….