Gekko

Configuration


Layout

Configuration is split by concern rather than kept in one monolithic file:

.gekko/
├── backends/
│   └── *.toml      # where to send requests, and how
├── models/
│   └── *.toml      # which model, on which backend operation
├── commands/
│   └── *.md        # the commands themselves (TOML frontmatter + prompt)
└── schemas/
    └── *.json      # JSON Schema contracts for structured output

Every directory is optional. A missing directory is not an error — it simply contributes nothing.


Scopes and precedence

The same layout can exist at three levels. They are read from broadest to most local, and the most local wins:

/etc/gekko                                           system-wide
      ↓
$XDG_CONFIG_HOME/gekko   (or $HOME/.config/gekko)    per user
      ↓
<walk-up>/.gekko                                     per project

If XDG_CONFIG_HOME is set and non-empty it replaces the $HOME-derived path; it does not add to it. On macOS, where XDG is not a native convention, the effective path is normally ~/.config/gekko. A scope directory that does not exist is skipped silently.

The project scope

The project scope is the nearest .gekko directory found by walking up from the current directory, so running gko from a subdirectory of a project still finds that project's .gekko. The search stops at the project's root — the first directory containing .git (a directory, or a file in a git worktree) — or at $HOME, whose own .gekko is checked but nothing above it.

--config-dir <DIR> or $GKO_CONFIG_DIR (the flag wins when both are set) name the project scope directly and skip the walk-up entirely:

$ gko --config-dir /path/to/.gekko config models

gko doctor reports the project scope it actually resolved, as an Ok line naming the directory.

On Windows, the same three tiers use their own environment variables instead: %ProgramData%\gekko (system-wide; skipped when %ProgramData% is unset), then %APPDATA%\gekko if set, else %USERPROFILE%\.config\gekko if set, else %HOME%\.config\gekko.

This lets a repository ship its own .gekko/ with project-specific commands, model aliases and backend overrides, without touching the machine or the user setup.

A shared .gekko/ cannot launch a program as a side effect of running a business command: runtime startup is an explicit operator action (gko backend serve), not part of the business pipeline. Treat command prompts and backend endpoints from a shared repository as untrusted configuration nonetheless.


Merge semantics

Merging is replacement, not deep merge. The replacement key is:

Kind Key
Backends the id field inside the file
Models the id field inside the file
Commands the full command path (git/review), derived from the file path

A backend with id = "ovms" defined in ./.gekko replaces the /etc/gekko one entirely. A field present in the broader definition and absent from the local one is not inherited — you get the local file, whole.

Entries whose keys differ simply accumulate, so a system-wide command and a project command coexist.

Resolution happens after merging, so a model defined in your project can reference a backend declared only in /etc/gekko.

Duplicate ids within one scope

Two files in the same scope declaring the same id are rejected, naming both paths: across scopes an override is intended, within one scope it is ambiguous.


Backends

A backend declares the runtime protocol, where to reach it, and which operations it exposes. gko config schema backend prints the JSON Schema of this file, and gko config schema model that of a model file; see gko config schema to have an editor validate them as you type.

# .gekko/backends/ovms.toml
id = "ovms"
type = "openai-compatible"
base_url = "http://127.0.0.1:8000"

[operations.chat]
method = "POST"
path = "/v3/chat/completions"
Key Required Notes
id yes the merge key, and how models refer to this backend
type yes "openai-compatible", the only supported value
base_url yes joined with an operation's path; a trailing / is handled either way
port no the listening port, declared once and read as {{ backend.port }} — see below
[operations.<name>] at least one method, path, and optionally protocol — see below
[runtime] no how gko backend serve starts this backend — see below

Unknown keys are rejected, with the file and line. A type other than "openai-compatible" and a method other than POST are both rejected at load time rather than silently ignored.

Operation protocols

An operation's name is yours; what it speaks is its protocol, "chat" when omitted:

[operations.embed]
method = "POST"
path = "/v1/embeddings"
protocol = "embeddings"
protocol Request Answer
chat {model, messages, ...} choices[0].message.content
embeddings {model, input}, the rendered prompt as input data[0].embedding, as a JSON array
transcriptions multipart/form-data: model, prompt (the rendered prompt, when not blank) and the input as file text

A command running an embeddings model must declare format = "json": its output is the vector, which [output].schema can constrain (its length, for one) and [output].extract can index. system, [[examples]], [generation], strip_reasoning and allow_truncated do not apply and are rejected, naming the command file and the model, when the command runs and by gko doctor. An embeddings model cannot declare a fallback (two models' vectors cannot be compared) nor a [generation] table, and a model's fallback must speak the same protocol as the model itself; both are rejected at load time, naming the model file.

A transcriptions operation (a whisper.cpp server, OVMS whisper) takes audio: a command running it declares [input] mode = "binary" and format = "text", and the same chat-only keys are rejected. The prompt, if any, is a hint on vocabulary and spelling and cannot reference {{ input }}. The upload is named after the FILE argument, whose extension servers often use to detect the format; read from stdin, it is named input. A transcriptions model takes no [generation] table; its fallback, if any, must be another transcriptions model.

port (optional)

A containerized backend would otherwise write its port twice — in -p and in base_url — and if the two diverge, gko doctor stays green (its probe reaches whatever answers on the base_url port, possibly another backend) while the real request fails with exit 3. port declares the number once:

port = 8001
base_url = "http://127.0.0.1:{{ backend.port }}"

[runtime]
type = "docker"
options = ["-p", "{{ backend.port }}:8000", "..."]

{{ backend.port }} is substituted at load time in base_url and in every [runtime] entry that can be resolved before the runtime exists: image, options and args for a Docker runtime, arguments and the [runtime.env] values for a process one — but not in a process runtime's command.

It is the only backend.* placeholder, and both halves are checked at load time: the placeholder without a port key is rejected, and so is a port key nothing references.

port = "auto"

"auto" lets Docker allocate the port: {{ backend.port }} becomes 0 in the [runtime] lists (-p 0:8000), Docker picks a free port, and gko reads it back with docker port whenever it needs the URL, so two backends can never collide.

[!IMPORTANT] port = "auto" makes Docker a prerequisite for executing commands on that backend, not just for its lifecycle. A fixed port never consults Docker at all, so this is strictly opt-in per backend. Two further constraints, both rejected at load time naming the file: "auto" requires a Docker runtime ([runtime] with type = "docker", there is nothing to read a port back from otherwise), and it requires base_url to read {{ backend.port }} (the allocated port would be unreachable otherwise).

The port changes on each gko backend serve. gko backend status prints the resolved URL, and a backend that is not started reports - there rather than failing the report.

When a fixed port is already taken

gko backend serve checks before starting anything and stops with exit 3, naming the backend and the port:

$ gko backend serve m
backend error: backend "probe": port 8001 is already in use by something else — change its "port" key, stop what is listening on it, or use port = "auto" to let Docker allocate one

A fixed port is never moved automatically, since something outside gko may depend on it; use "auto" when the number does not matter.

A backend that is already served is reported as such instead, since its container is what holds the port:

$ gko backend serve qwen3-8b
backend error: backend "ovms" is already served by container "gko-ovms" — `gko backend status` to see it, `gko backend stop qwen3-8b` to remove it

[timeouts] (optional)

[timeouts]
request_secs = 120

request_secs bounds a chat request against this backend, in seconds. Omitted, the backend falls back to the CLI's own default (120s — enough for a full max_tokens generation on a slow accelerator such as an NPU). request_secs = 0 is rejected at load time, naming the file.

max_concurrent (optional)

max_concurrent = 1

One request at a time to this backend, across every gko process on the machine: a git hook and an editor action launched together do not make one of them fail with a 5xx or a timeout on a single NPU. The second invocation waits, its spinner reading waiting for backend "ovms" (busy), and sends its request once the first one's answer is complete; waiting is not a failure, so it never triggers the model's fallback. With --no-wait a busy backend is a backend failure like any other: the model's fallback answers if it has one, otherwise the command fails at once (exit 3) naming the backend. gko config test and gko mcp serve always wait.

The lock is a file lock under the state directory, released as soon as the answer ends: a fallback model on the same backend takes it again rather than waiting for its own primary. It is keyed by backend id and backend file, so two scopes describing the same device do not wait for each other. Omitted, there is no limit. Only 1 is accepted; any other value is rejected at load time, naming the file. The state directory comes from XDG_STATE_HOME or HOME, as for the process records: with neither set, as on a stock Windows session, a max_concurrent backend fails with exit 1.

structured_output (optional)

structured_output = true

Declares that the server accepts an OpenAI response_format of type json_schema (OVMS, llama-server and vLLM do). A command's output schema is then sent with each request and constrains the model's answer. Omitted or false, nothing is sent and the answer is only validated after it arrives — a server that rejects unknown request fields never receives one.

[headers] (optional)

[headers]
Authorization = "Bearer {{ env.OPENAI_API_KEY }}"
X-Org = "acme"

A table of extra HTTP headers sent with every chat request to this backend — what makes it possible to talk to llama-server --api-key, LiteLLM, an Ollama behind a proxy, or a hosted OpenAI-compatible endpoint that requires authentication.


Starting a backend with Docker

A backend may declare how to start its own runtime, in a [runtime] table whose type picks the family: "docker" (below) or "process" (see Starting a backend as a process). Each family reads its own keys, and a key belonging to the other one is rejected by name.

gko backend serve <model> then runs it, and the family's prerequisite — Docker here — becomes an optional one: nothing changes for a configuration without this table.

# .gekko/backends/ovms.toml, continued
[runtime]
type = "docker"
image = "openvino/model_server:latest"
options = ["-p", "8000:8000", "-v", "{{ env.HOME }}/models:/models:rw"]
args = [
    "--source_model", "{{ args.model }}",
    "--model_repository_path", "/models",
    "--rest_port", "8000",
]
Key Required Notes
type yes "docker", which selects this family
image yes the container image to run
options no passed to docker run before the image: ports, volumes, devices
args no passed to the image after it: the server's own arguments

An unsupported family is rejected with the file named (unknown variant "podman"). Unknown keys inside [runtime] are rejected like everywhere else, including a key of the other family, such as image under type = "process".

The legacy [docker] table

A [docker] table, without type, is accepted as a synonym of [runtime] with type = "docker"; write [runtime] in new files. A backend declaring both is rejected at load time, naming the file and the backend.

options and args follow the grammar of docker run [OPTIONS] IMAGE [ARG...]. gko adds -d and --name gko-<backend-id> itself, and nothing else.

Every entry goes through the same templating as a prompt:

Declaring a Docker runtime also constrains the backend's id, which becomes the container name: ASCII letters, digits, _, . and -, starting with a letter or a digit. An id outside that set is rejected — with its file named — rather than mangled into something Docker accepts.

[!WARNING] Scope replacement is per whole backend, never field by field. A project scope that redefines base_url for ovms replaces the user scope's ovms entirely, [runtime] included. Repeat the table in the local file, or gko backend serve will report that the backend declares none.

For OpenVINO Model Server specifically: the -gpu image tag is the one to use for accelerators (there is no NPU-only image; that tag carries both plugins), with --device /dev/dri and --group-add <render gid> in options for the GPU, plus --device /dev/accel for the NPU.

When the served directory is a local export addressed with --model_name/--model_path, no device flag belongs in args: the device is baked into that export's graph.pbtxt by ovms --configure, and OVMS reads it from there. One export therefore serves one device, so running the same weights on the NPU and the GPU means two backends on two ports — which is also what lets several small models run at once.


Starting a backend as a process

The other runtime family starts a server directly on this machine, with no container and no daemon: llama.cpp's llama-server, an MLX server, a shell script of your own. Same table, same gko backend serve / stop / status / logs, different type.

# .gekko/backends/llamacpp.toml, continued
[runtime]
type = "process"
command = "llama-server"
arguments = [
    "--model", "{{ args.model }}",
    "--host", "127.0.0.1",
    "--port", "{{ backend.port }}",
]
startup_timeout_secs = 60

[runtime.env]
LLAMA_CACHE = "{{ env.HOME }}/.cache/llama.cpp"
Key Required Notes
type yes "process", which selects this family
command yes an absolute or relative path used as-is, or a bare name looked up on PATH; templated like the rest
arguments no the server's own arguments, one list entry per argument
[runtime.env] no variables layered over the environment gko itself runs in
startup_timeout_secs no readiness budget in seconds, 30 by default; 0 is rejected

arguments is a list of separate entries, never one string to be split, and no shell is involved: a model path containing a space stays one argument. Entries go through the same templating as a Docker runtime's — {{ args.model }}, {{ env.NAME }}, {{ backend.port }} — with {{ input }} rejected, since gko backend serve reads no input.

command is templated too — {{ args.model }} and {{ env.NAME }}, so a server living under a path only the environment knows can be named — but not {{ backend.port }}, which is rejected there naming the file. gko doctor checks that a command is available only when it is not templated.

The PATH lookup requires an executable file: a file with no execute bit is skipped and the scan continues, as a shell does.

[runtime.env] is an overlay, not a replacement: the child inherits gko's own environment and these values are layered on top. A server needing HOME, PATH or a proxy setting therefore does not have to redeclare them to gain one variable.

startup_timeout_secs is how long gko backend serve waits, after spawning the server, for it to answer on its base_url, so a base_url that does not parse into a host and a port is rejected at load time naming the file. When the budget runs out, the error carries the last probe error. Unlike the Docker family, this one does not return before the server is ready: a serve that succeeded means something answered.

Three constraints this family adds, all rejected at load time naming the file:

This family is Unix-only: on Windows a [runtime] type = "process" backend is rejected at load time naming the file. Use type = "docker" there, or start the server outside gko.

What gko remembers

A Docker runtime needs nothing persisted. For a process, gko backend serve writes a small JSON record, plus a .log file receiving the server's stdout and stderr. stop deletes the record; logs reads the file. Both are named <backend id>-<digest>, the digest being the first eight hex characters of the SHA-256 of the backend file the runtime was declared in. They live in $XDG_STATE_HOME/gekko/, or $HOME/.local/state/gekko/ when that variable is unset — and on macOS always in $HOME/Library/Application Support/gekko/state/.

The digest keeps projects apart: two projects each declaring llamacpp in their own ./.gekko get two records, two logs and two servers, and neither one's gko backend stop or gko backend logs can reach the other's.

The record names the backend file it was served from. A record naming another file, while its process is alive, is reported as foreign state and is never signalled, cleared or overwritten — serve and stop both refuse, naming both files. Once its process is gone, the next serve replaces it.

The record also holds the pid and its start time, and gko backend stop signals the process only when both still match: a record whose pid now belongs to another process is reported as stale state and forgotten, never signalled. The executable is recorded for information only, so a command ending on exec (a wrapper script, a virtualenv or uv/conda shim) is stopped correctly.

Changing a served backend's [runtime] family — or removing the table — leaves that record unreachable: stop, status and logs follow what the files say now, and a backend that now declares Docker is asked about a container. Run gko backend stop before changing the family. The record is plain JSON and holds the pid, so a forgotten one is still recoverable by hand.

[!WARNING] The spawned server is a plain child of the shell gko backend serve ran in. It is not detached into its own session, so a terminal hang-up takes it down with everything else in that session. Run gko backend serve from a session that outlives it (a service manager, nohup, a multiplexer) if the server is meant to stay up.

[!WARNING] gko signals the process it spawned, and only that one. A command whose process is the server — a binary, or a launcher ending on exec — is stopped correctly. A launcher that forks and waits instead (sh -c "server | tee log", conda run, anything that does not exec) has its wrapper signalled while the real server survives: gko backend stop reports success and deletes the record, a failed gko backend serve terminates the wrapper and abandons the rest, and the orphan keeps the port while every gko command reports the backend as never started. End your launcher on exec.


Models

A model is the bridge between a command and a backend capability.

# .gekko/models/qwen-fast.toml
id = "qwen-fast"
backend = "ovms"
operation = "chat"
model = "qwen-2.5-1.5b"

[generation]
temperature = 0.0
max_tokens = 512
Key Required Notes
id yes the merge key, and the name commands use
backend yes must match a backend id
operation yes must be an operation that backend exposes
model yes the concrete model identifier sent to the backend
fallback no another model id to retry against when this one fails — see below
[generation] no see Generation parameters below

A model naming an unknown backend, or an operation its backend does not expose, produces a configuration error listing what is available — when a command runs it, and in gko doctor. It is not a load-time error: a model nobody uses does not break the rest of the configuration.

Generation parameters

[generation]
temperature = 0.0
max_tokens = 512
seed = 42
top_p = 0.9
stop = ["\n\n", "###"]

[generation.extra]
chat_template_kwargs = { enable_thinking = false }
Key Notes
temperature float
max_tokens integer
seed integer, forwarded as-is for deterministic sampling
top_p float
stop a non-empty list of non-empty strings, with no upper bound
[generation.extra] free-form table, forwarded verbatim at the top level of the request, after the typed keys above

A value is only included in the request when present — no null is ever serialized for an absent field.

[generation.extra] carries engine-specific settings gko has no key for (for example chat_template_kwargs.enable_thinking = false on Qwen3). Its keys are converted from TOML to JSON as they are (tables become objects, arrays become arrays). Two things are rejected at load time, naming the file:

Per-command override (the one field-by-field merge)

A command's own frontmatter may declare its own [generation], in the same shape (extra included). Unlike every other configuration key, this one merges instead of replacing: each typed key the command sets overrides the model's; a key the command leaves unset keeps the model's value. extra merges the same way, key by key, at the top level only — a command setting chat_template_kwargs replaces the model's chat_template_kwargs wholesale, never merged one level deeper.

This is the one exception to Merge semantics: a command gets the model's defaults plus its own changes. When a model declares a fallback, the fallback uses its own base [generation] merged with the same command override — the override describes the command being run, not which model answers it.

gko describe prints the effective table (after the merge), so an agent sees exactly what will be sent. The diagnostic log line at info names the generation keys actually sent (including each extra key, as extra.<key>), never their values.

fallback (optional)

id = "qwen3-8b"
backend = "ovms"        # OVMS on the NPU
model = "qwen3-8b-int4-ov"
fallback = "qwen3-8b-gpu"

When the request fails with exit code 3 (a backend failure: unreachable, or a non-2xx answer), the same rendered prompt is sent once to the fallback model, on its own backend. Nothing else is retried — a 2 (bad configuration) and a 4 (the answer violated the output contract) are returned as-is.

The typical case: a model compiled for an Intel NPU has a static maximum prompt length, and OVMS refuses an over-long prompt at once with 400 ... Input length exceeds the maximum allowed length. The fallback sends that prompt to a GPU-served model instead.

Three properties worth knowing:

When both fail, the error names both models and both backends. gko config models shows the FALLBACK column so the routing is never invisible.

[!IMPORTANT] The fallback does not lift the model's context length. A GPU twin built by symlinking the primary's export shares its config.json, hence its context length (40960 tokens for qwen3-8b). The fallback recovers prompts sitting between the NPU's compiled shape and that ceiling; past it both fail, with two distinct messages — Input length exceeds the maximum allowed length is the NPU's static shape, Number of prompt tokens: N exceeds model max length: M is the model's context, and only the first one is recoverable.

[!NOTE] The NPU and the GPU need two separate backends, hence two containers and two ports: the target device is baked into the served export (OVMS reads it from graph.pbtxt), not chosen per request. The same is true of running several small models at once — one backend each. See Starting a backend with Docker.


When a broader scope is broken

A broken file in /etc/gekko, which may be out of your reach, does not disable your project: a broadly-scoped entry that is entirely shadowed by a more local one does not break anything. Backends and models differ from commands here:

Where the key lives Consequence
Backends, models the id field, inside the file An unparseable file cannot be matched to an id, so a parse error is always fatal. Validation of type and method applies only to the entries that win the merge.
Commands the file path Only winning files are read: a broken but shadowed command file is ignored.

Output schemas behave the same way: a schema that is missing or malformed on a command nobody invokes does not break gko --help. Checking every schema in every scope is the job of gko doctor.