Configuration
- Layout
- Scopes and precedence
- Merge semantics
- Backends
- Starting a backend with Docker
- Starting a backend as a process
- Models
- When a broader scope is broken
Layout
Configuration is split by concern rather than kept in one monolithic file:
.gekko/
├── backends/
│ └── *.toml # where to send requests, and how
├── models/
│ └── *.toml # which model, on which backend operation
├── commands/
│ └── *.md # the commands themselves (TOML frontmatter + prompt)
└── schemas/
└── *.json # JSON Schema contracts for structured output
Every directory is optional. A missing directory is not an error — it simply contributes nothing.
Scopes and precedence
The same layout can exist at three levels. They are read from broadest to most local, and the most local wins:
/etc/gekko system-wide
↓
$XDG_CONFIG_HOME/gekko (or $HOME/.config/gekko) per user
↓
<walk-up>/.gekko per project
If XDG_CONFIG_HOME is set and non-empty it replaces the $HOME-derived path; it does not add to
it. On macOS, where XDG is not a native convention, the effective path is normally
~/.config/gekko. A scope directory that does not exist is skipped silently.
The project scope
The project scope is the nearest .gekko directory found by walking up from the current
directory, so running gko from a subdirectory of a project still finds that project's .gekko.
The search stops at the project's root — the first directory containing .git (a directory, or a
file in a git worktree) — or at $HOME, whose own .gekko is checked but nothing above it.
--config-dir <DIR> or $GKO_CONFIG_DIR (the flag wins when both are set) name the project scope
directly and skip the walk-up entirely:
$ gko --config-dir /path/to/.gekko config modelsgko doctor reports the project scope it actually resolved, as an Ok line naming the directory.
On Windows, the same three tiers use their own environment variables instead:
%ProgramData%\gekko (system-wide; skipped when %ProgramData% is unset), then %APPDATA%\gekko
if set, else %USERPROFILE%\.config\gekko if set, else %HOME%\.config\gekko.
This lets a repository ship its own .gekko/ with project-specific commands, model aliases and
backend overrides, without touching the machine or the user setup.
A shared .gekko/ cannot launch a program as a side effect of running a business command:
runtime startup is an explicit operator action (gko backend serve), not part of the business
pipeline. Treat command prompts and backend endpoints from a shared repository as untrusted
configuration nonetheless.
Merge semantics
Merging is replacement, not deep merge. The replacement key is:
| Kind | Key |
|---|---|
| Backends | the id field inside the file |
| Models | the id field inside the file |
| Commands | the full command path (git/review), derived from the file path |
A backend with id = "ovms" defined in ./.gekko replaces the /etc/gekko one entirely. A field
present in the broader definition and absent from the local one is not inherited — you get the
local file, whole.
Entries whose keys differ simply accumulate, so a system-wide command and a project command coexist.
Resolution happens after merging, so a model defined in your project can reference a backend
declared only in /etc/gekko.
Duplicate ids within one scope
Two files in the same scope declaring the same id are rejected, naming both paths: across
scopes an override is intended, within one scope it is ambiguous.
Backends
A backend declares the runtime protocol, where to reach it, and which operations it exposes.
gko config schema backend prints the JSON Schema of this file, and gko config schema model
that of a model file; see gko config schema to have an editor
validate them as you type.
# .gekko/backends/ovms.toml
id = "ovms"
type = "openai-compatible"
base_url = "http://127.0.0.1:8000"
[operations.chat]
method = "POST"
path = "/v3/chat/completions"| Key | Required | Notes |
|---|---|---|
id |
yes | the merge key, and how models refer to this backend |
type |
yes |
"openai-compatible", the only supported value |
base_url |
yes | joined with an operation's path; a trailing / is handled either way |
port |
no | the listening port, declared once and read as {{ backend.port }} — see below |
[operations.<name>] |
at least one |
method, path, and optionally protocol — see below |
[runtime] |
no | how gko backend serve starts this backend — see below |
Unknown keys are rejected, with the file and line. A type other than "openai-compatible" and
a method other than POST are both rejected at load time rather than silently ignored.
Operation protocols
An operation's name is yours; what it speaks is its protocol, "chat" when omitted:
[operations.embed]
method = "POST"
path = "/v1/embeddings"
protocol = "embeddings"protocol |
Request | Answer |
|---|---|---|
chat |
{model, messages, ...} |
choices[0].message.content |
embeddings |
{model, input}, the rendered prompt as input
|
data[0].embedding, as a JSON array |
transcriptions |
multipart/form-data: model, prompt (the rendered prompt, when not blank) and the input as file
|
text |
A command running an embeddings model must declare format = "json": its output is the vector,
which [output].schema can constrain (its length, for one) and [output].extract can index.
system, [[examples]], [generation], strip_reasoning and allow_truncated do not apply
and are rejected, naming the command file and the model, when the command runs and by
gko doctor. An embeddings model cannot declare a fallback (two models' vectors cannot be
compared) nor a [generation] table, and a model's fallback must speak the same protocol as the
model itself; both are rejected at load time, naming the model file.
A transcriptions operation (a whisper.cpp server, OVMS whisper) takes audio: a command running
it declares [input] mode = "binary" and format = "text", and the same chat-only keys are
rejected. The prompt, if any, is a hint on vocabulary and spelling and cannot reference
{{ input }}. The upload is named after the FILE argument, whose extension servers often use to
detect the format; read from stdin, it is named input. A transcriptions model takes no
[generation] table; its fallback, if any, must be another transcriptions model.
port (optional)
A containerized backend would otherwise write its port twice — in -p and in base_url — and
if the two diverge, gko doctor stays green (its probe reaches whatever answers on the base_url
port, possibly another backend) while the real request fails with exit 3. port declares the
number once:
port = 8001
base_url = "http://127.0.0.1:{{ backend.port }}"
[runtime]
type = "docker"
options = ["-p", "{{ backend.port }}:8000", "..."]{{ backend.port }} is substituted at load time in base_url and in every [runtime] entry that
can be resolved before the runtime exists: image, options and args for a Docker runtime,
arguments and the [runtime.env] values for a process one — but not in a process runtime's
command.
It is the only backend.* placeholder, and both halves are checked at load time: the placeholder
without a port key is rejected, and so is a port key nothing references.
port = "auto""auto" lets Docker allocate the port: {{ backend.port }} becomes 0 in the [runtime] lists
(-p 0:8000), Docker picks a free port, and gko reads it back with docker port whenever it
needs the URL, so two backends can never collide.
[!IMPORTANT]
port = "auto"makes Docker a prerequisite for executing commands on that backend, not just for its lifecycle. A fixed port never consults Docker at all, so this is strictly opt-in per backend. Two further constraints, both rejected at load time naming the file:"auto"requires a Docker runtime ([runtime]withtype = "docker", there is nothing to read a port back from otherwise), and it requiresbase_urlto read{{ backend.port }}(the allocated port would be unreachable otherwise).
The port changes on each gko backend serve. gko backend status prints the resolved URL, and a backend that is
not started reports - there rather than failing the report.
When a fixed port is already taken
gko backend serve checks before starting anything and stops with exit 3, naming the backend and the
port:
$ gko backend serve m
backend error: backend "probe": port 8001 is already in use by something else — change its "port" key, stop what is listening on it, or use port = "auto" to let Docker allocate oneA fixed port is never moved automatically, since something outside gko may depend on it; use
"auto" when the number does not matter.
A backend that is already served is reported as such instead, since its container is what holds the port:
$ gko backend serve qwen3-8b
backend error: backend "ovms" is already served by container "gko-ovms" — `gko backend status` to see it, `gko backend stop qwen3-8b` to remove it
[timeouts] (optional)
[timeouts]
request_secs = 120request_secs bounds a chat request against this backend, in seconds. Omitted, the backend
falls back to the CLI's own default (120s — enough for a full max_tokens generation on a slow
accelerator such as an NPU). request_secs = 0 is rejected at load time, naming the file.
max_concurrent (optional)
max_concurrent = 1One request at a time to this backend, across every gko process on the machine: a git hook and
an editor action launched together do not make one of them fail with a 5xx or a timeout on a
single NPU. The second invocation waits, its spinner reading waiting for backend "ovms" (busy),
and sends its request once the first one's answer is complete; waiting is not a failure, so it
never triggers the model's fallback. With --no-wait a busy backend is a backend failure like
any other: the model's fallback answers if it has one, otherwise the command fails at once
(exit 3) naming the backend. gko config test and gko mcp serve always wait.
The lock is a file lock under the state directory, released as soon as the
answer ends: a fallback model on the same backend takes it again rather than waiting for its own
primary. It is keyed by backend id and backend file, so two scopes describing the same device do
not wait for each other. Omitted, there is no limit. Only 1 is accepted; any other value is rejected at load
time, naming the file. The state directory comes from XDG_STATE_HOME or HOME, as for the
process records: with neither set, as on a stock Windows session, a max_concurrent backend
fails with exit 1.
structured_output (optional)
structured_output = trueDeclares that the server accepts an OpenAI response_format of type json_schema (OVMS,
llama-server and vLLM do). A command's output schema is then sent with each request and
constrains the model's answer. Omitted or false, nothing is sent and the answer is only
validated after it arrives — a server that rejects unknown request fields never receives one.
[headers] (optional)
[headers]
Authorization = "Bearer {{ env.OPENAI_API_KEY }}"
X-Org = "acme"A table of extra HTTP headers sent with every chat request to this backend — what makes it
possible to talk to llama-server --api-key, LiteLLM, an Ollama behind a proxy, or a hosted
OpenAI-compatible endpoint that requires authentication.
- A value is a template accepting only
{{ env.NAME }}:{{ input }},{{ args.* }}and{{ schemas.* }}are configuration errors here, at load time, naming the file and the header — a header cannot depend on the command being run. -
Content-TypeandContent-Lengthare rejected (case-insensitively):gkosets both. - A header name must be a legal HTTP token (RFC 9110); two names colliding once case is ignored
(
Authorizationnext toauthorization) are rejected too — one of them would silently win. - The environment variable is resolved at preflight, before the command's input is read
(for the model's backend and, when the model declares a
fallback, for the fallback's backend too) — an undefined variable is a configuration error naming the file and the header, andgit diff | gko ...fails before the diff is consumed. - Header values are never logged, at any
--verboselevel, and never shown bygko describe: the trace line names the headers sent (with headers: Authorization, X-Org), never their values. -
gko doctor's TCP probe does not send headers — it only checks the port is open.
Starting a backend with Docker
A backend may declare how to start its own runtime, in a [runtime] table whose type picks the
family: "docker" (below) or "process" (see
Starting a backend as a process). Each family reads its own
keys, and a key belonging to the other one is rejected by name.
gko backend serve <model> then runs it, and the family's prerequisite — Docker here — becomes an
optional one: nothing changes for a configuration without this table.
# .gekko/backends/ovms.toml, continued
[runtime]
type = "docker"
image = "openvino/model_server:latest"
options = ["-p", "8000:8000", "-v", "{{ env.HOME }}/models:/models:rw"]
args = [
"--source_model", "{{ args.model }}",
"--model_repository_path", "/models",
"--rest_port", "8000",
]| Key | Required | Notes |
|---|---|---|
type |
yes |
"docker", which selects this family |
image |
yes | the container image to run |
options |
no | passed to docker run before the image: ports, volumes, devices |
args |
no | passed to the image after it: the server's own arguments |
An unsupported family is rejected with the file named (unknown variant "podman"). Unknown keys
inside [runtime] are rejected like everywhere else, including a key of the other family, such
as image under type = "process".
The legacy [docker] table
A [docker] table, without type, is accepted as a synonym of [runtime] with
type = "docker"; write [runtime] in new files. A backend declaring both is rejected at load
time, naming the file and the backend.
options and args follow the grammar of docker run [OPTIONS] IMAGE [ARG...]. gko adds -d
and --name gko-<backend-id> itself, and nothing else.
Every entry goes through the same templating as a prompt:
-
{{ args.model }}— themodelfield of the model being served. It is the only argument available here; any other name is rejected at load time, naming the file. -
{{ env.NAME }}— an environment variable, required to be defined atservetime. -
{{ input }}— rejected:gko backend servereads no input.
Declaring a Docker runtime also constrains the backend's id, which becomes the container name: ASCII
letters, digits, _, . and -, starting with a letter or a digit. An id outside that set is
rejected — with its file named — rather than mangled into something Docker accepts.
[!WARNING] Scope replacement is per whole backend, never field by field. A project scope that redefines
base_urlforovmsreplaces the user scope'sovmsentirely,[runtime]included. Repeat the table in the local file, orgko backend servewill report that the backend declares none.
For OpenVINO Model Server specifically: the -gpu image tag is the one to use for accelerators
(there is no NPU-only image; that tag carries both plugins), with --device /dev/dri and
--group-add <render gid> in options for the GPU, plus --device /dev/accel for the NPU.
When the served directory is a local export addressed with --model_name/--model_path, no
device flag belongs in args: the device is baked into that export's graph.pbtxt by
ovms --configure, and OVMS reads it from there. One export therefore serves one device, so
running the same weights on the NPU and the GPU means two backends on two ports — which is also
what lets several small models run at once.
Starting a backend as a process
The other runtime family starts a server directly on this machine, with no container and no
daemon: llama.cpp's llama-server, an MLX server, a shell script of your own. Same table, same
gko backend serve / stop / status / logs, different type.
# .gekko/backends/llamacpp.toml, continued
[runtime]
type = "process"
command = "llama-server"
arguments = [
"--model", "{{ args.model }}",
"--host", "127.0.0.1",
"--port", "{{ backend.port }}",
]
startup_timeout_secs = 60
[runtime.env]
LLAMA_CACHE = "{{ env.HOME }}/.cache/llama.cpp"| Key | Required | Notes |
|---|---|---|
type |
yes |
"process", which selects this family |
command |
yes | an absolute or relative path used as-is, or a bare name looked up on PATH; templated like the rest |
arguments |
no | the server's own arguments, one list entry per argument |
[runtime.env] |
no | variables layered over the environment gko itself runs in |
startup_timeout_secs |
no | readiness budget in seconds, 30 by default; 0 is rejected |
arguments is a list of separate entries, never one string to be split, and no shell is
involved: a model path containing a space stays one argument. Entries go
through the same templating as a Docker runtime's — {{ args.model }}, {{ env.NAME }},
{{ backend.port }} — with {{ input }} rejected, since gko backend serve reads no input.
command is templated too — {{ args.model }} and {{ env.NAME }}, so a server living under a
path only the environment knows can be named — but not {{ backend.port }}, which is rejected
there naming the file. gko doctor checks that a command is available only when it is not
templated.
The PATH lookup requires an executable file: a file with no execute bit is skipped and the
scan continues, as a shell does.
[runtime.env] is an overlay, not a replacement: the child inherits gko's own environment
and these values are layered on top. A server needing HOME, PATH or a proxy setting therefore
does not have to redeclare them to gain one variable.
startup_timeout_secs is how long gko backend serve waits, after spawning the server, for it to
answer on its base_url, so a base_url that does not parse into a host and a port is rejected at
load time naming the file. When the budget runs out, the error carries the last probe error.
Unlike the Docker family, this one does not return before the server is ready: a serve that
succeeded means something answered.
Three constraints this family adds, all rejected at load time naming the file:
-
port = "auto"is refused: only Docker can report the port it allocated. Declare a fixedport. - the backend
idmust be usable as a file name (ASCII letters, digits,_,.and-, starting with a letter or a digit), since it names the backend's state file. -
startup_timeout_secsmust be between1and86400.
This family is Unix-only: on Windows a [runtime] type = "process" backend is rejected at load
time naming the file. Use type = "docker" there, or start the server outside gko.
What gko remembers
A Docker runtime needs nothing persisted. For a process, gko backend serve writes a small JSON
record, plus a .log file receiving the server's stdout and stderr. stop deletes the record; logs reads the file. Both are named
<backend id>-<digest>, the digest being the first eight hex characters of the SHA-256 of the
backend file the runtime was declared in. They live in $XDG_STATE_HOME/gekko/, or
$HOME/.local/state/gekko/ when that variable is unset — and on macOS always in
$HOME/Library/Application Support/gekko/state/.
The digest keeps projects apart: two projects each declaring llamacpp in their own ./.gekko get
two records, two logs and two servers, and neither one's gko backend stop or gko backend logs
can reach the other's.
The record names the backend file it was served from. A record naming another file, while its
process is alive, is reported as foreign state and is never signalled, cleared or overwritten —
serve and stop both refuse, naming both files. Once its process is gone, the next serve
replaces it.
The record also holds the pid and its start time, and gko backend stop signals the process only
when both still match: a record whose pid now belongs to another process is reported as
stale state and forgotten, never signalled. The executable is recorded for information only, so a
command ending on exec (a wrapper script, a virtualenv or uv/conda shim) is stopped
correctly.
Changing a served backend's [runtime] family — or removing the table — leaves that record
unreachable: stop, status and logs follow what the files say now, and a backend that now
declares Docker is asked about a container. Run gko backend stop before changing the family.
The record is plain JSON and holds the pid, so a forgotten one is still recoverable by hand.
[!WARNING] The spawned server is a plain child of the shell
gko backend serveran in. It is not detached into its own session, so a terminal hang-up takes it down with everything else in that session. Rungko backend servefrom a session that outlives it (a service manager,nohup, a multiplexer) if the server is meant to stay up.
[!WARNING]
gkosignals the process it spawned, and only that one. Acommandwhose process is the server — a binary, or a launcher ending onexec— is stopped correctly. A launcher that forks and waits instead (sh -c "server | tee log",conda run, anything that does notexec) has its wrapper signalled while the real server survives:gko backend stopreports success and deletes the record, a failedgko backend serveterminates the wrapper and abandons the rest, and the orphan keeps the port while everygkocommand reports the backend as never started. End your launcher onexec.
Models
A model is the bridge between a command and a backend capability.
# .gekko/models/qwen-fast.toml
id = "qwen-fast"
backend = "ovms"
operation = "chat"
model = "qwen-2.5-1.5b"
[generation]
temperature = 0.0
max_tokens = 512| Key | Required | Notes |
|---|---|---|
id |
yes | the merge key, and the name commands use |
backend |
yes | must match a backend id
|
operation |
yes | must be an operation that backend exposes |
model |
yes | the concrete model identifier sent to the backend |
fallback |
no | another model id to retry against when this one fails — see below |
[generation] |
no | see Generation parameters below |
A model naming an unknown backend, or an operation its backend does not expose, produces a
configuration error listing what is available — when a command runs it, and in gko doctor.
It is not a load-time error: a model nobody uses does not break the rest of the configuration.
Generation parameters
[generation]
temperature = 0.0
max_tokens = 512
seed = 42
top_p = 0.9
stop = ["\n\n", "###"]
[generation.extra]
chat_template_kwargs = { enable_thinking = false }| Key | Notes |
|---|---|
temperature |
float |
max_tokens |
integer |
seed |
integer, forwarded as-is for deterministic sampling |
top_p |
float |
stop |
a non-empty list of non-empty strings, with no upper bound |
[generation.extra] |
free-form table, forwarded verbatim at the top level of the request, after the typed keys above |
A value is only included in the request when present — no null is ever serialized for an
absent field.
[generation.extra] carries engine-specific settings gko has no key for (for example
chat_template_kwargs.enable_thinking = false on Qwen3). Its keys are converted from TOML to JSON
as they are (tables become objects, arrays become arrays). Two things are rejected at load time,
naming the file:
- a key of
extrathat collides with a typed key (temperature,max_tokens,seed,top_p,stop, and the keysgkoitself controls:model,messages,stream,response_format) — use the typed key instead; - a TOML
Datetimeor a non-finite float (nan,inf, legal TOML float literals) anywhere insideextra, including nested in a table or an array, since JSON cannot represent either. A plain string that merely looks like a date ("2024-01-01") is unaffected: only a genuine TOML datetime value is rejected.
Per-command override (the one field-by-field merge)
A command's own frontmatter may declare its own [generation], in the same shape (extra
included). Unlike every other configuration key, this one merges instead of replacing: each typed key the command sets overrides the model's; a key the command leaves
unset keeps the model's value. extra merges the same way, key by key, at the top level only —
a command setting chat_template_kwargs replaces the model's chat_template_kwargs wholesale,
never merged one level deeper.
This is the one exception to Merge semantics: a command gets the model's
defaults plus its own changes. When a model declares a fallback, the fallback uses its own base [generation]
merged with the same command override — the override describes the command being run, not
which model answers it.
gko describe prints the effective table (after the merge), so an agent sees exactly what
will be sent. The diagnostic log line at info names the generation keys actually sent
(including each extra key, as extra.<key>), never their values.
fallback (optional)
id = "qwen3-8b"
backend = "ovms" # OVMS on the NPU
model = "qwen3-8b-int4-ov"
fallback = "qwen3-8b-gpu"When the request fails with exit code 3 (a backend failure: unreachable, or a non-2xx answer),
the same rendered prompt is sent once to the fallback model, on its own backend. Nothing else is
retried — a 2 (bad configuration) and a 4 (the answer violated the output contract) are
returned as-is.
The typical case: a model compiled for an Intel NPU has a static maximum prompt length, and OVMS
refuses an over-long prompt at once with 400 ... Input length exceeds the maximum allowed length.
The fallback sends that prompt to a GPU-served model instead.
Three properties worth knowing:
-
The retry is single hop. The fallback's own
fallbackis not followed, so a chain cannot form and no cycle is possible. -
It is blind to the reason.
gkocannot tell an over-long prompt from a stopped container, so the primary failure is always written to stderr atwarnlevel. Without it, a backend that has been down all day would look like a healthy fallback. -
It is checked at load time. A
fallbacknaming an unknown model, or naming its own model, is a configuration error (exit2) naming the file — not a surprise on the day the recovery is actually needed.
When both fail, the error names both models and both backends. gko config models shows the FALLBACK
column so the routing is never invisible.
[!IMPORTANT] The fallback does not lift the model's context length. A GPU twin built by symlinking the primary's export shares its
config.json, hence its context length (40960 tokens forqwen3-8b). The fallback recovers prompts sitting between the NPU's compiled shape and that ceiling; past it both fail, with two distinct messages —Input length exceeds the maximum allowed lengthis the NPU's static shape,Number of prompt tokens: N exceeds model max length: Mis the model's context, and only the first one is recoverable.
[!NOTE] The NPU and the GPU need two separate backends, hence two containers and two ports: the target device is baked into the served export (OVMS reads it from
graph.pbtxt), not chosen per request. The same is true of running several small models at once — one backend each. See Starting a backend with Docker.
When a broader scope is broken
A broken file in /etc/gekko, which may be out of your reach, does not disable your project: a
broadly-scoped entry that is entirely shadowed by a more local one does not break anything.
Backends and models differ from commands here:
| Where the key lives | Consequence | |
|---|---|---|
| Backends, models | the id field, inside the file |
An unparseable file cannot be matched to an id, so a parse error is always fatal. Validation of type and method applies only to the entries that win the merge. |
| Commands | the file path | Only winning files are read: a broken but shadowed command file is ignored. |
Output schemas behave the same way: a schema that is missing or malformed on a command
nobody invokes does not break gko --help. Checking every schema in every scope is the job of
gko doctor.