Access
Callers become identities
Tokens hashed at rest, per-caller isolation, request history and audit, instead of an open port everyone shares.
Anything written against the Ollama API, including the stock CLI, works when the host points at LM-Kit One. Then the application stack starts.
The dialect is the door, not the product: the verbs you already use land on a bigger server.
The unmodified ollama CLI, aimed at your server.
The Ollama Python library, with only the host changed.
The management verbs work too, which is the part compatibility shims usually skip.
Chat, generate, embeddings, tags, show, ps, version, pull, push, create, copy, delete and blob upload are all served; the coverage matrix is the authoritative list.
It worked on your machine. Now five applications and a team need it, and the model runner was never the hard part.
Access
Tokens hashed at rest, per-caller isolation, request history and audit, instead of an open port everyone shares.
Documents
OCR, conversion, structured extraction, search collections and answers with page citations, on the same server.
Agents
Named agents with skills, governed tools and memory, adopted by any client with one field instead of rebuilt per app.
Dialects
The same models also answer OpenAI and Anthropic clients and MCP assistants, so every tool in the house lands on one governed server.
Scale
Horizontal scaling for inference and sustained document workloads when the experiment becomes infrastructure.
Operations
Live telemetry, request attribution, model fit against hardware, alerts and logs, for the person who has to keep it up.
Two differences, stated here rather than discovered later.
Catalog
Model names resolve against the LM-Kit catalog and your imports, not Ollama's registry. Pull what you need, or bring your own GGUF through create and push.
Purpose
LM-Kit One is built to be operated: identities, policies and an admin console come with it. For personal single-machine use, a runner may stay the simpler tool.
LM-Kit One