Models

Last updated

A model is something that talks, draws pictures, writes music, reads text aloud, writes down what is said, turns text into vectors or makes videos, published at an address like @you/legal-llama. Terminus's own models live under @terminus; yours runs on your own server, and Terminus calls it there. Everyone uses a model the same way: by its address.

On this page⌄

A model's two files#

A model is a folder with terminus.json, its header, and model.json, which says where it runs and how much it reads. Its name, description, prices, who can see it and its key are set on the web, never in a file.

// terminus.json
{ "kind": "model", "modality": "chat", "id": "@you/legal-llama", "version": "0.0.1" }

// model.json
{ "server": "https://llm.example.com/v1", "model_name": "legal-llama-3-70b", "context": 131072 }
modality
chat, image, music, speech, transcription, embedding or video: what it does. It is fixed for the life of the address.
server
The base URL Terminus calls, public https, at each kind's own path in OpenAI's format: a chat model at {server}/chat/completions, an image model at /images/generations, a speech model at /audio/speech, a transcription model at /audio/transcriptions, an embedding model at /embeddings and a video model at /videos. A music model is called at /music.
model_name
The name your server expects in each request's model.
context
Chat: the tokens it reads in one call, from 1,024 to 10,000,000. Embedding: the tokens it reads of one text, from 128 to 1,000,000.
output
Chat: the most tokens one answer may write. 8,192 unless you say.
images, tools, reasoning
Chat: whether it reads pictures (no unless you say), calls tools (yes unless you say) and takes a reasoning effort (no unless you say).
voices
Speech: the voices it reads in, 1 to 64 names of letters, digits, spaces, ., _ or -. The first is its voice unless a call names another; a later release may add voices, never take one away.
dimensions
Embedding: the length of each vector it answers, from 1 to 8,192. It is fixed for the life of the address: vectors of two lengths never compare.
shortens
Embedding: whether it answers fewer dimensions when a call asks for them (no unless you say).

An image, music, transcription or video model needs nothing in model.json beyond where it runs. A speech model names its voices, and an embedding model its vectors:

// model.json, a speech model
{ "server": "https://voice.example.com/v1", "model_name": "narrator-1", "voices": ["alloy", "nova"] }

// model.json, an embedding model
{ "server": "https://vectors.example.com/v1", "model_name": "embed-1", "dimensions": 1536, "context": 8191, "shortens": true }

Make the model on the web first — Create on the desktop at /os, then Model — and fill in its Model form, or push the two files from a working copy with terminus push. Either way the draft is checked as it is saved, and a refusal says exactly what is wrong.

Its key, or a signed call#

If your server wants a key, set it as the model's API_KEY — with the key on its page, or terminus secrets set API_KEY. Terminus sends it as Authorization: Bearer with every call, and never shows it again. A key never goes in a file.

With no key, Terminus signs each call instead, once you show the server is yours: serve the token the page gives you at https://your-server/.well-known/terminus-model-verification, then press Verify.

What Terminus sends your server is only the request, never who it is for.

Prices per unit#

A model is priced per unit of what it does, in credits (a credit is a dollar), and its next release carries the prices you set:

Chat
Input and output per million tokens; cache read and cache write per million tokens too, or blank, when cached tokens bill as input.
Image
Per picture returned.
Music
Per minute of audio, billed by the second.
Speech
Per million characters read.
Transcription
Per minute of the audio, billed by the second.
Embedding
Per million tokens read.
Video
Per second of video: the video's own length, at most the seconds asked.

A model cannot publish before its prices are set. Each call people pay for earns you credits.

Use a model#

Every kind names a model by its address:

Norbert
Pick another chat model for a conversation from the model menu on his line. A model that runs on someone's own server says where before your first conversation goes there. He draws pictures, writes music, reads text aloud, transcribes recordings and makes videos with the models Terminus sets for him, asks before the first one costs you anything in a conversation, and keeps the files he makes in your Downloads.
Agents
models in an agent's terminus.json names the model it talks with, and tools: ["model:@terminus/gpt-image-2"] gives it a picture, music, speech, transcription or video model as a tool: generate_image, compose_music, generate_speech, transcribe_audio or generate_video, a second of a kind <slug>_generate_image. A chat or embedding model is no tool. The person talking to the agent pays.
Apps
models in the app's schema.ts names each model it calls by address; its handle's chat answers a chat model's completion and embed an embedding model's vectors, while generate, compose, speak and transcribe answer the job that makes the pictures or the video, the music, the speech or the transcript. See Models in the SDK reference.
Services
An operation lists the models it calls in models and calls them with ctx.models.call(address, input). The service's owner pays, within its daily budget; a picture, a song, speech or a video comes back finished, its files in the answer, a transcript with its words, and an embedding model's vectors as they are.
Chat completions
Terminus's OpenAI-compatible chat completions take a model's address, such as @terminus/claude-sonnet-4-6, as the model to call.
Embeddings
Terminus's OpenAI-compatible embeddings take an embedding model's address as the model, and answer in OpenAI's Embeddings format: one vector per text, up to 512 texts a call, with dimensions to ask for fewer where the model shortens them.

Connect your own provider key in System Settings › Connections, and a call that goes to that provider uses it: you pay the provider directly, and Terminus charges no tokens.

For your coding agent

https://www.terminus.build/docs/SKILL.md