Models
Last updated
A model is something that talks, draws pictures, writes music, reads text aloud, writes down what is said, turns text into vectors or makes videos, published at an address like @you/legal-llama. Terminus's own models live under @terminus; yours runs on your own server, and Terminus calls it there. Everyone uses a model the same way: by its address.
A model's two files#
A model is a folder with terminus.json, its header, and model.json, which says where it runs and how much it reads. Its name, description, prices, who can see it and its key are set on the web, never in a file.
// terminus.json
{ "kind": "model", "modality": "chat", "id": "@you/legal-llama", "version": "0.0.1" }
// model.json
{ "server": "https://llm.example.com/v1", "model_name": "legal-llama-3-70b", "context": 131072 }- modality
chat,image,music,speech,transcription,embeddingorvideo: what it does. It is fixed for the life of the address.- server
- The base URL Terminus calls, public
https, at each kind's own path in OpenAI's format: a chat model at{server}/chat/completions, an image model at/images/generations, a speech model at/audio/speech, a transcription model at/audio/transcriptions, an embedding model at/embeddingsand a video model at/videos. A music model is called at/music. - model_name
- The name your server expects in each request's
model. - context
- Chat: the tokens it reads in one call, from 1,024 to 10,000,000. Embedding: the tokens it reads of one text, from 128 to 1,000,000.
- output
- Chat: the most tokens one answer may write. 8,192 unless you say.
- images, tools, reasoning
- Chat: whether it reads pictures (no unless you say), calls tools (yes unless you say) and takes a reasoning effort (no unless you say).
- voices
- Speech: the voices it reads in, 1 to 64 names of letters, digits, spaces,
.,_or-. The first is its voice unless a call names another; a later release may add voices, never take one away. - dimensions
- Embedding: the length of each vector it answers, from 1 to 8,192. It is fixed for the life of the address: vectors of two lengths never compare.
- shortens
- Embedding: whether it answers fewer dimensions when a call asks for them (no unless you say).
An image, music, transcription or video model needs nothing in model.json beyond where it runs. A speech model names its voices, and an embedding model its vectors:
// model.json, a speech model
{ "server": "https://voice.example.com/v1", "model_name": "narrator-1", "voices": ["alloy", "nova"] }
// model.json, an embedding model
{ "server": "https://vectors.example.com/v1", "model_name": "embed-1", "dimensions": 1536, "context": 8191, "shortens": true }Make the model on the web first — Create on the desktop at /os, then Model — and fill in its Model form, or push the two files from a working copy with terminus push. Either way the draft is checked as it is saved, and a refusal says exactly what is wrong.
Its key, or a signed call#
If your server wants a key, set it as the model's API_KEY — with the key on its page, or terminus secrets set API_KEY. Terminus sends it as Authorization: Bearer with every call, and never shows it again. A key never goes in a file.
With no key, Terminus signs each call instead, once you show the server is yours: serve the token the page gives you at https://your-server/.well-known/terminus-model-verification, then press Verify.
What Terminus sends your server is only the request, never who it is for.
Prices per unit#
A model is priced per unit of what it does, in credits (a credit is a dollar), and its next release carries the prices you set:
- Chat
- Input and output per million tokens; cache read and cache write per million tokens too, or blank, when cached tokens bill as input.
- Image
- Per picture returned.
- Music
- Per minute of audio, billed by the second.
- Speech
- Per million characters read.
- Transcription
- Per minute of the audio, billed by the second.
- Embedding
- Per million tokens read.
- Video
- Per second of video: the video's own length, at most the seconds asked.
A model cannot publish before its prices are set. Each call people pay for earns you credits.
Use a model#
Every kind names a model by its address:
- Norbert
- Pick another chat model for a conversation from the model menu on his line. A model that runs on someone's own server says where before your first conversation goes there. He draws pictures, writes music, reads text aloud, transcribes recordings and makes videos with the models Terminus sets for him, asks before the first one costs you anything in a conversation, and keeps the files he makes in your Downloads.
- Agents
modelsin an agent's terminus.json names the model it talks with, andtools: ["model:@terminus/gpt-image-2"]gives it a picture, music, speech, transcription or video model as a tool:generate_image,compose_music,generate_speech,transcribe_audioorgenerate_video, a second of a kind<slug>_generate_image. A chat or embedding model is no tool. The person talking to the agent pays.- Apps
modelsin the app's schema.ts names each model it calls by address; its handle'schatanswers a chat model's completion andembedan embedding model's vectors, whilegenerate,compose,speakandtranscribeanswer the job that makes the pictures or the video, the music, the speech or the transcript. See Models in the SDK reference.- Services
- An operation lists the models it calls in
modelsand calls them withctx.models.call(address, input). The service's owner pays, within its daily budget; a picture, a song, speech or a video comes back finished, its files in the answer, a transcript with its words, and an embedding model's vectors as they are. - Chat completions
- Terminus's OpenAI-compatible chat completions take a model's address, such as
@terminus/claude-sonnet-4-6, as the model to call. - Embeddings
- Terminus's OpenAI-compatible embeddings take an embedding model's address as the model, and answer in OpenAI's Embeddings format: one vector per text, up to 512 texts a call, with
dimensionsto ask for fewer where the model shortens them.
Connect your own provider key in System Settings › Connections, and a call that goes to that provider uses it: you pay the provider directly, and Terminus charges no tokens.
For your coding agent
https://www.terminus.build/