typecastlm

One address, the Jev API in front, and the model behind it in one of three places. Pick a scheme, fill in where, apply: the old one keeps answering until the new one is ready.

…

1 · Weights here

clientClient() typecastlm-serve prompt · head · temperatures trunk on this GPUweights from the Hub, torch

Downloads the checkpoint from Hugging Face and runs it in this process. Needs a GPU with 16 GB and the [server] extra. Returns logits and the calibration.

2 · An embedding server

clientClient() typecastlm-serveprompt · headtemperaturesno torch vector llama-serverGGUF trunkor vLLM, any GPU

The trunk runs as a GGUF in llama-server (or vLLM over /v1/embeddings, unnormalised); this process applies the head. Same numbers within 0.014; this box needs no GPU.

3 · Jev, proxied

clientClient() typecastlm-servetranslates tfu, levelspasses the restno model here https api.typesafe.aiJevtheir key, their bill

Every request goes to TypeSafe's Jev under their key. Probabilities come back and no logits; tfu is asked as a choice with a third option and says native: false.

Now serving

…

The same facts as GET /health. Settings that worked are written to the config file and read before the environment at the next start, so what is chosen here is what the box comes back with.