Loom / Guides / vLLM and llama.cpp

Local · Self-hosted

vLLM and llama.cpp

Connect to a vLLM or llama.cpp server you run yourself.

Overview

Both servers speak the OpenAI API. The one thing to get right is tool calling: without the right flags the model just writes text where a tool call should go.

Set it up

  1. 1
    vLLM: start it with vllm serve <model> --host 0.0.0.0 --enable-auto-tool-choice --tool-call-parser <parser>. The parser must match your model, for example llama3_json for Llama 3.1 and 3.3, hermes for Qwen, or mistral. The default port is 8000.
  2. 2
    llama.cpp: start it with llama-server -m model.gguf --jinja --host 0.0.0.0 --port 8080. --jinja turns on the model’s chat template, which tool calling needs.
  3. 3
    Allow the port through your computer’s firewall for your private network, and note the computer’s address on the Wi-Fi.

Cost. Free. You use your own hardware.

Add it to Loom

  1. 1
    Open Loom, then Providers in the menu, and choose Add provider.
  2. 2
    Pick vLLM or llama.cpp (the search box finds it).
  3. 3
    Fill in the fields in the table below.
  4. 4
    Choose Connect. Loom checks the key and says how many models it found. If it cannot check, it says so, and you can Save anyway and try a message.
  5. 5
    Open the model chip below the message box and pick a model.
  • Endpoint: http://192.168.1.20:8000/v1 for vLLM, or http://192.168.1.20:8080/v1 for llama.cpp (your computer’s address).
  • Add a key only if you started the server with --api-key.
Field in LoomWhat to putNeeded
API keyOnly if the server or service asks for oneOptional
Endpointhttp://localhost:8000/v1Pre-filled
NameAnything you like. It is shown in the model pickerOptional

At a glance

Appears in Loom asvLLM or llama.cpp
Sign-inNone, or a key if you start the server with one
ModelsLoom fetches the list from your account
Tools (connectors, search)If started as above
PhotosDepends on the model
PDFsNo

If something goes wrong

What you seeWhat to do
The model writes tool calls as plain textStart the server with the tool flags above.
Can’t connectCheck the port, the --host 0.0.0.0 flag and the firewall.

Official documentation