checking…

BRIMIND Local AI

An Ollama server running on our own machine. This page is open to everyone; every API call needs an API key. It speaks both the Ollama API and the OpenAI API, so most tools plug in by changing one URL.

Base URL (OpenAI)
Base URL (Ollama)
Modelqwen2.5-coder:3b
Auth headerAuthorization: Bearer bk_…

Guide, step by step

1

Get an API key

Keys start with bk_. Ask the BRIMIND owner for one. Each app or person gets its own key, so one can be switched off without touching the others. Keep it out of your code: put it in an environment variable.

export BRIMIND_AI_KEY="bk_…"
2

Check the server is up

If this answers with a list of models, the machine is awake and reachable. It is the only data URL that needs no key.

curl ORIGIN/models.json
3

Send your first message

Use the OpenAI-style endpoint: it is what most libraries and tools expect. The answer is in choices[0].message.content. Ready-made code for curl, JavaScript, Python and the OpenAI SDK is under Use it.

4

Keep a conversation

The server remembers nothing between calls. To continue a chat, send the whole history each time, oldest first:

"messages": [
  {"role": "system",    "content": "You answer in French, briefly."},
  {"role": "user",      "content": "Bonjour !"},
  {"role": "assistant", "content": "Bonjour ! Comment puis-je aider ?"},
  {"role": "user",      "content": "Donne-moi 3 idées de slogan."}
]

The system message sets the tone and rules for the whole conversation.

5

Stream the answer word by word

Add "stream": true. With the OpenAI endpoint you receive data: {…} lines (Server-Sent Events) ending with data: [DONE]; with the Ollama endpoint /api/chat you receive one JSON object per line, the last one with "done": true. Streaming shows the first words in under a second, instead of waiting for the full reply.

6

Tune the reply

Useful fields in the request body:

FieldWhat it doesTypical
temperatureLower = precise and repeatable, higher = creative0.2 facts · 0.8 ideas
max_tokensUpper limit on the reply length (≈ ¾ of a word per token)300
response_format{"type":"json_object"} forces valid JSON; also say "answer in JSON" in the promptfor data extraction

Limits and good to know

TopicDetail
ContextUp to 32 000 tokens per request (prompt + history + reply), about 24 000 words. Longer input is cut from the start.
SpeedAbout 30–40 tokens/s for short prompts. A very long prompt is read first, so the first word can take several seconds.
One at a timeRequests are answered one after another. If several arrive together, the others wait their turn: set a client timeout of 120 s.
AvailabilityIt runs on a laptop. If the laptop sleeps, calls fail with 502; retry later.
Best atCode, short answers, rewriting, extracting data to JSON, in French and English. Weaker in Arabic and on long reasoning.
PrivacyPrompts are processed on our machine and are not sent to any AI company. They are not stored after the reply.

Errors

CodeMeaningFix
401No key sentAdd Authorization: Bearer <key>
403Key wrong or switched offCheck it has no spaces; ask for a new one
404 "model not found"Model name typoUse exactly qwen2.5-coder:3b
502The machine is asleep or restartingRetry in a minute
TimeoutLong prompt, or another request was being answeredRaise the timeout to 120 s, or use streaming

What needs a key

EndpointMethodKey
/ this page and its guideGETopen
/models.json the model list shown on this pageGETopen
/api/tags · /v1/models · /api/version · /api/psGETrequired
/v1/chat/completions · /v1/completions · /v1/embeddingsPOSTrequired
/api/chat · /api/generate · /api/embedPOSTrequired

Send the key as Authorization: Bearer <key> or x-api-key: <key>. No key → 401, wrong or disabled key → 403.

Use it

Try it

Models on this server

NameSizeParameters
loading…

Use the model above. The larger ones are stored here but are too heavy for this machine's memory budget.