An Ollama server running on our own machine. This page is open to everyone; every API call needs an API key. It speaks both the Ollama API and the OpenAI API, so most tools plug in by changing one URL.
qwen2.5-coder:3bAuthorization: Bearer bk_…Keys start with bk_. Ask the BRIMIND owner for one. Each app or person gets its own key, so one can be switched off without touching the others.
Keep it out of your code: put it in an environment variable.
export BRIMIND_AI_KEY="bk_…"
If this answers with a list of models, the machine is awake and reachable. It is the only data URL that needs no key.
curl ORIGIN/models.json
Use the OpenAI-style endpoint: it is what most libraries and tools expect. The answer is in
choices[0].message.content. Ready-made code for curl, JavaScript, Python and the OpenAI SDK is under Use it.
The server remembers nothing between calls. To continue a chat, send the whole history each time, oldest first:
"messages": [
{"role": "system", "content": "You answer in French, briefly."},
{"role": "user", "content": "Bonjour !"},
{"role": "assistant", "content": "Bonjour ! Comment puis-je aider ?"},
{"role": "user", "content": "Donne-moi 3 idées de slogan."}
]
The system message sets the tone and rules for the whole conversation.
Add "stream": true. With the OpenAI endpoint you receive data: {…} lines (Server-Sent Events) ending with data: [DONE];
with the Ollama endpoint /api/chat you receive one JSON object per line, the last one with "done": true.
Streaming shows the first words in under a second, instead of waiting for the full reply.
Useful fields in the request body:
| Field | What it does | Typical |
|---|---|---|
temperature | Lower = precise and repeatable, higher = creative | 0.2 facts · 0.8 ideas |
max_tokens | Upper limit on the reply length (≈ ¾ of a word per token) | 300 |
response_format | {"type":"json_object"} forces valid JSON; also say "answer in JSON" in the prompt | for data extraction |
| Topic | Detail |
|---|---|
| Context | Up to 32 000 tokens per request (prompt + history + reply), about 24 000 words. Longer input is cut from the start. |
| Speed | About 30–40 tokens/s for short prompts. A very long prompt is read first, so the first word can take several seconds. |
| One at a time | Requests are answered one after another. If several arrive together, the others wait their turn: set a client timeout of 120 s. |
| Availability | It runs on a laptop. If the laptop sleeps, calls fail with 502; retry later. |
| Best at | Code, short answers, rewriting, extracting data to JSON, in French and English. Weaker in Arabic and on long reasoning. |
| Privacy | Prompts are processed on our machine and are not sent to any AI company. They are not stored after the reply. |
| Code | Meaning | Fix |
|---|---|---|
401 | No key sent | Add Authorization: Bearer <key> |
403 | Key wrong or switched off | Check it has no spaces; ask for a new one |
404 "model not found" | Model name typo | Use exactly qwen2.5-coder:3b |
502 | The machine is asleep or restarting | Retry in a minute |
| Timeout | Long prompt, or another request was being answered | Raise the timeout to 120 s, or use streaming |
| Endpoint | Method | Key |
|---|---|---|
/ this page and its guide | GET | open |
/models.json the model list shown on this page | GET | open |
/api/tags · /v1/models · /api/version · /api/ps | GET | required |
/v1/chat/completions · /v1/completions · /v1/embeddings | POST | required |
/api/chat · /api/generate · /api/embed | POST | required |
| Name | Size | Parameters |
|---|---|---|
| loading… | ||