Applicaties → Ollama
Private AI models, on your own server.
Ollama runs open-source large language models locally on your ReefOffice server. No data ever leaves your server to be processed by an external AI provider — private chat, document Q&A and autonomous AI tasks, all local.
What Ollama does
Ollama handles the AI compute — it downloads, loads and runs open-source language models (LLaMA, Mistral, Qwen, Phi, Gemma and hundreds more) directly on your server. It powers private chat through Open WebUI, autonomous AI tasks through Hermes Agent, and document Q&A on Private AI plans.
Why Ollama on ReefOffice
1 Local AI model execution
Ollama runs hundreds of open-source models on your server's CPU or GPU — LLaMA, Mistral, Qwen, Gemma, Phi, DeepSeek and more. Models are downloaded, cached and executed locally. No API calls to external providers, no data leaving your infrastructure.
2 Private chat and document Q&A
Open WebUI provides a ChatGPT-style interface powered by Ollama. Ask questions about your documents, get answers with source references, and chat privately — all without sending data to OpenAI or Anthropic.
3 Private AI answers
On Private AI plans, Ollama generates answers from your document corpus. Ask "which invoices are unpaid?" or "summarise the latest contract" — the answer is generated on your server from your indexed documents, with source references.
4 Model flexibility
Choose the model that fits your use case. Run a small, fast model for everyday chat, a multilingual model for document Q&A, or a larger model for complex reasoning. Switch between models with a single command.
5 GPU acceleration ready
Private AI plans include dedicated GPU capacity on ReefOffice infrastructure in Germany. For local-only setups, Ollama uses CPU with optimised inference — slower but entirely private. Optional external model keys can supplement local capacity.
6 Connected to the full stack
Ollama connects to Open WebUI for chat, Hermes Agent for autonomous tasks, and can be accessed via API for custom integrations. It runs as a NixOS native service — always available, no container overhead.
Technical detail
How it works
Ollama wraps llama.cpp for efficient local inference with support for GPU acceleration, quantised models (GGUF format) and prompt caching. On ReefOffice it runs as a NixOS native service. Models are downloaded from the Ollama library, cached locally, and served via a local API that Open WebUI and Hermes connect to.
How it fits your stack
Ollama powers Open WebUI for private ChatGPT-style chat and Hermes Agent for autonomous AI tasks. It can also be accessed via its local API for custom integrations. The service is internal and API-protected — never exposed directly to the internet.
Frequently asked questions
Will local AI be fast enough for daily use? +
For document search and indexing, yes — embedding generation with bge-m3 runs efficiently on CPU and completes within seconds for typical document volumes. For chat and AI answers, performance depends on model size and hardware. Small models (3-8B parameters) give responsive answers on modern CPU. For larger models or faster generation, Private AI plans include GPU capacity on ReefOffice infrastructure.
How much storage do AI models use? +
Models vary from 1.5 GB (small quantised) to 40+ GB (full-precision 70B models). ReefOffice pre-installs the recommended models for your plan (typically bge-m3 for embeddings and Qwen 2.5 7B for generation). Additional models are downloaded on demand and cached. Storage for vector embeddings depends on document volume — roughly 1 GB per 100,000 pages.
Can I use an external AI provider instead of local models? +
Yes. On plans without local GPU, you can connect an external AI provider key (OpenAI, Anthropic, Mistral AI, or any OpenAI-compatible API). Open WebUI and Hermes will use the external API for generation while keeping your data private.
Ready for private AI?
Book a demo and see how Ollama on ReefOffice gives you private AI capabilities — chat, document Q&A and local model execution — all on your own server.
Boek een demo