🇬🇧 English · Home · Deutsche Version
Deep Dive – Technical Details for Professionals
Understand the technology behind Fleet Navigator in detail
Architecture overview
Fleet Navigator is built on a modular microservices architecture that allows maximum flexibility and scalability. All components communicate via secure REST APIs and can be updated individually.
One codebase. Five platforms. No forced cloud.
macOS (Apple Silicon), Windows (AMD64/ARM), Linux (AMD64/ARM) — local AI infrastructure with its own stack, several inference backends and edition-specific builds.
Core components
Inference Engine
- llama.cpp, vLLM, MLX (Apple Silicon)
- CPU and GPU inference depending on the hardware; edition-specific selection at build time
- Speculative decoding with draft models for compatible model pairs
- Slot-based parallelism for several loaded models
RAG Pipeline
- Document parsing: PDF, DOCX, ODT, EPUB, FB2/FBZ, HTML, Markdown, TXT — e-book and office import without external dependencies, DRM-free formats
- Chunking: sliding window, legal, timeline, map-reduce, code chunking.
- Embeddings: BGE-M3
- Vector DB: Go implementation (cosine + MMR reranking) on top of SQLite as storage, with a RAM cache.
- Retrieval: top-K with reranking
API Layer
- Native Go HTTP server with modular route registration
- More than 500 endpoints in 34 domain modules
- Streaming: server-sent events for chat, benchmark, PDF analysis and document import
- WebSocket: for Maat/Navigator sync and multi-device synchronisation
- Authentication: token + session + CSRF, optional Maat pairing via Ed25519
- Build-tag–based edition differentiation (Light/Maat/Enterprise)
Supported model formats
| Format | Description | Optimisation | Performance |
|---|---|---|---|
| GGUF | Native llama.cpp format | 4-bit to 8-bit quantisation | Excellent on CPU |
| MLX | Apple’s own engine | 4-bit / 6-bit / 8-bit (optionally BF16 uncompressed) | Maximum speed on Apple Silicon (Metal + unified memory) |
Performance metrics
Benchmark results (Mac Mini M4 Pro, 64GB RAM, 30B model):
- Reference setup: Mac Mini M4 Pro, 64 GB unified memory, Qwen3-30B-A3B-Instruct (MoE, 30B parameters / 3B active) as MLX 4-bit
- Tokens per second: 28–32 (generation, thanks to MoE)
- First-token latency: 120–180 ms for a short prompt with cache; 200–500 ms for a cold 1k prompt
- Context window: 128K technically available, 64K as the stable limit in practice (KV cache + model must fit into unified memory)
- Continuous batching: up to 4 parallel requests depending on model size
- Total memory required: ~17 GB model file, ~40 GB incl. 64K KV cache and activations
Security & compliance
Encryption
- TLS 1.2/1.3 for data in transit, with perfect forward secrecy (ECDHE cipher suites)
- Ed25519 pairing for device connections (Maat/Navigator sync)
- Data protection at rest via the operating system’s own disk encryption (FileVault / BitLocker / LUKS) — standard on modern devices
Audit & Logging
- Security-relevant actions are logged (safety refusals, licence operations, authentication)
- SIEM-ready log files (Filebeat/Fluentd compatible)
- Automatic log rotation
Compliance
- GDPR by architecture: in the default configuration data stays local, with no external processor
- Suitable for practice compliant with Section 203 German Criminal Code (StGB) (lawyer/doctor/tax adviser); professional responsibility remains with the practitioner
- Audit-ready: all security-relevant decisions can be traced in the log
- No cloud dependency in the default configuration (Captain edition: cloud API that can optionally be enabled, opt-in only)
Do you have further technical questions? Contact our developer team for a deep dive.