Beliebte Suchanfragen:

🇬🇧 English · Home · Deutsche Version

Deep Dive – Technical Details for Professionals

Understand the technology behind Fleet Navigator in detail


Architecture overview

Fleet Navigator is built on a modular microservices architecture that allows maximum flexibility and scalability. All components communicate via secure REST APIs and can be updated individually.

One codebase. Five platforms. No forced cloud.

                                              macOS (Apple Silicon), Windows (AMD64/ARM), Linux (AMD64/ARM) — local AI infrastructure with its own stack, several inference backends and edition-specific builds.

Core components

Inference Engine

  • llama.cpp, vLLM, MLX (Apple Silicon)
  • CPU and GPU inference depending on the hardware; edition-specific selection at build time
  • Speculative decoding with draft models for compatible model pairs
  • Slot-based parallelism for several loaded models

RAG Pipeline

  • Document parsing: PDF, DOCX, ODT, EPUB, FB2/FBZ, HTML, Markdown, TXT — e-book and office import without external dependencies, DRM-free formats
  • Chunking: sliding window, legal, timeline, map-reduce, code chunking.
  • Embeddings: BGE-M3
  • Vector DB: Go implementation (cosine + MMR reranking) on top of SQLite as storage, with a RAM cache.
  • Retrieval: top-K with reranking

API Layer

  • Native Go HTTP server with modular route registration
  • More than 500 endpoints in 34 domain modules
  • Streaming: server-sent events for chat, benchmark, PDF analysis and document import
  • WebSocket: for Maat/Navigator sync and multi-device synchronisation
  • Authentication: token + session + CSRF, optional Maat pairing via Ed25519
  • Build-tag–based edition differentiation (Light/Maat/Enterprise)


Supported model formats

FormatDescriptionOptimisationPerformance
GGUFNative llama.cpp format4-bit to 8-bit quantisationExcellent on CPU
MLXApple’s own engine
 4-bit / 6-bit / 8-bit (optionally BF16 uncompressed)

Maximum speed on Apple Silicon (Metal + unified memory) 

Performance metrics

Benchmark results (Mac Mini M4 Pro, 64GB RAM, 30B model):

  • Reference setup: Mac Mini M4 Pro, 64 GB unified memory, Qwen3-30B-A3B-Instruct (MoE, 30B parameters / 3B active) as MLX 4-bit
  • Tokens per second: 28–32 (generation, thanks to MoE)
  • First-token latency: 120–180 ms for a short prompt with cache; 200–500 ms for a cold 1k prompt
  • Context window: 128K technically available, 64K as the stable limit in practice (KV cache + model must fit into unified memory)
  • Continuous batching: up to 4 parallel requests depending on model size
  • Total memory required: ~17 GB model file, ~40 GB incl. 64K KV cache and activations

Security & compliance

Encryption

  • TLS 1.2/1.3 for data in transit, with perfect forward secrecy (ECDHE cipher suites)
  • Ed25519 pairing for device connections (Maat/Navigator sync)
  • Data protection at rest via the operating system’s own disk encryption (FileVault / BitLocker / LUKS) — standard on modern devices

Audit & Logging

  • Security-relevant actions are logged (safety refusals, licence operations, authentication)
  • SIEM-ready log files (Filebeat/Fluentd compatible)
  • Automatic log rotation

Compliance

  • GDPR by architecture: in the default configuration data stays local, with no external processor
  • Suitable for practice compliant with Section 203 German Criminal Code (StGB) (lawyer/doctor/tax adviser); professional responsibility remains with the practitioner
  • Audit-ready: all security-relevant decisions can be traced in the log
  • No cloud dependency in the default configuration (Captain edition: cloud API that can optionally be enabled, opt-in only)

Do you have further technical questions? Contact our developer team for a deep dive.

Book your appointment