01 / 16
OdyssAI · 2026

Companion — Universal AI Client

Companion
the OdyssAI client.

One client. Every engine — local cluster, MLX, cloud API. Persistent memory, projects, skills, and agents.

Projects Memory Skills MCP Agents Voice
Projects & organized conversations 01
LightRAG memory across sessions 02
Skills, MCP servers & extensions 03
Hermes, Pi & ComfyUI agents 04
Voice, inference control, multi-user 05
Overview

One client.
Every engine.

Companion pairs with OdyssAI-X, Telemak, Ollama, LM Studio, vLLM, MLX, OpenRouter, Anthropic, and OpenAI. It never ships its own model — you decide what's behind the chat window.

Local cluster MLX · Mac OpenRouter Anthropic OpenAI Your model
03 / 16
Organisation

Projects

Group conversations into Projects. Each project shares a system prompt, memory toggles, and an optional wiki injected as context.

Move a chat into a project and it inherits that context instantly. Long-running work stays coherent across many conversations.

Projects · System prompt · Wiki
03
Projects view
04 / 16
Memory

Némo
remembers.

A LightRAG knowledge layer survives conversations. Retrieved semantically per turn — only what's relevant to your current question.

Three tiers: User · Project · Company. Each conversation freezes a memory snapshot at creation.

LightRAG · Semantic retrieval · Per-turn injection
04
Memory in chat
05 / 16
Inference

Engine
Pairing

Connect OdyssAI-X in Gateway mode for direct, low-latency inference. Switch to Advanced for named model slots or Expert for the full provider catalogue.

Easy — auto-routes. Advanced — 4 named slots. Expert — every model, full control.

Gateway · OdyssAI-X · LiteLLM
05
Inference settings
06 / 16
Fine-grained Control

Custom
Inference

Temperature, top-P, top-K, max tokens, thinking mode, reasoning effort — all tunable per conversation. Save named presets and reload them instantly.

System prompt editor with import/export. The Custom tab overrides the engine defaults for power users.

Temperature · Top-P · Presets · System prompt
06
Custom inference settings
07 / 16
Smart Routing

Auto
Router

Inspects each incoming message and routes it to the best model pool automatically. Zero picker friction in Easy mode.

A chat-level router — distinct from CoeOS, the engine-side per-skill router used by agents. Two separate layers, each doing one thing.

Auto-route · Easy mode · Per-message
07
Auto Router settings
08 / 16
Extensions

Skills

Markdown instruction packages the agent loads on demand. The model always sees a compact catalog — and pulls the full body only when a skill matches your request.

Create by form, ask Némo to draft one in chat, or import any SKILL.md from the agentskills.io open library.

agentskills.io · SKILL.md · On-demand
08
Skills settings
09 / 16
Extensions

MCP
Servers

Wire in Notion, Linear, GitHub, Tavily, Obsidian, Filesystem — or any MCP-compatible server. Their tools merge into the agent's toolbox, prefixed to avoid collisions.

OAuth flows, bearer tokens, Streamable HTTP and SSE. Tools are gated on the per-conversation Agent mode toggle.

MCP · OAuth · Tools
09
MCP Servers
10 / 16
Agents tokens

Companion
as Brain

Mint an hms_… token and external agents — Cline, Claude Code, Cowork — call back into Companion's memory and tools over MCP.

Tokens are revocable, budget-capped, and scoped per agent. The same Némo memory your chat sees is exposed as a structured MCP endpoint.

hms_… tokens · MCP endpoint · Cline · Cowork
10
Agent Tokens
11 / 16
MCP API

The MCP
Brain

A full structured API for external agents to read and write Companion's state.

companion_remember Write a fact to Némo's memory
companion_search Semantic search in memory
companion_project_get Read project context
tools/list · /api/mcp · Full schema
11
MCP Brain tools
12 / 16
Coding agents

Hermes
& Pi

/hermes — real-machine agent. Reads files, writes files, runs shell, browses. Every action streams inline in the chat. Nothing is silent.

/pi — reflective TUI for slow thinking and planning. Omnigent — multi-agent orchestrator with plan → act → verify loop.

/hermes · /pi · Omnigent · Bridge
12
Hermes and Pi agents
13 / 16
Voice

Speak to
Companion

Push-to-talk on Space, or full Talk mode for hands-free conversations. Local OpenAI-compatible TTS + ASR, or Gemini Live for low-latency full-duplex.

Configure TTS endpoint, ASR endpoint, model name, and voice per user. Settings sync to every device you sign in to.

TTS · ASR · Space push-to-talk · Talk mode
13
Voice settings
14 / 16
Image Generation

ComfyUI
Imager

Type /comfyui in any chat. Pick a FLUX template, set dimensions, hit Generate — the image drops inline in the conversation.

Runs on your own GPU. The LLM can also call the bridge itself via comfyui_generate when an image would help the answer.

/comfyui · FLUX · Local GPU · Sovereign
14
ComfyUI Imager modal
15 / 16
Image Generation

Result
inline.

The image lands directly in the conversation — persisted as an attachment reference, visible on reload. Hover to save the original PNG.

Templates: image-rapide (fast, 12 steps), photo-article-tmb (full quality, 20 steps). Add your own by dropping a workflow JSON.

Inline · Saved · PNG · FLUX.1-dev
15
ComfyUI result in chat
16 / 17
Administration

Users &
Teams

Invite users, assign roles, and create teams that share a LightRAG memory collection. Guest tokens give time-limited, budget-capped chat sessions to outside collaborators.

Manage compute fleet, file sync, and model deployment from the same admin panel via the Inference engine.

Roles · Teams · Guest tokens · Admin
16
Users and Teams
OdyssAI · Companion

The complete
AI workspace.

Projects · Memory · Skills · MCP · Agents · Voice · Image — one client, your stack.

companion.odyssai.eu