Skip to content

YantrikDB — Cognitive Memory Database for AI Agents

Engine v0.14.0 · lexical fusion, reranked recall & the explain surface
Memory that reasons about what it knows.

Not a vector store. A cognitive memory database — temporal decay, semantic consolidation, contradiction detection, and a knowledge graph — in one embeddable Rust engine, a Python package, an MCP server, or a Raft-replicated cluster.

AGPL open source Rust engine · Python + MCP ~3s leader-kill recovery
Installpip install yantrikdbcargo add yantrikdbdocker pull ghcr.io/yantrikos/yantrikdbServer starsEngine starsMCP stars

YantrikDB ships as three independently-released pieces, and they carry different version numbers on purpose — the engine you embed, the server you deploy, and the MCP bridge your agent talks to each move on their own release train.

ComponentVersionInstallWhat it is
Core engineyantrikdb0.14.0pip install yantrikdb
cargo add yantrikdb
The embeddable memory engine (Rust + Python bindings). Semantic recall, knowledge graph, contradiction detection, temporal decay, consolidation. 0.13 added BM25 lexical fusion + cross-encoder reranked recall and the explain surface; 0.13.2–0.13.4 hardened encryption at rest. This is what most people mean by “YantrikDB”.
Serveryantrikdb-server0.15.1docker pull ghcr.io/yantrikos/yantrikdbThe network database that wraps the engine: multi-tenant HTTP API, YRP native replication (~3s leader-kill recovery), server-side clustered packs, encryption at rest.
MCPyantrikdb-mcp0.14.0pip install yantrikdb-mcpThe MCP server — plugs the engine into Claude Code, Cursor, Windsurf, Copilot, and any MCP-compatible agent. Runs embedded, or points at a server.

If you saw one version quoted somewhere and a different one elsewhere, that’s these three lines being read as one number. They aren’t. The server is at 0.15.1; the engine is at 0.14.0; the MCP bridge is at 0.14.0.


Every AI memory solution does the same thing:

Store everything. Embed. Retrieve top-k. Inject into context. Hope it helps.

That doesn’t model how memory works. It treats all memories as equal. Old memories never fade. Contradictions are never detected. Nothing is ever consolidated. The AI never proactively remembers anything.

YantrikDB fixes all of this.

Relevance-Conditioned Scoring

Relevance gates every other signal multiplicatively. A perfectly relevant old memory surfaces. An irrelevant high-importance memory doesn’t. This is the key insight — patented and proven.

Cognitive State Graph

Typed nodes (beliefs, goals, intents, preferences) with typed edges (supports, contradicts, causes, predicts). Your AI doesn’t just remember — it reasons about what it knows.

Autonomous Cognition

Consolidation merges related memories. Conflict detection flags contradictions. Pattern mining discovers recurring themes. All automatic via db.think().

Proactive Triggers

Decaying memories, unresolved conflicts, emerging patterns — YantrikDB tells your AI when to act, grounded in real data. Not engagement farming.

Five Unified Indexes

Vector (HNSW), graph, temporal, decay heap, and key-value — all in one embedded SQLite database. No server. No infrastructure. Just a file.

MCP Server

pip install yantrikdb-mcp — instant persistent memory for Claude Code, Cursor, Windsurf, and any MCP-compatible AI agent.

First-Class Skills

/v1/skills/{define, get, search, outcome, forget} — agent skills as a substrate primitive, not a convention. Strict shape validation, append-only outcome event log, schema-not-semantics design line. Ships in v0.8.11.

Cluster Mode

Multi-node deployment via YRP native replication (RFC 028) — leader election, quorum-durable writes, exactly-once keyed writes, linearizable reads, quarantine-not-wedge recovery, and beyond-GC engine backfill. Chaos-gated releases; a benchmarked leader kill under load re-elects and resumes writes in ~3 seconds. Witness quorum for safe 2-node HA. Live on a 4-member homelab cluster.

Bundled Embedder — now multilingual

pip install yantrikdb works out of the box — no ONNX, no pip install sentence-transformers. Default potion-base-2M static embedder ships with the engine. Optional potion-base-8M, potion-base-32M (~95% MiniLM), and potion-multilingual-128M — record in one language, recall in another (100+ languages, cross-lingual recall verified) — all via set_embedder_named(). v0.9.4 fixes the loader so the multilingual path is a one-call setup.


The real YantrikDB engine — Rust compiled to WebAssembly — running in your browser. No server. No API calls. Click through to see record(), recall(), relate(), and think() in action.


Hermes is an open-source agent runtime. YantrikDB is its cognitive memory — three layers that compose into the full ecosystem:

① The plugin (agent-side)

pip install yantrikdb-hermes-plugin → 3-line .env config → the agent autonomously calls yantrikdb_remember / yantrikdb_recall / yantrikdb_stats during conversations. Sub-millisecond on the embedded backend. Owner-scoping for multi-user Hermes gateways contributed by community member @wysie (v0.4.10).

Hermes plugin guide →

② The server (memory backend)

docker pull ghcr.io/yantrikos/yantrikdb:{v.server} → multi-tenant database with HTTP API, YRP native replication, automatic failover, encryption at rest, and server-side clustered packs (RFC 031). Switch the plugin’s YANTRIKDB_MODE=http and one agent’s memory becomes shared infrastructure for a fleet.

Server quickstart →

③ The dashboard (operator console for Hermes memory)

Built by community member @wysie: yantrikdb-hermes-dashboard — a FastAPI dashboard for browsing, configuring, and safely maintaining a Hermes agent’s YantrikDB memory: namespace/identity/space configuration, per-user memory toggles, recall debugger, contradiction review, entity graph visualiser, lifecycle housekeeping. Read-only browsing by default; Admin Mode opt-in for mutating ops, with an optional dashboard password.

Third-party code. Not maintained by the YantrikDB org. Audit before running against production data — see the security considerations on the guide.

Hermes dashboard guide →

v0.8.17 ships the unlock. Until v0.8.17 the dashboard only ran against an embedded SQLite file — one machine, one agent. The new /v1/identity-scope, /v1/memories, and /v1/memory/{rid} endpoints let it run against a clustered deployment, so an operator can inspect a multi-agent fleet’s shared memory from one URL. Token-derived auth means a tenant-pinned token sees exactly its namespace; a cluster-admin token sees the whole fleet. Read the dashboard guide →


Don’t take our word for it — see what the cognitive architecture actually produces. Scroll through the experiments below.

🤖 Multi-Agent Ops — Which Memory Is Stale?

Five AI agents watched the same Black Friday incident. Two were working from stale data. YantrikDB surfaced exactly which beliefs were live and which were 20 minutes out of date — with validity windows, source attribution, and confidence bands.

Belief management under contradiction. Not a vector store — a witness stand.

Read the full experiment →

🚗 Volkswagen Dieselgate

The car passed emissions — because it knew it was being tested. Twelve years of public record across five sources. Five polarity contradictions on one tuple.

Public claims vs internal engineering vs regulator findings vs DOJ plea, all preserved as coexisting claims.

Read the full experiment →

⚖ Legal Discovery — Testimony vs Logs

He said he never touched the repo. Badge, VPN, git, USB, and an email to the competitor’s recruiter all say he did. The sworn denial and the forensic record coexist in the claims ledger, with the contradiction as the query result.

Sworn testimony and machine evidence as first-class structured claims.

Read the full experiment →

💶 Wirecard — The €1.9B That Existed and Didn’t

The same number, reported across four sources, took four contradictory positions over six years. Eight polarity contradictions on one tuple. The temporal query flips the belief state between 2019 and 2020.

Financial forensics as contradiction reconstruction, not scandal reporting.

Read the full experiment →

📰 Investigative Journalism — Follow the Money

The candidate denied taking pharma money. Five hops through public registries — FEC, Delaware, PAC disclosures, industry classification — traced it anyway. The contradiction lives in the composition of sources, not any single one.

Entity-graph reconstruction across public records.

Read the full experiment →

🔍 The Rashomon Engine

Five witnesses to a data breach, some of them lying, plus badge and git logs as ground truth. The engine identifies the perpetrator, cites the exact lies, and explains its reasoning — with real queries into the claims ledger, not scripted narrative.

Memory as a reasoning substrate, not a search index.

Read the full experiment →

🏛 Watergate: What the Tapes Caught

Fifty years of declassified sources: Nixon’s public denials, the White House tapes, sworn Senate testimony. Six polarity contradictions on Nixon alone in one query — the work the Senate Watergate Committee spent two years building by hand.

Historical research as a tractable problem.

Read the full experiment →

🎭 Shakespeare: Bringing a Character Alive

207 first-person memories. 288 entities auto-extracted. Personality derived. 28 proactive triggers. Zero LLM calls. Then he wrote a letter to his wife — and the character came through because the memories made him specific.

Richer memory → richer character → richer output from any LLM.

Read the full experiment →


YantrikDB ships in three forms. Pick the one that fits your stack:

📦 Embeddable engine

Drop the Rust crate or Python package into your app. Zero servers, single-process, fastest possible. Best for desktop apps, agents that own their memory, and CLI tools.

Terminal window
cargo add yantrikdb
pip install yantrikdb

Quick Start →

🌐 Network database

Run yantrikdb serve and get a multi-tenant database with HTTP + wire protocol, replication, automatic failover, encryption at rest, and a psql-style REPL. Best for self-hosted agents, homelab clusters, and shared memory across services.

Terminal window
brew install yantrikos/tap/yantrikdb
docker pull ghcr.io/yantrikos/yantrikdb

Run the Server →

🔌 MCP server

Plug-and-play memory for Claude Code, Cursor, Windsurf and any MCP-compatible agent. 15 tools for remember/recall/relate/think. The fastest way to give an existing AI assistant persistent memory.

Terminal window
pip install yantrikdb-mcp

MCP Setup →

All three share the same underlying engine and convergent semantics. You can start embedded and migrate to clustered later — your data works the same way.


IndexWhat It DoesExample Query
Vector (HNSW)Semantic similarity search”What did the user say about work?”
GraphEntity relationships & reasoning”Who works at what company?”
TemporalTime-aware retrieval”What happened last Tuesday?”
Decay HeapImportance with biological time decayMemories fade like human memory
Key-ValueInstant fact lookup”User’s timezone is CST”

All five indexes query the same data. A single recall() call blends signals from all of them into one relevance-conditioned score.


Vector DBRAG PipelineYantrikDB
StorageFlat embeddingsChunked documentsTyped memories with metadata
RetrievalCosine top-kHybrid searchRelevance-conditioned scoring
TimeIgnoredIgnoredTemporal decay + recency
ContradictionsUndetectedUndetectedAutomatic conflict detection
ConsolidationNoneNoneAutonomous merging
ProactiveNeverNeverTrigger-based notifications
GraphSeparate systemNoneBuilt-in cognitive state graph

Benchmark: Token Savings vs File-Based Memory

Section titled “Benchmark: Token Savings vs File-Based Memory”

Benchmarked with 15 diverse queries across 4 scales. File-based memory (CLAUDE.md, memory files) loads everything into context every conversation. YantrikDB’s selective recall retrieves only the 3–5 memories relevant to the current task.

MemoriesFile-BasedYantrikDBSavingsPrecision
1001,770 tokens69 tokens96%66%
5009,807 tokens72 tokens99.3%77%
1,00019,988 tokens72 tokens99.6%84%
5,000101,739 tokens53 tokens99.9%88%

Selective recall cost is O(1). File-based memory cost is O(n).

At 500 memories, file-based memory already exceeds 32K context windows. At 5,000 memories, it doesn’t fit in any context window — not even 200K. YantrikDB stays at ~70 tokens per query with recall latency under 60ms. Precision improves with more data: the opposite of file-based memory, which degrades as context fills up.

Works with Claude Code, Cursor, Windsurf, Copilot, Kilo Code — any MCP-compatible agent. Run the benchmark yourself: python benchmarks/bench_token_savings.py


v0.13 — Retrieval that reads the way you meant it (engine, shipped)

Section titled “v0.13 — Retrieval that reads the way you meant it (engine, shipped)”

Released through v0.13.4. The retrieval overhaul, measured end-to-end on a production clone.

Lexical fusion + reranked recall

BM25 keyword scoring fused into the vector lane, then optional cross-encoder rerank. On a labeled production clone the MRR moved 0.519 → 0.600 → 0.762 — the first release past the embedding model’s own exact-cosine ceiling. A verbatim phrase is findable again instead of being averaged into a long record’s dominant topic.

The explain surface

recall(explain=True) returns the candidate pool — every record that entered selection, each with a stable id, the lanes that admitted it, its bm25 strength, and its rank before and after fusion. A lane that didn’t run says why (never_ran: expand_entities=false) rather than reporting an ambiguous zero. You can see exactly why a result came back — and why one didn’t.

Matched-window snippets + token diet

Recall returns the window that actually matched, not the whole record — ~58% fewer characters over a 400-hit run with the ranking unchanged, plus a relative score cutoff that trims the weak tail. The token win grows as your records get longer.

Encryption at rest, hardened

AES-256-GCM at rest with per-database keys. 0.13.2–0.13.4 sealed the write-ahead oplog and erased its freed pages (GHSA-84vx-5fgq-5p59) — “encrypted” now means no record content anywhere in the raw file, enforced by a byte-scan test that ships in CI with its own control.


v0.12 — The retrieval floor moved (engine, shipped)

Section titled “v0.12 — The retrieval floor moved (engine, shipped)”

Released as v0.12.1.

Long records, chunked

Records longer than an embedder window are chunked and indexed per-window, so a long note is findable by any part of it instead of being blurred into a single averaged vector.

Relevance outranks the calendar

The additive recency wall came down; MMR diversity and re-tuned scoring replaced it. On a production clone, end-to-end MRR moved 0.054 → 0.541 — a relevant old memory now beats a fresh irrelevant one, which was the whole premise of relevance-conditioned scoring.


A pack is a sealed, signed, measured YantrikDB file. Your local model mounts it and gains the knowledge, rules and skills inside — then gives them back, leaving your memory byte-for-byte as it was.

The marketplace

First-party packs live at packs.yantrikdb.com — post-cutoff APIs, breaking releases, domain corpora — all official, signed, and measured against control questions before publication.

Browse packs →

Settings travel with the pack

Since v0.11.3 the sealed manifest carries the retrieval settings its author measured (recommended_top_k, recommended_min_similarity), signed with everything else — a consumer never has to guess a similarity floor for a corpus they didn’t write.

Clustered packs (server v0.14.0)

RFC 031: packs mount server-side on a replicated cluster, so a fleet of agents shares one mounted corpus instead of each carrying its own copy.

Ask the packs

The Ask button on this site and on the marketplace is yantrikdb-assistant — a support widget that answers only from mounted packs, and refuses before any model is called when the packs don’t cover the question.

The packs guide →


Released as v0.10.0. Every guarantee below is enforced by release-blocking trace contracts — 13 of 13 implemented — and proven by crash-kill tests that run in CI.

Most memory systems will happily store a lie, count a retry as new evidence, and leave half a write behind on a crash. v0.10 made the write path something you can trust:

Anti-laundering provenance gate

A write that declares source=inference cannot claim kind=fact without a verification basis — refused at write time, before any side effect, on every write path (record, batch, corrections, replication apply). An agent’s guess can no longer quietly become your ground truth.

Repetition is not corroboration

Durable idempotency keys: the same write retried returns the original record — no duplicate, no importance inflation, no certainty bump — even mid-crash, even under full backpressure. A different payload under the same key is a typed conflict, never a silent merge.

Losers don't move state

A rejected write — backpressure, gate refusal, failed transaction — leaves the engine byte-for-byte untouched: no row, no oplog entry, no calibration drift, no counters. Proven by crash-kill tests that run in CI, not by comments.

Upgrading from 0.9? Two behavioral changes to know: recall() now excludes superseded records by default (pass include_superseded=True for history access — current truth is what your agent should act on), and record_text/record_batch now normalize blank namespaces exactly like record() always did. Errors worth branching on now arrive as typed exceptions (yantrikdb.IdempotencyConflict, .Backpressure, .RecallContended, …), all subclassing RuntimeError so existing handlers keep working.

Next up: attachable memory & skill packs — download a domain’s knowledge and skills as a standalone pack, mount it into your agent’s memory, remove it cleanly when done. The v0.10 machinery (provenance identity, import-by-origin, diff-aware idempotent updates) is the pack engine.


U.S. Patent Application No. 19/573,392 (filed March 2026) — covers relevance-conditioned scoring, the cognitive state graph, and the unified system architecture.

Open source under AGPL-3.0. The patent protects the methods, not the code. Use it freely. Read more →


ComponentDescriptionLicense
yantrikdbCognitive memory engine (Rust + Python bindings)AGPL-3.0
yantrikdb-serverMulti-tenant network database with replication, auto-failover, encryptionAGPL-3.0
yantrikdb-witnessVote-only daemon for 2-node Raft cluster failoverAGPL-3.0
yantrikdb-protocolWire protocol codec (frames, opcodes, MessagePack)AGPL-3.0
yqlInteractive REPL client (like psql for cognitive memory)MIT
yantrikdb-mcpMCP server for Claude Code, Cursor, Windsurf & moreMIT
yantrikdb-hermes-pluginHermes Agent plugin — embedded YantrikDB, self-maintaining memory, owner-scoped multi-user mode, agent-authored skill substrateMIT
CortexOpenClaw/ClawDBot plugin — personality traits, bond evolution, context assemblyMIT

Distribution: crates.io · Docker Hub (GHCR) · Homebrew tap · PyPI

Open source. Get started →