Projects

Personal / Open Source

BARA — Layered Memory Architecture

Behavior-adaptive retrieval for stateful conversational AI — four memory tiers behind 41 deterministic decision gates instead of always-on RAG.

Key results

  • 41 deterministic decision gates
  • 50+ tunable thresholds
  • 5,000+ chunk evaluation corpus
  • 358 automated tests

Fig. 1 BARA — retrieval as a decision

Read the sequence
  1. Four memory tiers: research memory, conversational state, semantic profile and episodic memory.
  2. Each message is classified for intent, conversation type and behavioural shift, then threaded to a topic.
  3. The policy's 41 deterministic gates decide whether to retrieve and from which tiers.
  4. An acknowledgement-style message retrieves nothing and goes straight to the LLM.
  5. A message that refers back to earlier decisions opens only the research and episodic tiers.
  6. Selected context is passed to the multi-provider LLM client, which produces the response.

Retrieval is a decision made by deterministic gates, not a reflex on every message.

Problem

A research prototype and reference implementation built around one question: can a structured, multi-tier memory architecture with deterministic retrieval gating measurably outperform standard sliding-window RAG in multi-turn conversations?

Production LLM systems tend to hit the same wall: context is a sliding window, “memory” is a buffer, and retrieval is an always-on reflex. BARA replaces that buffer with structured memory and makes retrieval a decision.

Architecture

flowchart TB
  Msg[User message] --> Behavior[behavior_engine<br/>intent · conversation type · shift]
  Behavior --> Topic[topic_threading]
  Topic --> Policy{policy.py<br/>41 gates · 50+ thresholds}
  Policy -->|selected tiers only| Retrieval[Selective retrieval]
  Policy -->|no retrieval| LLM
  Retrieval --> LLM[Multi-provider LLM client]
  LLM --> Response

  subgraph tiers [Memory tiers]
    R[Research memory]
    C[Conversational state]
    S[Semantic profile]
    E[Episodic memory · pgvector]
  end
  Retrieval -.-> tiers

Four memory tiers

TierScopeHolds
Research memoryPermanent, cross-threadDecisions, conclusions and hypotheses extracted from conversation, linked through a concept graph (research_insights, concept_links)
Conversational statePer conversationTone, precision mode, repetition patterns, active topic threads (conversation_state, conversation_threads)
Semantic profilePermanent, per userIdentity, preferences and expertise domains, refined over time without explicit configuration
Episodic memoryPermanentPast interactions with vector embeddings in pgvector for semantic similarity retrieval

What I Built

Cognitive subsystems

  • behavior_engine.py — classifies every message for intent, conversation type and behavioural shift before any retrieval decision.
  • topic_threading.py — tracks active topic threads across turns, detecting shifts and threading references back to prior context.
  • research_memory.py — background extraction of insights, hypotheses and concept links as conversations happen.
  • conversation_state.py — per-session behavioural state: tone calibration, repetition guard, precision mode.
  • policy.py — 41 deterministic decision gates with 50+ tunable thresholds controlling exactly when and what gets retrieved.

LLM orchestration

  • Multi-provider LLM client for Cerebras, OpenAI and Anthropic with automatic fallback.
  • Intent classification and response generation as separate pipeline stages.
  • profile_detector.py for passive expertise and preference inference.
  • A hook system (hooks.py, four hooks) that injects subsystem outputs into the prompt pipeline at the right moment.

Engineering Decisions

  • Deterministic gates, not a learned router. Retrieval is controlled by explicit gates with named thresholds, so every decision can be inspected and tuned.
  • Classify before retrieving. Behaviour classification and topic threading run first; retrieval only sees their output.
  • Make every decision visible. The React frontend shows a real-time pipeline timeline per request: which gates fired, which memory tier was hit, and what was retrieved.

Testing / Validation

  • A/B experiment framework comparing BARA against standard RAG over a 5,000+ chunk knowledge base: 52 complete IETF RFCs plus 14 technical documents.
  • 50-query evaluation suite with LLM-as-a-Judge relevance scoring.
  • 358 automated tests covering each subsystem independently.
  • cli.py for inspecting and querying cognitive state between sessions.

Technology

Python 3.12 · FastAPI 0.115+ · PostgreSQL · pgvector · numpy · React · Vite · Tailwind · Vercel AI SDK · Zustand · Docker

Jump to