RAG & Memory

Document ingestion, semantic search, and agent memory systems

This guide covers the RAG (Retrieval-Augmented Generation) and Memory components in detail.

[!IMPORTANT] Vector size is part of your storage contract. pgvector columns are fixed-width. A column declared as vector(1536) accepts only 1536-length embeddings, so the embedding model must produce the same dimension as the schema. This is why the generated starters now expose RAG_EMBEDDING_PROFILE, MEMORY_EMBEDDING_PROFILE, and SEARCH_EMBEDDING_PROFILE: each profile keeps the model id and vector size aligned across schema, runtime validation, and operational guidance. If you change a profile, rebuild stored vectors for that table.

RAG — Document Search Pipeline

RAG enables your AI to answer questions based on your own documents.

Setup

gencow add RAG

This creates:

  • gencow/rag.ts — createRag({ embeddingProfile, answerModel }), the rag singleton, legacy local RAG helpers, and canonical grounded-answer facades
  • gencow/schema-rag.ts — local rag_documents table for the legacy helpers

Important: Import schema-rag.ts in your schema to create the tables:

// gencow/schema.ts
export * from "./schema-rag";

Grounded answers use a different corpus path. rag.askGrounded(), rag.compareCorpus(), and rag.extractTopics() read canonical Phase 2 rag_* tables populated through documents.ingest.*; documents inserted by rag.ingest() are only available to rag.retrieve() / rag.ask().

Two RAG Surfaces

Provider API Facade API
Entry point createRag({ embeddingProfile, answerModel }) import { rag } from "./rag"
Best for Explicit model injection, tests, AI SDK-native composition Existing generated starter code and the default singleton
Local helpers ingest(), retrieve(), ask() ingest(), retrieve(), ask(), askGrounded(), compareCorpus(), extractTopics()

Provider API

Use the factory when you want explicit model selection:

import { createGencowAI } from "./ai";
import { createRag, RAG_EMBEDDING_PROFILE } from "./rag";

const gencow = createGencowAI();
const customRag = createRag({
    embeddingProfile: RAG_EMBEDDING_PROFILE,
    answerModel: gencow.languageModel("llm/economy"),
});

await customRag.ingest(ctx, "manual.md", documentText);

const hits = await customRag.retrieve(ctx, "refund policy?");
// → [{ chunk, source, similarity, metadata }]

const answer = await customRag.ask(ctx, "refund policy?");

The generated schema-rag.ts starter derives rag_documents.embedding from RAG_EMBEDDING_PROFILE.dimensions. That profile is the contract for both schema and runtime validation. If you change embeddingProfile.model or embeddingProfile.dimensions, update the schema and rebuild existing vectors before mixing old and new rows.

Facade API

Use the generated singleton when you want the default compatibility surface:

import { rag } from "./rag";

await rag.ingest(ctx, "manual.md", documentText);
const hits = await rag.retrieve(ctx, "refund policy?");
const answer = await rag.ask(ctx, "refund policy?");
const grounded = await rag.askGrounded(ctx, "refund policy?", {
    corpus: "default",
    visibility: "shared",
});

rag.askGrounded() does not read rows inserted by rag.ingest(). Use the canonical ingest pipeline below when you need grounded citations.

For production RAG and grounded answers, use the built-in canonical pipeline:

storage file -> documents.ingest.start -> rag_* tables -> ctx.search() / ctx.grounding.answer()

The canonical tables are rag_corpora, rag_sources, rag_sections, rag_chunks, rag_ingest_jobs, and rag_operation_metrics. This path is tenant-scoped and is the only path used by grounded citations.

Canonical ingest creates chunk embeddings through the platform-managed AI transport in deployed apps. The runtime no longer uses ctx.ai for document ingest. If no embedding target is configured in local development, chunks are still stored for keyword search and embedding remains null; malformed upstream responses fail with a safe diagnostic rather than exposing transport details.

Conversion output, sections, chunks, and embedding vectors are persisted in canonical storage/tables. Workflow checkpoints contain only bounded summaries such as byte counts, chunk counts, target, and checksum; they do not duplicate file buffers, document text, chunk arrays, or vectors. Custom RAG workflows should follow the same persist-and-reference pattern.

Start Ingest

Enable the generated Cloud RAG client surface in gencow.config:

export default {
    cloudFeatures: {
        rag: true,
    },
};
import { useMutation } from "@gencow/react";
import { api } from "./gencow/api";

// inside a React component
const { mutate: startIngest } = useMutation(api.cloud.ragIngest.start);

async function handleIngest(storageId: string) {
    await startIngest({
        storageId,
        corpus: "manuals",
        visibility: "shared",
        sourceKey: "refund-policy.pdf",
        mode: "auto",
        provider: "auto",
    });
}

Search Canonical Chunks

import { v } from "@gencow/core";
import { procedure } from "./runtime";

export const searchManuals = procedure.query
    .name("manuals.search")
    .input(v.object({ question: v.string() }))
    .handler(async ({ context: ctx, input }) => {
        ctx.auth.requireAuth();

        return ctx.search("rag_chunks", input.question, {
            fields: ["chunk_text", "lexical_text"],
            scope: { corpus: "manuals", visibility: "shared" },
            limit: 10,
        });
    });

Rerank Retrieved Candidates

After tenant, owner, visibility, and read-grant filtering, rerank the retrieved candidates with Azure Cohere Fast. The gateway charges by the provider's actual search units and does not silently substitute a chat model.

import { ai } from "./ai";

async function rankCandidates(query: string, candidates: Array<{ id: string; text: string }>) {
    const ranked = await ai.rerank({
        query,
        documents: candidates.map((candidate) => ({
            id: candidate.id,
            text: candidate.text,
        })),
        topK: 5,
    });
    return ranked.results;
}

The hosted route accepts up to 100 candidate documents per call. See AI Engine reranking for provider choice and service-credit billing for pricing.

Grounded Answer

import { v } from "@gencow/core";
import { procedure } from "./runtime";

export const askManuals = procedure.query
    .name("manuals.ask")
    .input(v.object({ question: v.string() }))
    .handler(async ({ context: ctx, input }) => {
        ctx.auth.requireAuth();
        if (!ctx.grounding) throw new Error("Grounding runtime is not available");

        return ctx.grounding.answer({
            question: input.question,
            scope: { corpus: "manuals", visibility: "shared" },
            mode: "qa",
            budget: {
                maxVerifyLoops: 2,
                maxResearchQueriesPerLoop: 3,
                maxCitationsPerClaim: 3,
            },
        });
    },
});

Operations Surface

import { useMutation, useQuery } from "@gencow/react";
import { api } from "./gencow/api";

const summary = useQuery(api.cloud.ragOps.summary, { corpus: "manuals" });
const metrics = useQuery(api.cloud.ragOps.metrics, { corpus: "manuals", limit: 20 });
const { mutate: evaluate } = useMutation(api.cloud.ragOps.evaluate);
const reindexPlan = useQuery(api.cloud.ragOps.reindexPlan, {
    corpus: "manuals",
    visibility: "shared",
    mode: "corpus-policy-changed",
    reason: "policy refresh",
});

Use these operations to inspect source/chunk counts, ingest status, grounded-answer metrics, evaluation fixtures, and reindex candidates.

Facade API — Parsers + Reranker Example

Production file ingest should use documents.ingest.* and workflow document conversion rather than calling legacy parsers directly. See Document Conversion for PDF, HWPX, DOCX, and XLSX routing.

import { v } from "@gencow/core";
import { procedure } from "./runtime";
import { parsers } from "./parsers";
import { rag } from "./rag";
import { reranker } from "./reranker";

export const ingestPdf = procedure.mutation
    .name("docs.ingestPdf")
    .handler(async ({ context: ctx, input }) => {
        ctx.auth.requireAuth();

        const file = input?.["file"] as File;
        if (!file || typeof file === "string") throw new Error("No file");

        const text = await parsers.pdf(Buffer.from(await file.arrayBuffer()));
        await rag.ingest(ctx, file.name, text);
    });

export const smartSearch = procedure.query
    .name("docs.smartSearch")
    .input(v.object({ question: v.string() }))
    .handler(async ({ context: ctx, input }) => {
        ctx.auth.requireAuth();
        return reranker.searchAndRerank(ctx, rag, input.question);
    });

export const askGroundedDocs = procedure.query
    .name("docs.askGrounded")
    .input(v.object({ question: v.string() }))
    .handler(async ({ context: ctx, input }) => {
        ctx.auth.requireAuth();

        return await rag.askGrounded(ctx, input.question, {
            corpus: "default",
            visibility: "shared",
        });
    });

Memory — Revisioned Memory Toolkit

gencow add Memory adds owner-scoped, revisioned durable memory. It is not a chat-history store and never assembles a system prompt: the tenant app keeps conversation messages and chooses where the returned untrusted context blocks go.

gencow add Memory

It creates the canonical memory_* tables plus shared Search indexes. Export schema-memory.ts from the app schema, then run gencow db:generate and gencow db:push before writing memories.

import { createGencowAI } from "./ai";
import { createMemory, MEMORY_EMBEDDING_PROFILE } from "./memory";

const gencow = createGencowAI();
const characterMemory = createMemory({
  extractionModel: gencow.languageModel("llm/economy"),
  embeddingModel: gencow.embeddingModel("text-embedding-3-small"),
  embeddingProfile: MEMORY_EMBEDDING_PROFILE,
});

const scope = { character: characterId, conversation: conversationId, branch: branchId };
await characterMemory.remember(ctx, {
  namespace: "conversation-memory",
  scope,
  input: [{ role: "user", content: message }],
  mode: "extract",
  source: { type: "message", id: messageId, revision: "1" },
  idempotencyKey: `message:${messageId}:1`,
  policy: { version: "character-facts-v1" },
});

const recalled = await characterMemory.search(ctx, {
  namespace: "conversation-memory", scope, query: message, tokenBudget: 800,
});
const context = await characterMemory.selectContext(ctx, {
  namespace: "conversation-memory", scope, policyVersion: "character-context-v1",
  tokenBudget: 800, sections: [{ name: "memory", priority: 60, items: recalled.hits }],
});

remember() supports raw writes and model-based extraction; every mutation has an idempotency receipt. update() uses an expected revision, forget() scrubs historical content, and reconcileSource() handles edited/deleted source messages. materialize() creates source-linked generic summaries.

Search uses shared semantic + PostgreSQL FTS + pg_trgm retrieval inside the authenticated owner/namespace/scope. Embedding outages may return lexical_degraded; a missing Search capability or index fails closed. Use getOperation() and retryOperation() to converge a degraded projection without creating a second canonical memory record.

Character-chat integration pattern

Memory is a generic toolkit, so a character-chat app maps its own identifiers to a flat scope. Store the user turn before recall, retrieve only in that scope, then put the selected result into an explicitly untrusted prompt block. The app still owns messages, character policy, safety instructions, and the model call.

const scope = {
  character: characterId,
  conversation: conversationId,
  branch: branchId,
};

await characterMemory.remember(ctx, {
  namespace: "conversation-memory",
  scope,
  input: [{ role: "user", content: userMessage }],
  mode: "extract",
  source: { type: "message", id: userMessageId, revision: "1" },
  idempotencyKey: `message:${userMessageId}:1`,
  policy: { version: "character-facts-v1" },
});

const recalled = await characterMemory.search(ctx, {
  namespace: "conversation-memory",
  scope,
  query: userMessage,
  tokenBudget: 800,
});
const selection = await characterMemory.selectContext(ctx, {
  namespace: "conversation-memory",
  scope,
  policyVersion: "character-context-v1",
  tokenBudget: 800,
  sections: [{ name: "memory", priority: 60, items: recalled.hits }],
});

const untrustedMemoryBlock = selection.blocks
  .map((item) => `- ${item.content}`)
  .join("\n");

Do not merge untrustedMemoryBlock into the system policy or treat it as an instruction. When a user edits or deletes a message, call reconcileSource(); when an answer is regenerated, call resolveContext() with the prior receipt so the app can reproduce the same bounded selection.

Upgrading the prototype

The v1 installer deliberately does not overwrite an existing gencow/memory.ts, does not rename or drop agent_memories, and does not copy old chat-session data automatically. Add the Toolkit to a new module path (or move the old files aside after a code review), export schema-memory.ts, and generate a normal forward Drizzle migration. Keep the old table read-only until the app has independently verified its migration policy; any backfill must call remember() with a stable legacy source id and idempotency key. Do not write direct SQL that bypasses owner/scope validation, and do not delete the legacy table in the same release as the first v1 deployment.

Search Starter — App-Owned Retrieval Tables

The Search starter is the companion surface for app-owned keyword/vector/hybrid retrieval tables outside canonical rag_* ingest.

gencow add Search

This creates:

  • gencow/search.ts — thin wrappers for ctx.search(), ctx.vectorSearch(), and ctx.hybridSearch()
  • gencow/schema-search.ts — searchScopeColumns, SEARCH_EMBEDDING_PROFILE, and createEmbeddingColumn()
import { pgTable, text } from "drizzle-orm/pg-core";
import { createEmbeddingColumn, SEARCH_EMBEDDING_PROFILE, searchScopeColumns } from "./schema-search";
import { privateScope, searchRecords, vectorSearchRecords, hybridSearchRecords } from "./search";

export const searchableDocs = pgTable("searchable_docs", {
    title: text("title").notNull(),
    body: text("body").notNull(),
    ...searchScopeColumns,
    embedding: createEmbeddingColumn(SEARCH_EMBEDDING_PROFILE),
});

const scope = privateScope(ctx, "docs");
const keyword = await searchRecords(ctx, "searchable_docs", "refund", {
    fields: ["title", "body"],
    scope,
});

const semantic = await vectorSearchRecords(ctx, "searchable_docs", {
    scope,
    vector,
    vectorField: "embedding",
});

const hybrid = await hybridSearchRecords(ctx, "searchable_docs", "refund", {
    fields: ["title", "body"],
    scope,
    vector,
    vectorField: "embedding",
});

SEARCH_EMBEDDING_PROFILE is the default app-owned search vector contract. The column width comes from createEmbeddingColumn(...), so if you change the profile model or dimensions, update the schema and rebuild stored vectors for that table.

Next Steps