RAG & Memory
Document ingestion, semantic search, and agent memory systems
This guide covers the RAG (Retrieval-Augmented Generation) and Memory components in detail.
[!IMPORTANT] Vector size is part of your storage contract. pgvector columns are fixed-width. A column declared as
vector(1536)accepts only 1536-length embeddings, so the embedding model must produce the same dimension as the schema. This is why the generated starters now exposeRAG_EMBEDDING_PROFILE,MEMORY_EMBEDDING_PROFILE, andSEARCH_EMBEDDING_PROFILE: each profile keeps the model id and vector size aligned across schema, runtime validation, and operational guidance. If you change a profile, rebuild stored vectors for that table.
RAG — Document Search Pipeline
RAG enables your AI to answer questions based on your own documents.
Setup
gencow add RAGThis creates:
gencow/rag.ts—createRag({ embeddingProfile, answerModel }), theragsingleton, legacy local RAG helpers, and canonical grounded-answer facadesgencow/schema-rag.ts— localrag_documentstable for the legacy helpers
Important: Import
schema-rag.tsin your schema to create the tables:// gencow/schema.ts export * from "./schema-rag";
Grounded answers use a different corpus path.
rag.askGrounded(),rag.compareCorpus(), andrag.extractTopics()read canonical Phase 2rag_*tables populated throughdocuments.ingest.*; documents inserted byrag.ingest()are only available torag.retrieve()/rag.ask().
Two RAG Surfaces
| Provider API | Facade API | |
|---|---|---|
| Entry point | createRag({ embeddingProfile, answerModel }) |
import { rag } from "./rag" |
| Best for | Explicit model injection, tests, AI SDK-native composition | Existing generated starter code and the default singleton |
| Local helpers | ingest(), retrieve(), ask() |
ingest(), retrieve(), ask(), askGrounded(), compareCorpus(), extractTopics() |
Provider API
Use the factory when you want explicit model selection:
import { createGencowAI } from "./ai";
import { createRag, RAG_EMBEDDING_PROFILE } from "./rag";
const gencow = createGencowAI();
const customRag = createRag({
embeddingProfile: RAG_EMBEDDING_PROFILE,
answerModel: gencow.languageModel("llm/economy"),
});
await customRag.ingest(ctx, "manual.md", documentText);
const hits = await customRag.retrieve(ctx, "refund policy?");
// → [{ chunk, source, similarity, metadata }]
const answer = await customRag.ask(ctx, "refund policy?");The generated schema-rag.ts starter derives rag_documents.embedding from
RAG_EMBEDDING_PROFILE.dimensions. That profile is the contract for both schema
and runtime validation. If you change embeddingProfile.model or
embeddingProfile.dimensions, update the schema and rebuild existing vectors
before mixing old and new rows.
Facade API
Use the generated singleton when you want the default compatibility surface:
import { rag } from "./rag";
await rag.ingest(ctx, "manual.md", documentText);
const hits = await rag.retrieve(ctx, "refund policy?");
const answer = await rag.ask(ctx, "refund policy?");
const grounded = await rag.askGrounded(ctx, "refund policy?", {
corpus: "default",
visibility: "shared",
});rag.askGrounded() does not read rows inserted by rag.ingest(). Use the
canonical ingest pipeline below when you need grounded citations.
Canonical RAG Foundation (Recommended)
For production RAG and grounded answers, use the built-in canonical pipeline:
storage file -> documents.ingest.start -> rag_* tables -> ctx.search() / ctx.grounding.answer()The canonical tables are rag_corpora, rag_sources, rag_sections, rag_chunks,
rag_ingest_jobs, and rag_operation_metrics. This path is tenant-scoped and is
the only path used by grounded citations.
Canonical ingest creates chunk embeddings through the platform-managed AI
transport in deployed apps. The runtime no longer uses ctx.ai for document
ingest. If no embedding target is configured in local development, chunks are
still stored for keyword search and embedding remains null; malformed
upstream responses fail with a safe diagnostic rather than exposing transport
details.
Conversion output, sections, chunks, and embedding vectors are persisted in canonical storage/tables. Workflow checkpoints contain only bounded summaries such as byte counts, chunk counts, target, and checksum; they do not duplicate file buffers, document text, chunk arrays, or vectors. Custom RAG workflows should follow the same persist-and-reference pattern.
Start Ingest
Enable the generated Cloud RAG client surface in gencow.config:
export default {
cloudFeatures: {
rag: true,
},
};import { useMutation } from "@gencow/react";
import { api } from "./gencow/api";
// inside a React component
const { mutate: startIngest } = useMutation(api.cloud.ragIngest.start);
async function handleIngest(storageId: string) {
await startIngest({
storageId,
corpus: "manuals",
visibility: "shared",
sourceKey: "refund-policy.pdf",
mode: "auto",
provider: "auto",
});
}Search Canonical Chunks
import { v } from "@gencow/core";
import { procedure } from "./runtime";
export const searchManuals = procedure.query
.name("manuals.search")
.input(v.object({ question: v.string() }))
.handler(async ({ context: ctx, input }) => {
ctx.auth.requireAuth();
return ctx.search("rag_chunks", input.question, {
fields: ["chunk_text", "lexical_text"],
scope: { corpus: "manuals", visibility: "shared" },
limit: 10,
});
});Rerank Retrieved Candidates
After tenant, owner, visibility, and read-grant filtering, rerank the retrieved candidates with Azure Cohere Fast. The gateway charges by the provider's actual search units and does not silently substitute a chat model.
import { ai } from "./ai";
async function rankCandidates(query: string, candidates: Array<{ id: string; text: string }>) {
const ranked = await ai.rerank({
query,
documents: candidates.map((candidate) => ({
id: candidate.id,
text: candidate.text,
})),
topK: 5,
});
return ranked.results;
}The hosted route accepts up to 100 candidate documents per call. See AI Engine reranking for provider choice and service-credit billing for pricing.
Grounded Answer
import { v } from "@gencow/core";
import { procedure } from "./runtime";
export const askManuals = procedure.query
.name("manuals.ask")
.input(v.object({ question: v.string() }))
.handler(async ({ context: ctx, input }) => {
ctx.auth.requireAuth();
if (!ctx.grounding) throw new Error("Grounding runtime is not available");
return ctx.grounding.answer({
question: input.question,
scope: { corpus: "manuals", visibility: "shared" },
mode: "qa",
budget: {
maxVerifyLoops: 2,
maxResearchQueriesPerLoop: 3,
maxCitationsPerClaim: 3,
},
});
},
});Operations Surface
import { useMutation, useQuery } from "@gencow/react";
import { api } from "./gencow/api";
const summary = useQuery(api.cloud.ragOps.summary, { corpus: "manuals" });
const metrics = useQuery(api.cloud.ragOps.metrics, { corpus: "manuals", limit: 20 });
const { mutate: evaluate } = useMutation(api.cloud.ragOps.evaluate);
const reindexPlan = useQuery(api.cloud.ragOps.reindexPlan, {
corpus: "manuals",
visibility: "shared",
mode: "corpus-policy-changed",
reason: "policy refresh",
});Use these operations to inspect source/chunk counts, ingest status, grounded-answer metrics, evaluation fixtures, and reindex candidates.
Facade API — Parsers + Reranker Example
Production file ingest should use documents.ingest.* and workflow document
conversion rather than calling legacy parsers directly. See
Document Conversion for PDF, HWPX, DOCX, and
XLSX routing.
import { v } from "@gencow/core";
import { procedure } from "./runtime";
import { parsers } from "./parsers";
import { rag } from "./rag";
import { reranker } from "./reranker";
export const ingestPdf = procedure.mutation
.name("docs.ingestPdf")
.handler(async ({ context: ctx, input }) => {
ctx.auth.requireAuth();
const file = input?.["file"] as File;
if (!file || typeof file === "string") throw new Error("No file");
const text = await parsers.pdf(Buffer.from(await file.arrayBuffer()));
await rag.ingest(ctx, file.name, text);
});
export const smartSearch = procedure.query
.name("docs.smartSearch")
.input(v.object({ question: v.string() }))
.handler(async ({ context: ctx, input }) => {
ctx.auth.requireAuth();
return reranker.searchAndRerank(ctx, rag, input.question);
});
export const askGroundedDocs = procedure.query
.name("docs.askGrounded")
.input(v.object({ question: v.string() }))
.handler(async ({ context: ctx, input }) => {
ctx.auth.requireAuth();
return await rag.askGrounded(ctx, input.question, {
corpus: "default",
visibility: "shared",
});
});Memory — Revisioned Memory Toolkit
gencow add Memory adds owner-scoped, revisioned durable memory. It is not a
chat-history store and never assembles a system prompt: the tenant app keeps
conversation messages and chooses where the returned untrusted context blocks go.
gencow add MemoryIt creates the canonical memory_* tables plus shared Search indexes. Export
schema-memory.ts from the app schema, then run gencow db:generate and
gencow db:push before writing memories.
import { createGencowAI } from "./ai";
import { createMemory, MEMORY_EMBEDDING_PROFILE } from "./memory";
const gencow = createGencowAI();
const characterMemory = createMemory({
extractionModel: gencow.languageModel("llm/economy"),
embeddingModel: gencow.embeddingModel("text-embedding-3-small"),
embeddingProfile: MEMORY_EMBEDDING_PROFILE,
});
const scope = { character: characterId, conversation: conversationId, branch: branchId };
await characterMemory.remember(ctx, {
namespace: "conversation-memory",
scope,
input: [{ role: "user", content: message }],
mode: "extract",
source: { type: "message", id: messageId, revision: "1" },
idempotencyKey: `message:${messageId}:1`,
policy: { version: "character-facts-v1" },
});
const recalled = await characterMemory.search(ctx, {
namespace: "conversation-memory", scope, query: message, tokenBudget: 800,
});
const context = await characterMemory.selectContext(ctx, {
namespace: "conversation-memory", scope, policyVersion: "character-context-v1",
tokenBudget: 800, sections: [{ name: "memory", priority: 60, items: recalled.hits }],
});remember() supports raw writes and model-based extraction; every mutation has
an idempotency receipt. update() uses an expected revision, forget()
scrubs historical content, and reconcileSource() handles edited/deleted source
messages. materialize() creates source-linked generic summaries.
Search uses shared semantic + PostgreSQL FTS + pg_trgm retrieval inside the
authenticated owner/namespace/scope. Embedding outages may return
lexical_degraded; a missing Search capability or index fails closed. Use
getOperation() and retryOperation() to converge a degraded projection
without creating a second canonical memory record.
Character-chat integration pattern
Memory is a generic toolkit, so a character-chat app maps its own identifiers to a flat scope. Store the user turn before recall, retrieve only in that scope, then put the selected result into an explicitly untrusted prompt block. The app still owns messages, character policy, safety instructions, and the model call.
const scope = {
character: characterId,
conversation: conversationId,
branch: branchId,
};
await characterMemory.remember(ctx, {
namespace: "conversation-memory",
scope,
input: [{ role: "user", content: userMessage }],
mode: "extract",
source: { type: "message", id: userMessageId, revision: "1" },
idempotencyKey: `message:${userMessageId}:1`,
policy: { version: "character-facts-v1" },
});
const recalled = await characterMemory.search(ctx, {
namespace: "conversation-memory",
scope,
query: userMessage,
tokenBudget: 800,
});
const selection = await characterMemory.selectContext(ctx, {
namespace: "conversation-memory",
scope,
policyVersion: "character-context-v1",
tokenBudget: 800,
sections: [{ name: "memory", priority: 60, items: recalled.hits }],
});
const untrustedMemoryBlock = selection.blocks
.map((item) => `- ${item.content}`)
.join("\n");Do not merge untrustedMemoryBlock into the system policy or treat it as an
instruction. When a user edits or deletes a message, call reconcileSource();
when an answer is regenerated, call resolveContext() with the prior receipt
so the app can reproduce the same bounded selection.
Upgrading the prototype
The v1 installer deliberately does not overwrite an existing
gencow/memory.ts, does not rename or drop agent_memories, and does not copy
old chat-session data automatically. Add the Toolkit to a new module path (or
move the old files aside after a code review), export schema-memory.ts, and
generate a normal forward Drizzle migration. Keep the old table read-only until
the app has independently verified its migration policy; any backfill must call
remember() with a stable legacy source id and idempotency key. Do not write
direct SQL that bypasses owner/scope validation, and do not delete the legacy
table in the same release as the first v1 deployment.
Search Starter — App-Owned Retrieval Tables
The Search starter is the companion surface for app-owned keyword/vector/hybrid
retrieval tables outside canonical rag_* ingest.
gencow add SearchThis creates:
gencow/search.ts— thin wrappers forctx.search(),ctx.vectorSearch(), andctx.hybridSearch()gencow/schema-search.ts—searchScopeColumns,SEARCH_EMBEDDING_PROFILE, andcreateEmbeddingColumn()
import { pgTable, text } from "drizzle-orm/pg-core";
import { createEmbeddingColumn, SEARCH_EMBEDDING_PROFILE, searchScopeColumns } from "./schema-search";
import { privateScope, searchRecords, vectorSearchRecords, hybridSearchRecords } from "./search";
export const searchableDocs = pgTable("searchable_docs", {
title: text("title").notNull(),
body: text("body").notNull(),
...searchScopeColumns,
embedding: createEmbeddingColumn(SEARCH_EMBEDDING_PROFILE),
});
const scope = privateScope(ctx, "docs");
const keyword = await searchRecords(ctx, "searchable_docs", "refund", {
fields: ["title", "body"],
scope,
});
const semantic = await vectorSearchRecords(ctx, "searchable_docs", {
scope,
vector,
vectorField: "embedding",
});
const hybrid = await hybridSearchRecords(ctx, "searchable_docs", "refund", {
fields: ["title", "body"],
scope,
vector,
vectorField: "embedding",
});SEARCH_EMBEDDING_PROFILE is the default app-owned search vector contract. The
column width comes from createEmbeddingColumn(...), so if you change the
profile model or dimensions, update the schema and rebuild stored vectors for
that table.
Next Steps
- CLI Reference — All CLI commands
- React Hooks — Frontend API reference
- Core API — Backend API reference