Components
Add AI capabilities with gencow add — AI, speech, RAG, Memory, Tools, Guardrails, and more
The gencow add command installs pre-built AI components into your project. Each component adds files to gencow/ and installs required dependencies.
Usage
# Add a single component
bunx gencow@latest add AI
# Add multiple at once
bunx gencow@latest add AI RAG Reranker
# Dependencies auto-resolve (RAG requires AI → AI is added automatically)
bunx gencow@latest add RAG # → also installs AIAvailable Components
| # | Component | Description | Requires |
|---|---|---|---|
| 1 | AI | Vercel AI SDK wrapper (OpenAI chat, embeddings, structured output, images, and speech; Azure Cohere reranking; feature-flagged cloud text streaming) | — |
| 2 | Agent | Durable workflow agent starter with injectable AI SDK model | AI |
| 3 | Tools | AI Tool Calling with ctx integration |
AI |
| 4 | RAG | Document ingestion + semantic search with injectable embedding/answer models | AI |
| 5 | Memory | Agent memory (episodic/semantic/procedural) with injectable extraction/embedding models | AI |
| 6 | Reranker | Dedicated Azure Cohere Fast reranking for retrieved candidates | AI |
| 7 | Guardrails | Input/output safety filters (PII, topic blocking) with injectable model | AI |
| 8 | Prompts | Reusable prompt templates | — |
| 9 | Parsers | PDF/HTML/CSV file parsing | — |
| 10 | Analytics | LLM call tracking | AI (coming soon) |
Component Details
AI — Core Engine
bunx gencow@latest add AICreates gencow/ai.ts (and related generated files). Import from gencow/ai.ts —
see AI Engine for model catalog, billing, structured output,
and the full runtime contract.
gencow/ai.ts exposes two APIs. Prefer the provider API for new code; the
facade API is a narrower compatibility layer for simple handlers and other
Gencow components.
Provider API (recommended):
import { embed, generateText } from "ai";
import { createGencowAI } from "./ai";
const gencow = createGencowAI();
await generateText({
model: gencow.languageModel("llm/economy"),
messages: [{ role: "user", content: "Hello" }],
});
await embed({
model: gencow.embeddingModel("text-embedding-3-small"),
value: "Hello, world!",
});pgvector schema match required: if you later store these embeddings in pgvector, the embedding model dimension must match the DB column size. The generated RAG and Memory starter schemas use
vector(1536).
Facade API:
import { ai } from "./ai";
ai.chat({ system, messages, model }) // Non-streaming response
ai.stream({ system, messages, model }) // Text streaming when local direct mode or cloud streaming is enabled
ai.embed(text) // Generate embeddings
ai.image.generate({ prompt, model }) // Generate images with GPT Image
ai.vision.extractText({ image, mediaType }) // Extract text from image input
ai.rerank({ query, documents, topK }) // Azure Cohere Fast reranking in cloud
ai.speech.transcribe({ storageId, model }) // Start async STT from a private storage object
ai.speech.waitForTranscript(jobId) // Poll until the STT job reaches a terminal state
ai.speech.getTranscriptResult(jobId) // Read the private canonical transcript
ai.speech.synthesize({ input, model: "gpt-4o-mini-tts", voice: "alloy" }) // Private TTS audioInstalled deps: ai, @ai-sdk/openai
Env required: OPENAI_API_KEY (local only — auto in cloud)
Speech helpers are proxy-backed even in development because they need Gencow
storage lookup, media preprocessing, async job state, and service-credit billing.
For speech-to-text, upload audio or video to private app storage first and pass
only the resulting storageId; do not send raw multipart audio to the AI proxy.
OpenAI speech-to-text jobs continue independently when the app sleeps.
azure-speech-fast-transcription is the supported Azure STT exception for
timestamp-capable transcripts. Other Azure speech routes remain unavailable.
OpenAI gpt-4o-mini-tts supports direct synthesis and async jobs with
request-level usage billing. See Azure Speech for the
Fast Transcription exception and AI Engine for billing.
Speech availability requires a supported route and billable usage meter, not only a catalog entry. Token-billed OpenAI speech settles from request usage; Azure Fast Transcription uses its audio-duration meter. Preserve the original job ID and poll its status after an interrupted request instead of creating another job. See AI Engine billing.
Generated AI Factories
Generated AI-dependent components now follow the same dual-surface pattern as
gencow/ai.ts:
- Provider API: call
createX(...)and inject models fromcreateGencowAI(). - Facade API: import the generated singleton (
rag,memory,guardrails,reranker,runAgent) for compatibility.
Use the factory when a handler needs explicit model selection, tests need mock models, or you are composing with AI SDK helpers directly. Existing singleton imports remain stable for generated starter code.
Agent — Workflow Agent
gencow add AgentCreates gencow/agent.ts.
Provider API (recommended):
import { createGencowAI } from "./ai";
import { createWorkflowAgent } from "./agent";
const gencow = createGencowAI();
export const customRunAgent = createWorkflowAgent({
model: gencow.languageModel("llm/economy"),
});Use the factory when you want explicit model selection or test doubles.
Facade API:
import { runAgent } from "./agent";
// Backward-compatible singleton workflow.
// Existing generated apps can keep exporting or starting runAgent directly.
export { runAgent };Use useWorkflow(api.workflows.get, run.id) and workflows.signal to observe
or resume durable agent runs.
Tools — AI Function Calling
gencow add ToolsCreates gencow/tools.ts with defineTools() — a thin wrapper around AI SDK's
tool() that injects ctx into handlers.
Provider API (use AI SDK tool() directly):
import { generateText, tool } from "ai";
import { z } from "zod";
import { createGencowAI } from "./ai";
const gencow = createGencowAI();
const getWeather = tool({
description: "Get weather for a city",
parameters: z.object({ city: z.string() }),
execute: async ({ city }) => `Weather in ${city}: 22°C, sunny`,
});
const result = await generateText({
model: gencow.languageModel("llm/economy"),
messages: [{ role: "user", content: "What's the weather in Seoul?" }],
tools: { getWeather },
});Facade API (with defineTools + ai.chat()):
import { z } from "zod";
import { ai } from "./ai";
import { defineTools } from "./tools";
const tools = defineTools(ctx, {
getWeather: {
description: "Get weather for a city",
parameters: z.object({ city: z.string() }),
handler: async (ctx, { city }) => {
return `Weather in ${city}: 22°C, sunny`;
},
},
});
const result = await ai.chat({
messages: [{ role: "user", content: "What's the weather in Seoul?" }],
tools,
});
console.log(result.text);RAG — Document Search
gencow add RAGCreates gencow/rag.ts + gencow/schema-rag.ts.
Provider API (recommended):
import { createGencowAI } from "./ai";
import { createRag, RAG_EMBEDDING_PROFILE } from "./rag";
const gencow = createGencowAI();
const customRag = createRag({
embeddingProfile: RAG_EMBEDDING_PROFILE,
answerModel: gencow.languageModel("llm/economy"),
});
await customRag.ingest(ctx, "manual.md", "Document text content...");
const hits = await customRag.retrieve(ctx, "What is Gencow?");
// → [{ chunk, source, similarity, metadata }]Use the factory when you want explicit embedding-profile and answer-model control.
The generated RAG starter now treats the embedding profile as the storage contract.
RAG_EMBEDDING_PROFILE carries both the model id and dimensions for
rag_documents.embedding, and the generated schema derives the vector width from
that same profile. If you change embeddingProfile.model or
embeddingProfile.dimensions, update the schema and rebuild stored vectors for
that table.
Facade API:
import { rag } from "./rag";
await rag.ingest(ctx, "manual.md", "Document text content...");
const answer = await rag.ask(ctx, "What is Gencow?");
const grounded = await rag.askGrounded(ctx, "What is Gencow?", {
corpus: "default",
visibility: "shared",
});Important: Import
schema-rag.tsin your mainschema.tsto create the required DB tables.
rag.ingest()writes to the localrag_documentstable. Grounded helpers such asrag.askGrounded()andreranker.answerGrounded()read canonical Phase 2rag_*corpora populated throughdocuments.ingest.*.
[!IMPORTANT] Embedding model and vector size must stay aligned with the pgvector column. The default starter profile uses
text-embedding-3-smallwith1536dimensions. If you change the model, update the profile dimensions, schema, and stored vectors to match. See AI Engine for the general rule and RAG & Memory for starter-specific guidance.
Memory — Memory Toolkit
gencow add MemoryCreates gencow/memory.ts, gencow/schema-memory.ts, and the supporting
scope, policy, error, and type modules. It is an owner-scoped, revisioned
Memory Toolkit; it does not create a chat session store or assemble a system prompt.
import { createGencowAI } from "./ai";
import { createMemory, MEMORY_EMBEDDING_PROFILE } from "./memory";
const gencow = createGencowAI();
const customMemory = createMemory({
extractionModel: gencow.languageModel("llm/economy"),
embeddingProfile: MEMORY_EMBEDDING_PROFILE,
});
await customMemory.remember(ctx, {
namespace: "conversation-memory",
scope: { character: characterId, conversation: conversationId, branch: branchId },
input: [{ role: "user", content: query }],
mode: "extract",
source: { type: "message", id: messageId, revision: "1" },
idempotencyKey: `memory:${messageId}:1`,
policy: { version: "character-facts-v1" },
});
const recalled = await customMemory.search(ctx, {
namespace: "conversation-memory",
scope: { character: characterId, conversation: conversationId, branch: branchId },
query,
});
const context = await customMemory.selectContext(ctx, {
namespace: "conversation-memory",
scope: { character: characterId, conversation: conversationId, branch: branchId },
policyVersion: "character-context-v1",
tokenBudget: 1200,
sections: [{ name: "relevant", priority: 60, items: recalled.hits }],
});Use update() for a user correction, forget() for privacy deletion, and
reconcileSource() when the app edits or deletes a source message. The Toolkit
returns role-free data blocks from selectContext(); the app must place them in
an untrusted context section after its own system and safety policy.
MEMORY_EMBEDDING_PROFILE applies to memory_search_documents.embedding.
Changing its model or dimensions requires schema alignment and a reindex.
For character chat, map character, conversation, and branch to scope;
store each message with a stable source id and idempotency key; then call
search() and selectContext() before generating the reply. Memory returns
untrusted data only—the app retains ownership of chat history, system policy,
and prompt construction. See RAG & Memory for the full
edit/delete and answer-regeneration pattern.
The default search path is semantic + PostgreSQL FTS + pg_trgm. An embedding outage can use lexical retrieval inside the same owner and scope; missing Search capability or indexes fail closed instead of scanning another user's memory.
Lifecycle API:
| API | Purpose |
|---|---|
remember() |
Raw or extracted additive memory with an idempotency receipt |
get() / list() / history() |
Canonical owner-scoped reads and revision audit |
search() |
Scope-isolated hybrid recall with lexical degradation |
materialize() |
Source-bound generic summary memory |
selectContext() / resolveContext() |
Deterministic token budget and replay receipt |
getOperation() / retryOperation() |
Same-identity operation observation and projection recovery |
Search — App-Owned Retrieval Tables
gencow add SearchCreates gencow/search.ts + gencow/schema-search.ts.
import { pgTable, text } from "drizzle-orm/pg-core";
import { createEmbeddingColumn, SEARCH_EMBEDDING_PROFILE, searchScopeColumns } from "./schema-search";
import { privateScope, searchRecords, vectorSearchRecords, hybridSearchRecords } from "./search";
export const searchableDocs = pgTable("searchable_docs", {
title: text("title").notNull(),
body: text("body").notNull(),
...searchScopeColumns,
embedding: createEmbeddingColumn(SEARCH_EMBEDDING_PROFILE),
});
const scope = privateScope(ctx, "docs");
const keyword = await searchRecords(ctx, "searchable_docs", "refund", {
fields: ["title", "body"],
scope,
});
const semantic = await vectorSearchRecords(ctx, "searchable_docs", {
scope,
vector,
vectorField: "embedding",
});
const hybrid = await hybridSearchRecords(ctx, "searchable_docs", "refund", {
fields: ["title", "body"],
scope,
vector,
vectorField: "embedding",
});SEARCH_EMBEDDING_PROFILE is the default app-owned search vector contract.
createEmbeddingColumn(...) derives the column width from that profile, so if
you change the model or dimensions, update the schema and rebuild stored
vectors for that table.
Reranker — Result Quality
gencow add RerankerCreates gencow/reranker.ts. In cloud, its dedicated rerank path uses Azure
Cohere Cohere-rerank-v4.0-fast; the gateway does not silently substitute a
chat model. Retrieve and authorize candidates before reranking them. The
generated starter also supports an explicitly supplied fallback model for
application-owned compatibility behavior. See AI Engine reranking
for an ai.rerank() example and search-unit credit pricing.
Guardrails — Safety Filters
gencow add GuardrailsCreates gencow/guardrails.ts.
Provider API (recommended):
import { createGencowAI } from "./ai";
import { createGuardrails } from "./guardrails";
const gencow = createGencowAI();
const customGuardrails = createGuardrails({
model: gencow.languageModel("llm/economy"),
});
const safe = await customGuardrails.validateInput(userText, {
maskPII: true,
blockTopics: ["politics"],
});
const checked = await customGuardrails.validateOutput(aiResponse, {
maxLength: 2000,
bannedPatterns: [/api[_-]?key/i],
});Facade API:
import { guardrails } from "./guardrails";
const safe = await guardrails.validateInput(userText, {
maskPII: true,
blockTopics: ["politics"],
});
const wrapped = await guardrails.wrap(
async (sanitized) => myFunction(sanitized),
userText,
{ maskPII: true },
{ maxLength: 2000 },
);Prompts — Reusable Templates
gencow add PromptsCreates gencow/prompts.ts. Pre-built prompt templates:
import { ragQAPrompt, summarizePrompt, classifyPrompt } from "./prompts";
// RAG Q&A
const prompt = ragQAPrompt({
question: "What is Gencow?",
context: "Retrieved context here...",
});
// Summarization
const prompt = summarizePrompt({ text: "Long text to summarize..." });
// Classification
const prompt = classifyPrompt({
text: "I love this product!",
categories: ["positive", "negative", "neutral"],
});Parsers — File Parsing
gencow add ParsersCreates gencow/parsers.ts. Parse various file formats:
// PDF
const text = await parsers.pdf(buffer);
// HTML
const text = await parsers.html(htmlString);
// CSV
const rows = await parsers.csv(csvString);
// Auto-detect by filename
const text = await parsers.auto("document.pdf", buffer);
// RAG integration
await rag.ingest(ctx, "manual.pdf", await parsers.pdf(buffer));Installed deps: pdf-parse
Dependency Resolution
Components automatically install their dependencies:
gencow add RAG
# → Also installs AI (because RAG requires AI)gencow add RAG Reranker Memory
# → Installs AI first, then RAG, Reranker, MemoryAfter Installation
# README is auto-updated with new component docs
gencow devEach component's usage is added to the auto-generated gencow/README.md.
Next Steps
- AI Engine — Provider vs facade APIs, models, billing, structured output
- Azure Speech — Fast Transcription timestamp option
- RAG & Memory — Deep dive into document search and agent memory
- CLI Reference — All CLI commands