Components

Add AI capabilities with gencow add — AI, speech, RAG, Memory, Tools, Guardrails, and more

The gencow add command installs pre-built AI components into your project. Each component adds files to gencow/ and installs required dependencies.

Usage

# Add a single component
bunx gencow@latest add AI

# Add multiple at once
bunx gencow@latest add AI RAG Reranker

# Dependencies auto-resolve (RAG requires AI → AI is added automatically)
bunx gencow@latest add RAG   # → also installs AI

Available Components

# Component Description Requires
1 AI Vercel AI SDK wrapper (OpenAI chat, embeddings, structured output, images, and speech; Azure Cohere reranking; feature-flagged cloud text streaming) —
2 Agent Durable workflow agent starter with injectable AI SDK model AI
3 Tools AI Tool Calling with ctx integration AI
4 RAG Document ingestion + semantic search with injectable embedding/answer models AI
5 Memory Agent memory (episodic/semantic/procedural) with injectable extraction/embedding models AI
6 Reranker Dedicated Azure Cohere Fast reranking for retrieved candidates AI
7 Guardrails Input/output safety filters (PII, topic blocking) with injectable model AI
8 Prompts Reusable prompt templates —
9 Parsers PDF/HTML/CSV file parsing —
10 Analytics LLM call tracking AI (coming soon)

Component Details

AI — Core Engine

bunx gencow@latest add AI

Creates gencow/ai.ts (and related generated files). Import from gencow/ai.ts — see AI Engine for model catalog, billing, structured output, and the full runtime contract.

gencow/ai.ts exposes two APIs. Prefer the provider API for new code; the facade API is a narrower compatibility layer for simple handlers and other Gencow components.

Provider API (recommended):

import { embed, generateText } from "ai";
import { createGencowAI } from "./ai";

const gencow = createGencowAI();

await generateText({
    model: gencow.languageModel("llm/economy"),
    messages: [{ role: "user", content: "Hello" }],
});

await embed({
    model: gencow.embeddingModel("text-embedding-3-small"),
    value: "Hello, world!",
});

pgvector schema match required: if you later store these embeddings in pgvector, the embedding model dimension must match the DB column size. The generated RAG and Memory starter schemas use vector(1536).

Facade API:

import { ai } from "./ai";

ai.chat({ system, messages, model })    // Non-streaming response
ai.stream({ system, messages, model })  // Text streaming when local direct mode or cloud streaming is enabled
ai.embed(text)                          // Generate embeddings
ai.image.generate({ prompt, model })    // Generate images with GPT Image
ai.vision.extractText({ image, mediaType }) // Extract text from image input
ai.rerank({ query, documents, topK })        // Azure Cohere Fast reranking in cloud
ai.speech.transcribe({ storageId, model })  // Start async STT from a private storage object
ai.speech.waitForTranscript(jobId)          // Poll until the STT job reaches a terminal state
ai.speech.getTranscriptResult(jobId)        // Read the private canonical transcript
ai.speech.synthesize({ input, model: "gpt-4o-mini-tts", voice: "alloy" }) // Private TTS audio

Installed deps: ai, @ai-sdk/openai Env required: OPENAI_API_KEY (local only — auto in cloud)

Speech helpers are proxy-backed even in development because they need Gencow storage lookup, media preprocessing, async job state, and service-credit billing. For speech-to-text, upload audio or video to private app storage first and pass only the resulting storageId; do not send raw multipart audio to the AI proxy. OpenAI speech-to-text jobs continue independently when the app sleeps. azure-speech-fast-transcription is the supported Azure STT exception for timestamp-capable transcripts. Other Azure speech routes remain unavailable. OpenAI gpt-4o-mini-tts supports direct synthesis and async jobs with request-level usage billing. See Azure Speech for the Fast Transcription exception and AI Engine for billing.

Speech availability requires a supported route and billable usage meter, not only a catalog entry. Token-billed OpenAI speech settles from request usage; Azure Fast Transcription uses its audio-duration meter. Preserve the original job ID and poll its status after an interrupted request instead of creating another job. See AI Engine billing.


Generated AI Factories

Generated AI-dependent components now follow the same dual-surface pattern as gencow/ai.ts:

  • Provider API: call createX(...) and inject models from createGencowAI().
  • Facade API: import the generated singleton (rag, memory, guardrails, reranker, runAgent) for compatibility.

Use the factory when a handler needs explicit model selection, tests need mock models, or you are composing with AI SDK helpers directly. Existing singleton imports remain stable for generated starter code.


Agent — Workflow Agent

gencow add Agent

Creates gencow/agent.ts.

Provider API (recommended):

import { createGencowAI } from "./ai";
import { createWorkflowAgent } from "./agent";

const gencow = createGencowAI();

export const customRunAgent = createWorkflowAgent({
    model: gencow.languageModel("llm/economy"),
});

Use the factory when you want explicit model selection or test doubles.

Facade API:

import { runAgent } from "./agent";

// Backward-compatible singleton workflow.
// Existing generated apps can keep exporting or starting runAgent directly.
export { runAgent };

Use useWorkflow(api.workflows.get, run.id) and workflows.signal to observe or resume durable agent runs.


Tools — AI Function Calling

gencow add Tools

Creates gencow/tools.ts with defineTools() — a thin wrapper around AI SDK's tool() that injects ctx into handlers.

Provider API (use AI SDK tool() directly):

import { generateText, tool } from "ai";
import { z } from "zod";
import { createGencowAI } from "./ai";

const gencow = createGencowAI();

const getWeather = tool({
    description: "Get weather for a city",
    parameters: z.object({ city: z.string() }),
    execute: async ({ city }) => `Weather in ${city}: 22°C, sunny`,
});

const result = await generateText({
    model: gencow.languageModel("llm/economy"),
    messages: [{ role: "user", content: "What's the weather in Seoul?" }],
    tools: { getWeather },
});

Facade API (with defineTools + ai.chat()):

import { z } from "zod";
import { ai } from "./ai";
import { defineTools } from "./tools";

const tools = defineTools(ctx, {
    getWeather: {
        description: "Get weather for a city",
        parameters: z.object({ city: z.string() }),
        handler: async (ctx, { city }) => {
            return `Weather in ${city}: 22°C, sunny`;
        },
    },
});

const result = await ai.chat({
    messages: [{ role: "user", content: "What's the weather in Seoul?" }],
    tools,
});
console.log(result.text);

gencow add RAG

Creates gencow/rag.ts + gencow/schema-rag.ts.

Provider API (recommended):

import { createGencowAI } from "./ai";
import { createRag, RAG_EMBEDDING_PROFILE } from "./rag";

const gencow = createGencowAI();
const customRag = createRag({
    embeddingProfile: RAG_EMBEDDING_PROFILE,
    answerModel: gencow.languageModel("llm/economy"),
});

await customRag.ingest(ctx, "manual.md", "Document text content...");
const hits = await customRag.retrieve(ctx, "What is Gencow?");
// → [{ chunk, source, similarity, metadata }]

Use the factory when you want explicit embedding-profile and answer-model control.

The generated RAG starter now treats the embedding profile as the storage contract. RAG_EMBEDDING_PROFILE carries both the model id and dimensions for rag_documents.embedding, and the generated schema derives the vector width from that same profile. If you change embeddingProfile.model or embeddingProfile.dimensions, update the schema and rebuild stored vectors for that table.

Facade API:

import { rag } from "./rag";

await rag.ingest(ctx, "manual.md", "Document text content...");
const answer = await rag.ask(ctx, "What is Gencow?");
const grounded = await rag.askGrounded(ctx, "What is Gencow?", {
    corpus: "default",
    visibility: "shared",
});

Important: Import schema-rag.ts in your main schema.ts to create the required DB tables.

rag.ingest() writes to the local rag_documents table. Grounded helpers such as rag.askGrounded() and reranker.answerGrounded() read canonical Phase 2 rag_* corpora populated through documents.ingest.*.

[!IMPORTANT] Embedding model and vector size must stay aligned with the pgvector column. The default starter profile uses text-embedding-3-small with 1536 dimensions. If you change the model, update the profile dimensions, schema, and stored vectors to match. See AI Engine for the general rule and RAG & Memory for starter-specific guidance.


Memory — Memory Toolkit

gencow add Memory

Creates gencow/memory.ts, gencow/schema-memory.ts, and the supporting scope, policy, error, and type modules. It is an owner-scoped, revisioned Memory Toolkit; it does not create a chat session store or assemble a system prompt.

import { createGencowAI } from "./ai";
import { createMemory, MEMORY_EMBEDDING_PROFILE } from "./memory";

const gencow = createGencowAI();
const customMemory = createMemory({
    extractionModel: gencow.languageModel("llm/economy"),
    embeddingProfile: MEMORY_EMBEDDING_PROFILE,
});

await customMemory.remember(ctx, {
    namespace: "conversation-memory",
    scope: { character: characterId, conversation: conversationId, branch: branchId },
    input: [{ role: "user", content: query }],
    mode: "extract",
    source: { type: "message", id: messageId, revision: "1" },
    idempotencyKey: `memory:${messageId}:1`,
    policy: { version: "character-facts-v1" },
});

const recalled = await customMemory.search(ctx, {
    namespace: "conversation-memory",
    scope: { character: characterId, conversation: conversationId, branch: branchId },
    query,
});

const context = await customMemory.selectContext(ctx, {
    namespace: "conversation-memory",
    scope: { character: characterId, conversation: conversationId, branch: branchId },
    policyVersion: "character-context-v1",
    tokenBudget: 1200,
    sections: [{ name: "relevant", priority: 60, items: recalled.hits }],
});

Use update() for a user correction, forget() for privacy deletion, and reconcileSource() when the app edits or deletes a source message. The Toolkit returns role-free data blocks from selectContext(); the app must place them in an untrusted context section after its own system and safety policy.

MEMORY_EMBEDDING_PROFILE applies to memory_search_documents.embedding. Changing its model or dimensions requires schema alignment and a reindex.

For character chat, map character, conversation, and branch to scope; store each message with a stable source id and idempotency key; then call search() and selectContext() before generating the reply. Memory returns untrusted data only—the app retains ownership of chat history, system policy, and prompt construction. See RAG & Memory for the full edit/delete and answer-regeneration pattern.

The default search path is semantic + PostgreSQL FTS + pg_trgm. An embedding outage can use lexical retrieval inside the same owner and scope; missing Search capability or indexes fail closed instead of scanning another user's memory.

Lifecycle API:

API Purpose
remember() Raw or extracted additive memory with an idempotency receipt
get() / list() / history() Canonical owner-scoped reads and revision audit
search() Scope-isolated hybrid recall with lexical degradation
materialize() Source-bound generic summary memory
selectContext() / resolveContext() Deterministic token budget and replay receipt
getOperation() / retryOperation() Same-identity operation observation and projection recovery

Search — App-Owned Retrieval Tables

gencow add Search

Creates gencow/search.ts + gencow/schema-search.ts.

import { pgTable, text } from "drizzle-orm/pg-core";
import { createEmbeddingColumn, SEARCH_EMBEDDING_PROFILE, searchScopeColumns } from "./schema-search";
import { privateScope, searchRecords, vectorSearchRecords, hybridSearchRecords } from "./search";

export const searchableDocs = pgTable("searchable_docs", {
    title: text("title").notNull(),
    body: text("body").notNull(),
    ...searchScopeColumns,
    embedding: createEmbeddingColumn(SEARCH_EMBEDDING_PROFILE),
});

const scope = privateScope(ctx, "docs");
const keyword = await searchRecords(ctx, "searchable_docs", "refund", {
    fields: ["title", "body"],
    scope,
});

const semantic = await vectorSearchRecords(ctx, "searchable_docs", {
    scope,
    vector,
    vectorField: "embedding",
});

const hybrid = await hybridSearchRecords(ctx, "searchable_docs", "refund", {
    fields: ["title", "body"],
    scope,
    vector,
    vectorField: "embedding",
});

SEARCH_EMBEDDING_PROFILE is the default app-owned search vector contract. createEmbeddingColumn(...) derives the column width from that profile, so if you change the model or dimensions, update the schema and rebuild stored vectors for that table.


Reranker — Result Quality

gencow add Reranker

Creates gencow/reranker.ts. In cloud, its dedicated rerank path uses Azure Cohere Cohere-rerank-v4.0-fast; the gateway does not silently substitute a chat model. Retrieve and authorize candidates before reranking them. The generated starter also supports an explicitly supplied fallback model for application-owned compatibility behavior. See AI Engine reranking for an ai.rerank() example and search-unit credit pricing.


Guardrails — Safety Filters

gencow add Guardrails

Creates gencow/guardrails.ts.

Provider API (recommended):

import { createGencowAI } from "./ai";
import { createGuardrails } from "./guardrails";

const gencow = createGencowAI();
const customGuardrails = createGuardrails({
    model: gencow.languageModel("llm/economy"),
});

const safe = await customGuardrails.validateInput(userText, {
    maskPII: true,
    blockTopics: ["politics"],
});

const checked = await customGuardrails.validateOutput(aiResponse, {
    maxLength: 2000,
    bannedPatterns: [/api[_-]?key/i],
});

Facade API:

import { guardrails } from "./guardrails";

const safe = await guardrails.validateInput(userText, {
    maskPII: true,
    blockTopics: ["politics"],
});

const wrapped = await guardrails.wrap(
    async (sanitized) => myFunction(sanitized),
    userText,
    { maskPII: true },
    { maxLength: 2000 },
);

Prompts — Reusable Templates

gencow add Prompts

Creates gencow/prompts.ts. Pre-built prompt templates:

import { ragQAPrompt, summarizePrompt, classifyPrompt } from "./prompts";

// RAG Q&A
const prompt = ragQAPrompt({
    question: "What is Gencow?",
    context: "Retrieved context here...",
});

// Summarization
const prompt = summarizePrompt({ text: "Long text to summarize..." });

// Classification
const prompt = classifyPrompt({
    text: "I love this product!",
    categories: ["positive", "negative", "neutral"],
});

Parsers — File Parsing

gencow add Parsers

Creates gencow/parsers.ts. Parse various file formats:

// PDF
const text = await parsers.pdf(buffer);

// HTML
const text = await parsers.html(htmlString);

// CSV
const rows = await parsers.csv(csvString);

// Auto-detect by filename
const text = await parsers.auto("document.pdf", buffer);

// RAG integration
await rag.ingest(ctx, "manual.pdf", await parsers.pdf(buffer));

Installed deps: pdf-parse

Dependency Resolution

Components automatically install their dependencies:

gencow add RAG
# → Also installs AI (because RAG requires AI)
gencow add RAG Reranker Memory
# → Installs AI first, then RAG, Reranker, Memory

After Installation

# README is auto-updated with new component docs
gencow dev

Each component's usage is added to the auto-generated gencow/README.md.

Next Steps