Documentation Index Fetch the complete documentation index at: https://mintlify.com/cloudflare/agents/llms.txt
Use this file to discover all available pages before exploring further.
Build AI-powered chat interfaces with AIChatAgent and useAgentChat. Messages are automatically persisted to SQLite, streams resume on disconnect, and tool calls work across server and client.
Overview
@cloudflare/ai-chat provides two main exports:
Export Import Purpose AIChatAgent@cloudflare/ai-chatServer-side agent class with message persistence and streaming useAgentChat@cloudflare/ai-chat/reactReact hook for building chat UIs
Built on the AI SDK and Cloudflare Durable Objects, you get:
Automatic message persistence — conversations stored in SQLite, survive restarts
Resumable streaming — disconnected clients resume mid-stream without data loss
Real-time sync — messages broadcast to all connected clients via WebSocket
Tool support — server-side, client-side, and human-in-the-loop tool patterns
Data parts — attach typed JSON (citations, progress, usage) to messages alongside text
Row size protection — automatic compaction when messages approach SQLite limits
Quick Start
Install dependencies
npm install @cloudflare/ai-chat agents ai workers-ai-provider
Create server agent
import { AIChatAgent } from "@cloudflare/ai-chat" ;
import { createWorkersAI } from "workers-ai-provider" ;
import { streamText , convertToModelMessages } from "ai" ;
export class ChatAgent extends AIChatAgent {
async onChatMessage () {
const workersai = createWorkersAI ({ binding: this . env . AI });
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages )
});
return result . toUIMessageStreamResponse ();
}
}
Create client UI
import { useAgent } from "agents/react" ;
import { useAgentChat } from "@cloudflare/ai-chat/react" ;
function Chat () {
const agent = useAgent ({ agent: "ChatAgent" });
const { messages , sendMessage , status } = useAgentChat ({ agent });
return (
< div >
{ messages . map (( msg ) => (
< div key = { msg . id } >
< strong > { msg . role } : </ strong >
{ msg . parts . map (( part , i ) =>
part . type === "text" ? < span key = { i } > { part . text } </ span > : null
) }
</ div >
)) }
< form
onSubmit = { ( e ) => {
e . preventDefault ();
const input = e . currentTarget . elements . namedItem (
"input"
) as HTMLInputElement ;
sendMessage ({ text: input . value });
input . value = "" ;
} }
>
< input name = "input" placeholder = "Type a message..." />
< button type = "submit" disabled = { status === "streaming" } >
Send
</ button >
</ form >
</ div >
);
}
Configure Wrangler
{
"ai" : { "binding" : "AI" },
"durable_objects" : {
"bindings" : [{ "name" : "ChatAgent" , "class_name" : "ChatAgent" }]
},
"migrations" : [{ "tag" : "v1" , "new_sqlite_classes" : [ "ChatAgent" ] }]
}
The new_sqlite_classes migration is required — AIChatAgent uses SQLite for message persistence and stream chunk buffering.
How It Works
┌──────────┐ WebSocket ┌──────────────┐
│ Client │ ◀──────────────────────────────────▶ │ AIChatAgent │
│ │ │ │
│ useAgent │ CF_AGENT_USE_CHAT_REQUEST ──────▶ │ onChatMessage│
│ Chat │ │ │
│ │ ◀────── CF_AGENT_USE_CHAT_RESPONSE │ streamText │
│ │ (UIMessageChunk stream) │ │
│ │ │ SQLite │
│ │ ◀────── CF_AGENT_CHAT_MESSAGES │ (messages, │
│ │ (broadcast to all clients) │ chunks) │
└──────────┘ └──────────────┘
Client sends message
The client sends a message via WebSocket
Agent persists and calls handler
AIChatAgent persists messages to SQLite and calls your onChatMessage method
Stream response
Your method returns a streaming Response (typically from streamText)
Real-time chunks
Chunks stream back over WebSocket in real-time
Broadcast final message
When the stream completes, the final message is persisted and broadcast to all connections
Server API
AIChatAgent
Extends Agent from the agents package. Manages conversation state, persistence, and streaming.
import { AIChatAgent } from "@cloudflare/ai-chat" ;
export class ChatAgent extends AIChatAgent {
// Access current messages
// this.messages: UIMessage[]
// Limit stored messages (optional)
maxPersistedMessages = 200 ;
async onChatMessage ( onFinish ? , options ? ) {
// onFinish: optional callback for streamText (cleanup is automatic)
// options.abortSignal: cancel signal
// options.body: custom data from client
// Return a Response (streaming or plain text)
}
}
onChatMessage
This is the main method you override. It receives the conversation context and should return a Response.
Streaming response
Plain text response
Custom body data
async onChatMessage () {
const workersai = createWorkersAI ({ binding: this . env . AI });
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
system: "You are a helpful assistant." ,
messages: await convertToModelMessages ( this . messages )
});
return result . toUIMessageStreamResponse ();
}
this.messages
The current conversation history, loaded from SQLite. This is an array of UIMessage objects from the AI SDK. Messages are automatically persisted after each interaction.
maxPersistedMessages
Cap the number of messages stored in SQLite. When the limit is exceeded, the oldest messages are deleted. This controls storage only — it does not affect what is sent to the LLM.
export class ChatAgent extends AIChatAgent {
maxPersistedMessages = 200 ;
}
To control what is sent to the model, use the AI SDK’s pruneMessages():
import { streamText , convertToModelMessages , pruneMessages } from "ai" ;
async onChatMessage () {
const workersai = createWorkersAI ({ binding: this . env . AI });
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: pruneMessages ({
messages: await convertToModelMessages ( this . messages ),
reasoning: "before-last-message" ,
toolCalls: "before-last-2-messages"
})
});
return result . toUIMessageStreamResponse ();
}
Controls whether AIChatAgent waits for MCP server connections to settle before calling onChatMessage. This ensures this.mcp.getAITools() returns the full set of tools, especially after Durable Object hibernation when connections are being restored in the background.
Value Behavior { timeout: 10_000 }Wait up to 10 seconds (default) { timeout: N }Wait up to N milliseconds trueWait indefinitely until all connections ready falseDo not wait (old behavior before 0.2.0)
export class ChatAgent extends AIChatAgent {
// Default — waits up to 10 seconds
// waitForMcpConnections = { timeout: 10_000 };
// Wait forever
waitForMcpConnections = true ;
// Disable waiting
waitForMcpConnections = false ;
}
Request Cancellation
When a user clicks “stop” in the chat UI, the client sends a CF_AGENT_CHAT_REQUEST_CANCEL message. The server propagates this to the abortSignal in options:
async onChatMessage ( _onFinish , options ) {
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages ),
abortSignal: options ?. abortSignal // Pass through for cancellation
});
return result . toUIMessageStreamResponse ();
}
If you do not pass abortSignal to streamText, the LLM call will continue running in the background even after the user cancels. Always forward it when possible.
Client API
useAgentChat
React hook that connects to an AIChatAgent over WebSocket. Wraps the AI SDK’s useChat with a native WebSocket transport.
import { useAgent } from "agents/react" ;
import { useAgentChat } from "@cloudflare/ai-chat/react" ;
function Chat () {
const agent = useAgent ({ agent: "ChatAgent" });
const {
messages ,
sendMessage ,
clearHistory ,
addToolOutput ,
addToolApprovalResponse ,
setMessages ,
status
} = useAgentChat ({ agent });
// ...
}
Options
Option Type Default Description agentReturnType<typeof useAgent>Required Agent connection from useAgent onToolCall({ toolCall, addToolOutput }) => void— Handle client-side tool execution autoContinueAfterToolResultbooleantrueAuto-continue conversation after client tool results and approvals resumebooleantrueEnable automatic stream resumption on reconnect bodyobject | () => object— Custom data sent with every request prepareSendMessagesRequest(options) => { body?, headers? }— Advanced per-request customization toolsRecord<string, AITool>— Dynamic client-defined tools for SDK/platform use cases. Schemas are sent to the server automatically getInitialMessages(options) => Promise<UIMessage[]> or null— Custom initial message loader. Set to null to skip the HTTP fetch entirely (useful when providing messages directly)
Return Values
Property Type Description messagesUIMessage[]Current conversation messages sendMessage(message) => voidSend a message clearHistory() => voidClear conversation (client and server) addToolOutput({ toolCallId, output }) => voidProvide output for a client-side tool addToolApprovalResponse({ id, approved }) => voidApprove or reject a tool requiring approval setMessages(messages | updater) => voidSet messages directly (syncs to server) statusstring"idle", "submitted", "streaming", or "error"
AIChatAgent supports three tool patterns, all using the AI SDK’s tool() function:
Pattern Where it runs When to use Server-side Server (automatic) API calls, database queries, computations Client-side Browser (via onToolCall) Geolocation, clipboard, camera, local storage Approval Server (after user approval) Payments, deletions, external actions
Tools with an execute function run automatically on the server:
import { streamText , convertToModelMessages , tool , stepCountIs } from "ai" ;
import { z } from "zod" ;
async onChatMessage () {
const workersai = createWorkersAI ({ binding: this . env . AI });
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages ),
tools: {
getWeather: tool ({
description: "Get weather for a city" ,
inputSchema: z . object ({ city: z . string () }),
execute : async ({ city }) => {
const data = await fetchWeather ( city );
return { temperature: data . temp , condition: data . condition };
}
})
},
stopWhen: stepCountIs ( 5 )
});
return result . toUIMessageStreamResponse ();
}
Define a tool on the server without execute, then handle it on the client with onToolCall. Use this for tools that need browser APIs:
tools : {
getLocation : tool ({
description: "Get the user's location from the browser" ,
inputSchema: z . object ({})
// No execute — the client handles it
});
}
When the LLM invokes getLocation, the stream pauses. The onToolCall callback fires, your code provides the output, and the conversation continues.
For SDKs and platforms where tools are defined dynamically by the embedding application at runtime, use the tools option on useAgentChat and createToolsFromClientSchemas() on the server:
import { createToolsFromClientSchemas } from "@cloudflare/ai-chat" ;
async onChatMessage ( _onFinish , options ) {
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages ),
tools: createToolsFromClientSchemas ( options ?. clientTools )
});
return result . toUIMessageStreamResponse ();
}
For most apps, server-side tools with tool() and onToolCall are simpler and provide full Zod type safety. Use dynamic client tools when the server does not know the tool surface at deploy time.
Use needsApproval for tools that require user confirmation before executing:
tools : {
processPayment : tool ({
description: "Process a payment" ,
inputSchema: z . object ({
amount: z . number (),
recipient: z . string ()
}),
needsApproval : async ({ amount }) => amount > 100 ,
execute : async ({ amount , recipient }) => charge ( amount , recipient )
});
}
Data Parts
Data parts let you attach typed JSON to messages alongside text — progress indicators, source citations, token usage, or any structured data your UI needs.
Writing Data Parts (Server)
Use createUIMessageStream with writer.write() to send data parts from the server:
import {
streamText ,
convertToModelMessages ,
createUIMessageStream ,
createUIMessageStreamResponse
} from "ai" ;
export class ChatAgent extends AIChatAgent {
async onChatMessage () {
const workersai = createWorkersAI ({ binding: this . env . AI });
const stream = createUIMessageStream ({
execute : async ({ writer }) => {
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages )
});
// Merge the LLM stream
writer . merge ( result . toUIMessageStream ());
// Write a data part — persisted to message.parts
writer . write ({
type: "data-sources" ,
id: "src-1" ,
data: { query: "agents" , status: "searching" , results: [] }
});
// Later: update the same part in-place (same type + id)
writer . write ({
type: "data-sources" ,
id: "src-1" ,
data: {
query: "agents" ,
status: "found" ,
results: [ "Agents SDK docs" , "Durable Objects guide" ]
}
});
}
});
return createUIMessageStreamResponse ({ stream });
}
}
Three Patterns
Pattern How Persisted? Use case Reconciliation Same type + id → updates in-place Yes Progressive state (searching → found) Append No id, or different id → appends Yes Log entries, multiple citations Transient transient: true → not added to message.partsNo Ephemeral status (thinking indicator)
Reading Data Parts (Client)
Non-transient data parts appear in message.parts. Use the UIMessage generic to type them:
import { useAgentChat } from "@cloudflare/ai-chat/react" ;
import type { UIMessage } from "ai" ;
type ChatMessage = UIMessage <
unknown ,
{
sources : { query : string ; status : string ; results : string [] };
usage : { model : string ; inputTokens : number ; outputTokens : number };
}
>;
const { messages } = useAgentChat < unknown , ChatMessage >({ agent });
// Typed access — no casts needed
for ( const msg of messages ) {
for ( const part of msg . parts ) {
if ( part . type === "data-sources" ) {
console . log ( part . data . results ); // string[]
}
}
}
Resumable Streaming
Streams automatically resume when a client disconnects and reconnects. No configuration is needed — it works out of the box.
When streaming is active:
All chunks are buffered in SQLite as they are generated
If the client disconnects, the server continues streaming and buffering
When the client reconnects, it receives all buffered chunks and resumes live streaming
Disable with resume: false:
const { messages } = useAgentChat ({ agent , resume: false });
For more details, see Resumable Streaming .
Storage Management
Row Size Protection
SQLite rows have a maximum size of 2 MB. When a message approaches this limit (for example, a tool returning a very large output), AIChatAgent automatically compacts the message:
Tool output compaction — Large tool outputs are replaced with an LLM-friendly summary that instructs the model to suggest re-running the tool
Text truncation — If the message is still too large after tool compaction, text parts are truncated with a note
Compacted messages include metadata.compactedToolOutputs so clients can detect and display this gracefully.
Controlling LLM Context vs Storage
Storage (maxPersistedMessages) and LLM context are independent:
Concern Control Scope How many messages SQLite stores maxPersistedMessagesPersistence What the model sees pruneMessages()LLM context Row size limits Automatic compaction Per-message
export class ChatAgent extends AIChatAgent {
maxPersistedMessages = 200 ; // Storage limit
async onChatMessage () {
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: pruneMessages ({
// LLM context limit
messages: await convertToModelMessages ( this . messages ),
reasoning: "before-last-message" ,
toolCalls: "before-last-2-messages"
})
});
return result . toUIMessageStreamResponse ();
}
}
Using Different AI Providers
AIChatAgent works with any AI SDK-compatible provider. The server code determines which model to use — the client does not need to change.
Workers AI (Cloudflare)
OpenAI
Anthropic
import { createWorkersAI } from "workers-ai-provider" ;
const workersai = createWorkersAI ({ binding: this . env . AI });
const result = streamText ({
model: workersai ( "@cf/zai-org/glm-4.7-flash" ),
messages: await convertToModelMessages ( this . messages )
});
Multi-Client Sync
When multiple clients connect to the same agent instance, messages are automatically broadcast to all connections. If one client sends a message, all other connected clients receive the updated message list.
Client A ──── sendMessage("Hello") ────▶ AIChatAgent
│
persist + stream
│
Client A ◀── CF_AGENT_USE_CHAT_RESPONSE ──────┤
Client B ◀── CF_AGENT_CHAT_MESSAGES ──────────┘
The originating client receives the streaming response. All other clients receive the final messages via a CF_AGENT_CHAT_MESSAGES broadcast.