Talk to us
Talk to us
menu

Best AI Agent SDKs With RAG in 2026: From Text Agents to Real-Time Voice

Best AI Agent SDKs With RAG in 2026: From Text Agents to Real-Time Voice

Best AI Agent SDKs With RAG in 2026: From Text Agents to Real-Time Voice

Retrieval-augmented generation (RAG) is what turns a chat model into an assistant that can answer questions about your product, your policies, and your data instead of guessing. That is why RAG support has become a baseline expectation when teams evaluate an AI agent SDK in 2026.

The problem is that most SDK comparisons stop at the text layer. They compare retrieval hooks, memory, and tool calling, then declare a winner. But the moment a user taps a microphone button, a different set of constraints takes over: speech-to-text streaming, turn detection, barge-in, and end-to-end latency measured in hundreds of milliseconds. An SDK that is excellent at orchestrating a text agent can be a poor foundation for a voice agent, and vice versa.

This guide does two things. First, it compares the SDKs most often cited for RAG agent development in 2026 — OpenAI Agents SDK, Vercel AI SDK, Claude Agent SDK, and Inworld — across the dimensions that actually decide an architecture. Second, it walks through the engineering work required to turn a text RAG agent into a real-time voice agent, and shows where ZEGOCLOUD’s AI Agent server fits as the real-time voice layer underneath your existing retrieval stack.

What a RAG-based AI agent SDK actually does

A RAG-based AI agent SDK is a toolkit for building agents that retrieve relevant knowledge from an external source and inject it into the model’s context before generating a response. The result is an answer grounded in specific, current documents rather than in the model’s training data alone.

In practice the SDK has to cover three responsibilities:

  • Retrieval — turning a user question into a query, hitting a vector store, keyword index, or hybrid search endpoint, and shaping the returned chunks into model-ready context.
  • Generation and orchestration — running the agent loop: calling the model, deciding whether more retrieval or another tool call is needed, and deciding when the answer is finished.
  • Tool calling and memory — letting the agent act on the outside world, and carrying conversation state forward so follow-up questions like “and for the enterprise tier?” resolve correctly.

The most important thing to understand before comparing SDKs is that most of them do not ship a retriever. They ship the loop, the tool abstraction, and the memory plumbing, and they expect you to bring chunking, embeddings, a vector store, and a ranking strategy. A small number ship a managed retrieval tool tied to their own vector store. What is rare is a single SDK that ships both a general retriever and a production-grade real-time audio pipeline, so in the set compared below retrieval and real-time audio are treated as separate concerns chosen separately.

The 2026 landscape at a glance

The table below compares the SDKs most commonly cited for RAG agent development, plus ZEGOCLOUD’s AI Agent server as the real-time voice layer. Where a vendor does not document a capability, the cell says so rather than inferring it from a broader product claim.

SDK Primary language / runtime RAG & retrieval Memory Tool calling Real-time voice Deployment model
OpenAI Agents SDK Python-first (production-ready) No native retriever, chunker, or vector-store abstraction. FileSearchTool retrieves from OpenAI Vector Stores; WebSearchTool handles web search. First-class Sessions: SQLite, Redis, SQLAlchemy, MongoDB, Dapr, OpenAI Conversations, plus a custom Session protocol. @function_tool decorator with Pydantic validation; tool namespaces, deferred loading, tool search, and hosted programmatic tool calling. Native. RealtimeAgent + RealtimeRunner over the OpenAI Realtime API, with semantic VAD turn detection and interrupt_response. Server-side library you run yourself; realtime sessions are server-managed.
Vercel AI SDK TypeScript (React, Next.js, Vue, Svelte, Node.js); Python in beta No native retriever or vector store. Provides embed, embedMany, and cosineSimilarity primitives, plus a RAG template built on language-model middleware. No first-party session store. Documented options are provider-defined tools (for example the Anthropic memory tool), third-party memory providers, or a custom tool. tool({ description, inputSchema, execute }); ToolLoopAgent is the documented recommended way to build an agent. Experimental. experimental_useRealtime, experimental_realtime.getToken(), and experimental_getRealtimeToolDefinitions(); token-based realtime over WebSockets with a provider dependency. Runs on your server and in the browser; optional AI Gateway normalizes providers.
Claude Agent SDK Python and TypeScript only No native RAG API. Retrieval is assembled by you and exposed to the agent as MCP or custom tools. CLAUDE.md context re-injected on every request, session transcripts, session forking, and a SessionStore adapter for cross-machine use; automatic compaction near the context limit. In-process MCP servers created with create_sdk_mcp_server; per-agent allowedTools; context-isolated subagents. Not documented. The only documented voice surface is CLI voice dictation, which is input-only transcription requiring a Claude.ai account. Local library (Python/TypeScript) or CLI subprocess; hosted Managed Agents is a separate product.
Inworld Hosted API with Python, Node.js, Unity, and Unreal clients Speech-first platform (Realtime TTS, Realtime STT, Realtime Router, Realtime API). No hosted retrieval API: the legacy Character Studio knowledge feature is retired, and the docs map that capability to Knowledge primitives you compose in the Runtime SDKs. The Router API’s documented grounding feature is web search. Server-extracted durable facts plus a rolling summary injected into the system prompt, with turn_interval, max_facts, and max_memory_length controls. Sessions are ephemeral. Functions registered mid-session via session.update, plus built-in web search. Native and central. Speech-to-speech over WebSocket or WebRTC, with semantic VAD, adjustable eagerness, and graceful barge-in. Hosted API only, with regional data residency options.
ZEGOCLOUD AI Agent (Server) Server REST API plus client SDKs for Android, iOS, Web, and Flutter Retrieval is owned by your custom LLM service. The agent server calls it over the OpenAI protocol and streams the grounded answer back as audio. Short-term memory provided externally or bound to In-app Chat (ZIM) history, with configurable sync mode and context window size. Tool use lives in your LLM service; the agent server adds proactive LLM and TTS invocation and client-side agent control. Native and the core product. Documented response latency as low as 1s, natural voice interruption detected in 500ms, and AI ANS, AI VAD, and AI AEC tuned for agent conversations. Server APIs plus SDK integration; the AI Agent service is enabled per project.

A note on LangGraph. LangGraph is a graph-based orchestration framework that appears in most 2026 roundups of this category, and ZEGOCLOUD’s own RAG guidance lists LangChain among the retrieval options you can plug in. Because its retrieval and audio behavior depend on the components you assemble around it rather than on a single documented SDK surface, it is discussed here as an orchestration choice rather than scored as a row in the capability table.

Reading the table: what each SDK actually optimizes for

The differences above are not quality differences. Each SDK is optimized for a different center of gravity, and that center of gravity is what determines whether it fits your product.

OpenAI Agents SDK: a complete text-agent loop with native realtime

The OpenAI Agents SDK is a Python-first framework built around agents, handoffs, guardrails, and sessions. For RAG work, its practical retrieval options are the built-in hosted tools — FileSearchTool against OpenAI Vector Stores and WebSearchTool — plus ordinary function tools that call whatever retrieval service you already run. There is no documented chunking pipeline, embedding abstraction, or bring-your-own-vector-store retriever interface, so hybrid search, reranking, and multi-tenant index routing remain your responsibility.

Its memory story is the broadest of the group. Sessions are a first-class abstraction with adapters for SQLite, Redis, SQLAlchemy, MongoDB, Dapr, and OpenAI-hosted conversations, and you can implement the Session protocol yourself. One documented constraint matters for RAG agents: you cannot combine a session with conversation_id or previous_response_id, so pick one state strategy rather than mixing them.

Voice is native rather than bolted on. The Realtime quickstart documents RealtimeAgent and RealtimeRunner over the OpenAI Realtime API, with PCM16 audio, audio.input.turn_detection semantic VAD, and interrupt_response for barge-in. Two limitations are worth planning around. First, the documented transports are server-side WebSocket and SIP/telephony — the Python SDK does not provide a browser WebRTC transport, so a browser-based voice agent needs a different path for the last hop. Second, the docs warn that automatic conversation compaction can block streaming, and recommend disabling it when you need fast turn-taking.

Vercel AI SDK: the TypeScript default, with realtime still experimental

The Vercel AI SDK is the natural choice if your application is already TypeScript and React or Next.js. It provides provider-agnostic primitives for model calls, structured output, and embeddings (embed, embedMany, cosineSimilarity), and its documented approach to building agents is ToolLoopAgent, which owns the tool-execution loop, context management, and stopping conditions.

For RAG, the SDK gives you embeddings and a reference RAG template but no retriever or vector store. For memory, the memory documentation explicitly offers three routes — provider-side memory tools, third-party memory providers such as Letta, Hindsight, or a MongoDB-backed store, or a custom tool — which means persistent agent memory is an integration you own.

Realtime voice is where the positioning becomes important. The realtime documentation exposes experimental_useRealtime in @ai-sdk/react and a server-side experimental_realtime.getToken(). The browser connects to the model provider using a short-lived token minted by your server, and continuous sessions may require an app-owned WebSocket relay. The API is prefixed experimental_, and it depends on the provider you route to. If your roadmap includes production voice, treat this as a promising but moving foundation rather than a settled one.

Claude Agent SDK: a strong agent loop, no documented real-time voice

The Claude Agent SDK packages the same tools, agent loop, and context management that power Claude Code into Python and TypeScript libraries. It is a strong choice when your agent needs deep file and tool autonomy, long-running tasks, or subagents that isolate context for parallel work. Custom tools are registered as in-process MCP servers with create_sdk_mcp_server, and subagents run as separate agent instances.

Two facts shape how it fits a RAG voice product. First, there is no native RAG API: retrieval is something you build and expose as tools. Second, there is no documented real-time voice capability. The only documented voice surface is voice dictation in Claude Code, which is input-only speech-to-text, streams audio to Anthropic’s servers for transcription, requires a Claude.ai account, and is explicitly unavailable when Claude Code is configured with an API key directly or through Bedrock, Google Cloud’s Agent Platform, or Microsoft Foundry. That is a developer convenience, not a speech-to-speech runtime.

Inworld: a speech-first platform, not a general RAG toolkit

Inworld is the outlier in this comparison because it is built for real-time conversation from the ground up. Its products are Realtime TTS, Realtime STT, Realtime Router, and the Realtime API, and its speech-to-speech endpoint follows the OpenAI Realtime protocol with extensions, available over WebSocket or WebRTC. It documents semantic VAD with adjustable eagerness, graceful barge-in, and back-channeling — the interaction details that make a voice agent feel natural.

Its knowledge story is narrower than its voice story. Inworld’s documented grounding path is Router web search, and the legacy hosted knowledge product has been retired, so teams with private corpora should expect to own retrieval themselves. Memory is server-extracted facts plus a rolling summary, bounded by max_facts and max_memory_length, and sessions are explicitly ephemeral — cross-session continuity is the developer’s responsibility. Inworld is also a hosted API only, priced per million characters of TTS and per hour of STT, with concurrency tiers that scale from 5 to 500 depending on plan. It is a strong choice when you want a managed speech stack; it is a weaker fit when your differentiator is your own retrieval and orchestration.

ZEGOCLOUD AI Agent: the real-time voice layer under your RAG stack

ZEGOCLOUD’s AI Agent takes the opposite approach to the SDKs above. It does not ask you to move your retrieval into a new framework. Instead, it runs the real-time conversation and calls your LLM service, which is where retrieval already lives.

The documented capabilities are the ones that matter for voice: response latency as low as 1s with full streaming, natural voice interruption detected within 500ms, recognition and interruption accuracy above 95% including double-talk and background music conditions, and dedicated AI ANS, AI VAD, and AI AEC processing for agent audio. It supports multiple LLM vendors and multiple TTS vendors, and it can be paired with a digital human avatar when the product needs a face.

Its boundaries are equally important. Retrieval is not part of the product — the RAG best-practices guide assumes you already have a retrieval service. Registration and instance creation are server-side REST calls that require a signed request, so you own that backend. And by default an account can run at most 10 agent instances concurrently; higher limits require contacting ZEGOCLOUD. For a first production deployment those constraints are usually fine, but they should be part of your capacity planning rather than a surprise at launch.

How to turn a RAG text agent into a real-time voice agent

The most common 2026 architecture is a text RAG agent that works well in a chat window, plus a product requirement to make it talk. The following sequence describes how that transformation works when the retrieval layer stays exactly where it is.

Step 1: Keep retrieval in your own LLM service and expose it over the OpenAI protocol

The cleanest pattern is to wrap your existing retrieval pipeline behind an OpenAI-compatible chat completions endpoint. Your service receives the user’s question, runs intent recognition and question enhancement if you need them, retrieves the relevant chunks, combines the fragments with the latest question, and returns a streaming response.

ZEGOCLOUD’s RAG best-practices guide documents exactly this flow. The user’s audio is published to a ZEGOCLOUD RTC room through the Express SDK; the AI Agent backend converts the audio to text and sends a chat completion request over the OpenAI protocol to your custom LLM service; your service performs retrieval and calls the model; the backend converts the streamed response back into audio and pushes it to the client. The guide’s worked example uses RAGFlow as the retrieval engine, with parameters such as similarity_threshold, vector_similarity_weight, and top_k, and it lists LangChain, LlamaIndex, LightRAG, and RAGFlow as supported retrieval approaches.

Two operational requirements are easy to miss. Your custom LLM service must be reachable on the public network — localhost and LAN addresses will not work — and during the two-week test period the placeholder credentials zego_test only work with the specific vendor endpoints and models listed in the API reference.

Step 2: Register the agent once, then create instances per room

ZEGOCLOUD separates the agent definition from the agent instance. You call RegisterAgent once to define the agent: an AgentId, a display name, and the LLM configuration pointing at your retrieval service. The LLM.Url must be compatible with the OpenAI Chat Completions API when Vendor is OpenAIChat (the default) or the OpenAI Responses API when it is OpenAIResponses. This is where you put the system prompt that instructs the agent to answer from retrieved knowledge and to say so plainly when the knowledge base has no answer.

You then call CreateAgentInstance for each conversation, passing the room ID, the real user’s ID and stream ID, and the agent’s own stream and user IDs. Each instance logs into the room, publishes its stream, and pulls the user’s stream. Three details cause most first-integration failures: AgentStreamId and AgentUserId must be unique across all instances or the later instance fails to publish or kicks the earlier one out; the agent’s user ID must never collide with a real user’s ID; and an instance is destroyed automatically if the user is absent longer than MaxIdleTime, which defaults to 120 seconds.

Step 3: Tune turn detection and interruption for natural conversation

Voice agents live or die on turn-taking. ZEGOCLOUD exposes this through the instance’s VAD configuration. TurnDetectConfig.SilenceSegmentation sets how long a silence ends a turn, with a documented range of 200–2000ms and a 500ms default; PauseInterval enables multi-sentence concatenation when it is set higher than the silence threshold. Interruption behavior is controlled by AdvancedConfig.InterruptMode: 0 interrupts the agent immediately when the user speaks, and 1 lets the agent finish its current sentence.

When false triggers are a problem, the voice interruption sensitivity guide exposes SensitiveConfig with a level from 0 to 3. Level 1 is the conservative setting at an energy threshold of 0.4 and a 100ms minimum speech duration; level 2 is the aggressive setting at 0.1 and 0ms; level 0 is the medium default; level 3 lets you set MinSpeechDur and EnergyThreshold yourself. The docs describe the trade-off directly: lower sensitivity avoids cutting the agent off on noise but misses genuine interruptions, while higher sensitivity catches real interruptions but reacts to background noise. If you are weighing this against a provider-side semantic VAD, note the difference in kind — ZEGOCLOUD exposes numeric energy and duration thresholds you tune per deployment, whereas semantic VAD models infer the end of a turn from meaning. That comparison is an architectural observation, not a benchmark.

Step 4: Decide where memory lives before you launch

RAG agents need two kinds of state: retrieved knowledge, which is stateless, and conversation history, which is not. ZEGOCLOUD’s short-term memory guide splits this into three stages — initial memory loaded at call start from ZIM chat history or an external context, in-call memory managed through the message list APIs, and post-call memory synchronized back when SyncMode is 0 and a valid ZIM.RobotId is configured.

The parameter that most affects quality and cost is WindowSize, the number of historical messages carried into context. It accepts 0–500 with a default of 20, and the documentation recommends 10–30 for most agents. Larger windows improve continuity but increase latency and token cost, and every model has a hard context ceiling that retrieval will compete with for space.

Step 5: Set a latency budget and measure each hop

Latency is the difference between a voice agent people use and one they abandon. A widely used rule of thumb for conversational turn-taking is a few hundred milliseconds from the end of the user’s speech to the first audio of the response, but that number is a design heuristic rather than a documented specification — no vendor in this comparison publishes it as a guaranteed figure. What matters more is knowing that it is the sum of your retrieval time, your model’s time to first token, the text-to-speech time to first byte, and the network path in both directions, and that you own most of those hops.

ZEGOCLOUD documents several levers on its side of that budget: full streaming from speech recognition through model output to synthesized audio, response latency as low as 1s for the complete agent turn, interruption detection within 500ms, and an RTC pull-streaming optimization in the v2.11.0 release that reduced voice latency by 50–100ms. The practical discipline is to instrument each hop separately — retrieval, LLM first token, TTS first byte, and transport — because the fix for a slow turn is almost never in the layer you assume.

Step 6: Choose the transport path deliberately

This is where SDK choice and voice quality intersect, and it is worth stating plainly. ZEGOCLOUD runs the conversation over its own real-time network with the Express SDK, so the agent and the user are both participants in an RTC room and the audio path is optimized end to end. Several agent SDKs take the opposite approach: the Python OpenAI Agents SDK documents server-side WebSocket and SIP transports but no browser WebRTC transport, and the Vercel AI SDK’s realtime APIs are experimental and connect the browser to the model provider through a token. If your product is a browser or mobile voice experience that must survive real-world networks, a purpose-built RTC layer is doing work that a WebSocket to a model provider does not.

One more capability worth knowing about: from v2.13.0 the Express SDK can control the agent instance directly through room signaling — triggering TTS or LLM output, interrupting the agent, and starting or stopping listening — without a server-side relay. For client-side interactions such as a “stop talking” button or a push-to-talk affordance, that removes a round trip through your backend.

Which SDK should you choose?

The decision is easier if you separate the two questions teams usually conflate: where does retrieval live, and where does the conversation live.

If your situation is… Start with Because
Python backend, text-first agent, OpenAI-hosted retrieval is acceptable, voice may come later OpenAI Agents SDK A production-ready agent loop with the widest set of documented session adapters, plus native realtime when you need it — subject to the documented transport limits.
TypeScript or React product, provider flexibility matters, voice is exploratory Vercel AI SDK Idiomatic for the stack, provider-agnostic, with ToolLoopAgent as the documented agent path. Treat realtime as experimental.
Long-running autonomous tasks, file and tool autonomy, or subagent-based decomposition Claude Agent SDK Purpose-built for deep agent autonomy. Plan to build retrieval and to source real-time voice elsewhere.
Managed speech stack with strong interaction quality, public or web-searchable knowledge Inworld Speech-to-speech with semantic VAD and graceful barge-in, without operating an audio pipeline yourself.
Private corpus you must own, and a real-time voice experience on browser or mobile Your existing RAG stack + ZEGOCLOUD AI Agent Retrieval stays in your service behind an OpenAI-compatible endpoint; the RTC layer handles streaming audio, turn detection, barge-in, and noise processing.

The pattern behind this table is that retrieval and real-time conversation are separable concerns, and the strongest 2026 architectures treat them that way. Choose the agent SDK that matches your language and autonomy requirements, keep retrieval in a service you control, and put a real-time media layer underneath when the product needs to speak.

Key takeaways

  • Most RAG agent SDKs do not include retrieval. They provide the agent loop, tool abstractions, and memory plumbing, and expect you to bring chunking, embeddings, a vector store, and ranking. The exception is hosted tools such as OpenAI’s FileSearchTool, which retrieves only from OpenAI Vector Stores.
  • Memory is a differentiator, not a checkbox. The OpenAI Agents SDK ships the widest set of session adapters; the Vercel AI SDK documents three integration routes instead of a first-party store; the Claude Agent SDK leans on CLAUDE.md, transcripts, and a SessionStore adapter; Inworld bounds memory with explicit fact and length limits and treats sessions as ephemeral.
  • Real-time voice support is uneven and often experimental. OpenAI’s Python SDK has native realtime but documents server-side WebSocket and SIP transports rather than browser WebRTC. Vercel’s realtime APIs are prefixed experimental_. The Claude Agent SDK’s only voice surface is input-only dictation. Inworld and ZEGOCLOUD are the two entries in this comparison built around speech.
  • You can keep your RAG stack and still get a voice agent. Exposing retrieval behind an OpenAI-compatible chat completions endpoint is the documented integration path for ZEGOCLOUD’s AI Agent server, so your vector store, reranker, and prompt strategy do not have to move.
  • Turn-taking parameters matter more than model choice for perceived quality. Silence segmentation, interruption sensitivity, and interaction mode determine whether an agent feels responsive or rude, and all three are tunable without changing your model.
  • Plan for the documented limits. Ten agent instances per account by default, a 120-second idle timeout, a 0–500 message history window, a 200–2000ms silence range, and public-network reachability for your LLM service are all specified constraints, not edge cases.

Frequently asked questions

What is a RAG-based AI agent SDK?

A RAG-based AI agent SDK is a toolkit for building agents that retrieve relevant documents from an external knowledge source and inject them into the model’s context before generating a response, so answers are grounded in your content rather than in training data alone. In practice the SDK provides the agent loop, tool calling, and conversation memory, while you supply the retrieval pipeline — chunking, embeddings, a vector store, and a ranking strategy. Only some SDKs ship a managed retrieval tool, and those are typically tied to that vendor’s own vector store.

Do RAG agent SDKs support real-time voice?

Some do, but support varies widely and is worth verifying per SDK rather than assuming. The OpenAI Agents SDK has native realtime agents over the OpenAI Realtime API, with documented server-side WebSocket and SIP transports; the Vercel AI SDK exposes realtime through experimental_ APIs that depend on your model provider; the Claude Agent SDK’s only documented voice surface is input-only dictation; Inworld provides speech-to-speech over WebSocket or WebRTC. For most text-oriented SDKs, real-time voice is a separate layer you add — a streaming speech-to-text and text-to-speech pipeline plus a low-latency transport with barge-in, which is exactly the layer ZEGOCLOUD’s AI Agent server provides.

Which AI agent SDKs are most referenced in 2026?

The three most often cited for RAG agent development are the OpenAI Agents SDK, the Vercel AI SDK, and the Claude Agent SDK, which between them cover Python, TypeScript, and deep agent autonomy. Inworld is commonly referenced as an alternative when the product is conversational rather than text-first, because its speech stack is the center of the platform. The right choice depends less on popularity than on your runtime, whether you need to own retrieval, and whether the product must speak in real time.

What latency actually matters for a real-time voice agent?

The number that determines whether a conversation feels natural is turn-taking latency: the time from when the user stops speaking to the first audio of the agent’s reply. A commonly cited rule of thumb is a few hundred milliseconds, and it is the sum of several hops — retrieval, the model’s time to first token, text-to-speech time to first byte, and the network path in both directions. Interruption responsiveness matters just as much, because an agent that cannot be cut off feels broken. ZEGOCLOUD documents response latency as low as 1s for a full agent turn and natural voice interruption detected within 500ms, and both are affected by how you configure silence segmentation and interruption sensitivity.

Conclusion

The RAG agent SDKs that lead in 2026 are strong at exactly what they were designed for: orchestrating model calls, tools, and memory. What they are not is a real-time media layer. Browser WebRTC transport is missing from OpenAI’s Python realtime path, Vercel’s realtime APIs are still experimental, and the Claude Agent SDK’s voice surface is limited to developer dictation.

That leaves a clean architectural split. Keep retrieval in the LLM service you already control, expose it over the OpenAI protocol, and put a purpose-built RTC layer underneath for the audio. ZEGOCLOUD’s AI Agent server is built for that role, with documented response latency as low as 1s, interruption detection within 500ms, and AI noise, voice-activity, and echo processing tuned for agent conversations — so your retrieval strategy stays yours and your voice experience does not have to be assembled from experimental parts.

If you want to see how the conversation layer performs before you commit, start with a proof of concept that connects your existing retrieval service to a registered agent and measures each hop of the latency budget. For background on why real-time delivery matters in this category, see Real-Time AI Voice Agents, and for the audio-processing capabilities that keep agent speech clean in noisy conditions, see AI Effects.

Read more: Best AI Agent SDK with RAG

Let’s Build APP Together

Start building with real-time video, voice & chat SDK for apps today!

Talk to us

Take your apps to the next level with our voice, video and chat APIs

Free Trial
  • 10,000 minutes for free
  • 4,000+ corporate clients
  • 3 Billion daily call minutes

Stay updated with us by signing up for our newsletter!

Don't miss out on important news and updates from ZEGOCLOUD!

* You may unsubscribe at any time using the unsubscribe link in the digest email. See our privacy policy for more information.