monetzly
My knees hurt after a run
Ice them, then try Recovery Gel
↑ contextual placement · sponsored
Weaving ads into the conversation…
monetzly
ProductHow it WorksPricing
Login
Join waitlist
  1. Home
  2. How-to
  3. Inject Ads Into a Streaming Response in OpenAI Assistants API

How-to guide

How to Inject Ads Into a Streaming Response in OpenAI Assistants API

This guide shows how to inject ads into a streaming response in a OpenAI Assistants API app, the goal being to wrap the LLM token stream so contextual ads appear inside the assistant's reply. It builds on the Monetzly server SDK, so the approach is specific to how OpenAI Assistants API produces and streams responses.

Overview

This is the core of monetizing a OpenAI Assistants API app: instead of returning the raw model stream, you pass it through Monetzly, which injects contextual, labelled ads into the token stream and hands you back the enhanced stream to forward to the client.

You keep your existing OpenAI Assistants API model call untouched. Injection is additive — a wrapper around the stream you already produce, keyed on the session and the live prompt so the ad matches what the user is asking about.

OpenAI Assistants API apps are usually built for hosted chat assistants, file-search agents, and tool-calling assistants, so inject ads into a streaming response typically comes up while a user is mid-conversation, the moment where monetization has to be additive rather than disruptive.

How this works on OpenAI Assistants API

The Node OpenAI SDK streams run/completion events. Map each text delta event to { content: delta } and feed the generator into sdk.inject(). The Monetzly side is standard; the only framework detail to confirm is the exact delta event field for your chosen mode (Assistants run stream vs Chat Completions).

Because Monetzly's inject() accepts any async token stream, the OpenAI Assistants API side of this task is just mapping your output to it, no rewrite of your OpenAI Assistants API model call.

Steps

  1. Produce your OpenAI Assistants API response as a token stream as you already do.
  2. Pass that stream into the Monetzly injection call with the session id and prompt.
  3. Forward the enhanced tokens to the client (SSE, WebSocket, or your existing transport).

Integration snippet

import OpenAI from "openai";
import { MonetzlySDK } from "@monetzly/server-sdk";

const openai = new OpenAI();
const sdk = new MonetzlySDK({
  apiKey: process.env.MONETZLY_API_KEY!,
  serverAddress: process.env.MONETZLY_SERVER_ADDRESS!,
});
await sdk.connect();

const stream = await openai.chat.completions.create({
  model: "gpt-4o-mini", messages, stream: true,
});

async function* asChunks() {
  // TODO_VERIFY: for the Assistants run stream, read the text delta from the
  // run event instead of choices[0].delta.content.
  for await (const ev of stream) yield { content: ev.choices[0]?.delta?.content ?? "" };
}

for await (const t of sdk.inject(asChunks(), { prompt })) send(t.content ?? "");
await sdk.disconnect();

Frequently asked questions

Start monetizing in about 5 minutes

Wrap your existing LLM response stream with the Monetzly SDK and earn on every session, no paywall required.

Get startedCompare models

Related

OpenAI Assistants API monetization guide
All how-to guides
Inject Ads Into a Streaming Response in LangChain
Inject Ads Into a Streaming Response in Vercel AI SDK (Next.js)
Inject Ads Into a Streaming Response in LlamaIndex.TS
Set Up the Monetzly SDK in OpenAI Assistants API
Pass Session Context to Ads in OpenAI Assistants API
Handle Ad-Injection Errors in OpenAI Assistants API
monetzly

Monetization for AI-native apps.

PRODUCT
OverviewHow it WorksUse CasesPricingFor Advertisers
RESOURCES
DocsGuidesFree ToolsChangelogStatus
COMPANY
AboutBlogContact
SOCIAL
Twitter / XLinkedIn
© 2026 Monetzly, Inc. All rights reserved.