How-to guide
How to Inject Ads Into a Streaming Response in OpenAI Assistants API
This guide shows how to inject ads into a streaming response in a OpenAI Assistants API app, the goal being to wrap the LLM token stream so contextual ads appear inside the assistant's reply. It builds on the Monetzly server SDK, so the approach is specific to how OpenAI Assistants API produces and streams responses.
Overview
This is the core of monetizing a OpenAI Assistants API app: instead of returning the raw model stream, you pass it through Monetzly, which injects contextual, labelled ads into the token stream and hands you back the enhanced stream to forward to the client.
You keep your existing OpenAI Assistants API model call untouched. Injection is additive — a wrapper around the stream you already produce, keyed on the session and the live prompt so the ad matches what the user is asking about.
OpenAI Assistants API apps are usually built for hosted chat assistants, file-search agents, and tool-calling assistants, so inject ads into a streaming response typically comes up while a user is mid-conversation, the moment where monetization has to be additive rather than disruptive.
How this works on OpenAI Assistants API
The Node OpenAI SDK streams run/completion events. Map each text delta event to { content: delta } and feed the generator into sdk.inject(). The Monetzly side is standard; the only framework detail to confirm is the exact delta event field for your chosen mode (Assistants run stream vs Chat Completions).
Because Monetzly's inject() accepts any async token stream, the OpenAI Assistants API side of this task is just mapping your output to it, no rewrite of your OpenAI Assistants API model call.
Steps
- Produce your OpenAI Assistants API response as a token stream as you already do.
- Pass that stream into the Monetzly injection call with the session id and prompt.
- Forward the enhanced tokens to the client (SSE, WebSocket, or your existing transport).
Integration snippet
import OpenAI from "openai";
import { MonetzlySDK } from "@monetzly/server-sdk";
const openai = new OpenAI();
const sdk = new MonetzlySDK({
apiKey: process.env.MONETZLY_API_KEY!,
serverAddress: process.env.MONETZLY_SERVER_ADDRESS!,
});
await sdk.connect();
const stream = await openai.chat.completions.create({
model: "gpt-4o-mini", messages, stream: true,
});
async function* asChunks() {
// TODO_VERIFY: for the Assistants run stream, read the text delta from the
// run event instead of choices[0].delta.content.
for await (const ev of stream) yield { content: ev.choices[0]?.delta?.content ?? "" };
}
for await (const t of sdk.inject(asChunks(), { prompt })) send(t.content ?? "");
await sdk.disconnect();Frequently asked questions
Start monetizing in about 5 minutes
Wrap your existing LLM response stream with the Monetzly SDK and earn on every session, no paywall required.