How-to guide
How to Inject Ads Into a Streaming Response in Gradio
This guide shows how to inject ads into a streaming response in a Gradio app, the goal being to wrap the LLM token stream so contextual ads appear inside the assistant's reply. It builds on the Monetzly gRPC service (no Python SDK exists for this stack), so the approach is specific to how Gradio produces and streams responses.
Overview
This is the core of monetizing a Gradio app: instead of returning the raw model stream, you pass it through Monetzly, which injects contextual, labelled ads into the token stream and hands you back the enhanced stream to forward to the client.
You keep your existing Gradio model call untouched. Injection is additive — a wrapper around the stream you already produce, keyed on the session and the live prompt so the ad matches what the user is asking about.
Gradio apps are usually built for model demos, chat interfaces, and shareable AI prototypes, so inject ads into a streaming response typically comes up while a user is mid-conversation, the moment where monetization has to be additive rather than disruptive.
How this works on Gradio
Gradio is Python. Drive the shipped tps_alter.proto gRPC service (or a Node sidecar) and yield injected tokens from a gr.ChatInterface generator fn for streaming output. Confirm the non-JS auth handshake before publishing.
Since Gradio is a Python stack with no official SDK, this task runs through the shipped tps_alter.proto gRPC service (or a Node sidecar); the auth handshake for a direct client is still TODO_VERIFY.
Steps
- Produce your Gradio response as a token stream as you already do.
- Pass that stream into the Monetzly injection call with the session id and prompt.
- Forward the enhanced tokens to the client (SSE, WebSocket, or your existing transport).
Integration snippet
# No Python SDK. Generate stubs from tps_alter.proto and drive ProcessStream
# (see the FastAPI page for the full bidi loop), then in a streaming chat fn:
#
# async def respond(message, history):
# out = ""
# async for token in inject(llm_tokens, message, session_id, api_key):
# out += token
# yield out
#
# TODO_VERIFY: auth handshake (API key placement) for a direct gRPC client.Frequently asked questions
Start monetizing in about 5 minutes
Wrap your existing LLM response stream with the Monetzly SDK and earn on every session, no paywall required.