SD-Turbo (~1B) · SDXL-Turbo (~3.5B) · 512×512 · CUDA fp16 on NVIDIA GeForce RTX 5080
In a streaming context AI-generated frames are the personalisation axis: per-viewer generated content breaks the cached-edge / multicast model and pushes delivery back to expensive unicast. Language Lab AI paper ↗
ⓘ How to read AI energy in a streaming context (click to expand)
Framing from the Greening of Streaming Language Lab AI position paper (Jan 2026), “Distinguishing Impact from Innovation”:
AI is neither inherently sustainable nor unsustainable — type, size and deployment context decide net impact.
The type of AI matters enormously — small specialised CNNs (per-title encoding, scene classification, super-resolution) are orders of magnitude cheaper than general-purpose LLMs and diffusion models. Streaming mostly uses the former; these tabs measure the latter.
Data volume ≠ energy consumption.
OWL measures the energy AI adds — not the infrastructure energy AI avoids through better compression, caching or routing. Both halves are needed for net impact; OWL has the first.
Inference cost only — no amortised training cost.
Watch for rebound effects: efficiency gains can be offset by expanded use (more variations, more personalisation).
Measures the wall-power cost of generating one AI image from text. SD-Turbo: CPU 8 steps (~12s) or GPU batch of 5 × 20 steps (~10s). Note: solo-mode GPU over-samples (native is 1–4 steps) to keep runtime above the P110 polling floor. SDXL-Turbo: GPU only, 4 steps (native), batch of 15 (~10s). Compare Models ⚡: both run at 4 steps (native for each), 512×512, same seed — SD-Turbo batch 30, SDXL-Turbo batch 15. Model size is the only variable.
Each run appends a random colour/mood modifier — live proof the image is generated, not replayed.