In a streaming context Chat-style LLMs have limited direct use in streaming workflows (which lean on small specialised models) — this tab measures the expensive end of the spectrum as an upper bound, not the typical case. Language Lab AI paper ↗
ⓘ How to read AI energy in a streaming context (click to expand)
Framing from the Greening of Streaming Language Lab AI position paper (Jan 2026), “Distinguishing Impact from Innovation”:
AI is neither inherently sustainable nor unsustainable — type, size and deployment context decide net impact.
The type of AI matters enormously — small specialised CNNs (per-title encoding, scene classification, super-resolution) are orders of magnitude cheaper than general-purpose LLMs and diffusion models. Streaming mostly uses the former; these tabs measure the latter.
Data volume ≠ energy consumption.
OWL measures the energy AI adds — not the infrastructure energy AI avoids through better compression, caching or routing. Both halves are needed for net impact; OWL has the first.
Inference cost only — no amortised training cost.
Watch for rebound effects: efficiency gains can be offset by expanded use (more variations, more personalisation).
Run a language model on a fixed prompt and measure energy per token.
Models span small → large: TinyLlama 1.1B · Mistral 7B · Gemma 3 12B. CPU + ROCm GPU (via Ollama).
Cold mode unloads the model first; warm mode reuses a loaded model. Batch mode runs N inferences with a rest between.
Primary metric: mWh per output token · P110 polled at 1s intervals.
Scope: device layer only — no amortised training cost included.