⚿ Sign in
OWLOWL  ← Home

LLM Inference Energy Test BETA

Greening of Streaming · OWL · GoS1
NEW · Compare 6 models on one prompt → energy per correct answer ↗  CR-048 hybrid: 3-prompt showcase + member “Try your own”.
In a streaming context
Chat-style LLMs have limited direct use in streaming workflows (which lean on small specialised models) — this tab measures the expensive end of the spectrum as an upper bound, not the typical case. Language Lab AI paper ↗
ⓘ How to read AI energy in a streaming context (click to expand)
Framing from the Greening of Streaming Language Lab AI position paper (Jan 2026), “Distinguishing Impact from Innovation”:
First time here? Try the Guided Tour →
ⓘ About this test (click to expand)
Run a language model on a fixed prompt and measure energy per token.
Models span small → large: TinyLlama 1.1B · Mistral 7B · Gemma 3 12B. CPU + ROCm GPU (via Ollama).
Cold mode unloads the model first; warm mode reuses a loaded model. Batch mode runs N inferences with a rest between.
Primary metric: mWh per output token · P110 polled at 1s intervals.
Scope: device layer only — no amortised training cost included.
Model

Qwen3 1.7B

1.7B · 1.4 GB

released 2025-04 · cutoff —

Qwen3 4B

4B · 2.5 GB

released 2025-04 · cutoff —

Qwen3 8B

8B · 5.2 GB

released 2025-04 · cutoff —

Mistral Nemo 12B

12B · 7.1 GB

released 2024-07 · cutoff —

Phi-4

14B · 9.1 GB

released 2024-12 · cutoff 2024-06

GPT-OSS 20B

20B · 13 GB

released 2025-08 · cutoff 2024-06

Task
🔒 Edit prompt — Members only · Join GoS ↗
✎ Edit prompt
Backend
Both: CPU then GPU with new baseline — full side-by-side comparison
Mode
Cold: unload model before baseline · Warm: model stays loaded
Repeats
Batch: load once, 10s rest between runs
🔒 Batch / compare — Members only · Join GoS ↗
· CPU · GPU