⚿ Sign in
OWLOWL  ← Home ← /llm

LLM · Compare across models

Energy cost of correct answers. Device layer only (GoS1). Network and CPE excluded.
Prompt
Expected: · Panel: 5 models, mean of 3 reps each (2026-05-26 probe)
Cheapest correct answer
The size ≠ smarts finding
Same prompt, all models — ranked by energy of a correct answer
modelanswer✓/✗conf tokenswall mWh/tokWh vs best
🔒 Try your own prompt — Members only · Join GoS ↗

Try your own prompt

Runs your prompt across all 6 models sequentially with P110 power measurement. Each model gets its own clean baseline. Between models the runner actively polls power and waits for the system to return to within ±3 W of model #1's baseline (max 120 s cap) so heat from a verbose model can't contaminate the next reading. Total wall time depends on how hot each model leaves the GPU.
Prompt
Expected answer (used to grade ✓/✗)
Substring match, case-insensitive. For numeric answers, leading-integer match.
Methodology & scope

What's measured (live runs): Tapo P110 wall-power at 1 Hz, baseline before each model (10 polls), ΔW × Δt → Wh, mWh/output_token, Traffic Light Confidence per the CR-028 CI model. Same protocol as the existing /llm endpoint.

What's estimated (showcase cards): Wh = wall_s × 25 W / 3600. The 25 W figure is a rough average power delta observed on GoS1 across Ollama 1B–20B models. Production P110 backfill of the three showcase cards is a follow-up (CR-048 phase 2).

Why "Wh per correct answer" is the headline: mWh/token rewards verbose models — a model that "thinks out loud" for 100 tokens to reach a 1-token answer looks more efficient per token than a model that answers in 1 token, even though it burned 100× the energy. The headline metric is the energy cost of correctness. mWh/token stays as a supporting column because it's the canonical inference-cost figure operators recognise.

6-model panel: Qwen3 1.7B (1.7B) · Qwen3 4B (4B) · Qwen3 8B (8B) · Mistral Nemo 12B (12B) · Phi-4 (14B) · GPT-OSS 20B (20B). All on GoS1 (Ryzen 9 7900 + NVIDIA GeForce RTX 5080, Ollama 0.20.2). Panel reflects the current llm.MODELS dict; the showcase tabs are frozen on the older 5-model probe and will be re-baselined when a fresh probe of the new panel runs.

Grading: Tolerant — substring match (case-insensitive) or leading-integer match. Same rule used by the 2026-05-26 probe to select these prompts.

← /llm (single-model lab)
· CPU · GPU