| model | answer | ✓/✗ | conf | tokens | wall | mWh/tok | Wh | vs best |
|---|
What's measured (live runs): Tapo P110 wall-power at 1 Hz, baseline before each model (10 polls), ΔW × Δt → Wh, mWh/output_token, Traffic Light Confidence per the CR-028 CI model. Same protocol as the existing /llm endpoint.
What's estimated (showcase cards): Wh = wall_s × 25 W / 3600. The 25 W figure is a rough average power delta observed on GoS1 across Ollama 1B–20B models. Production P110 backfill of the three showcase cards is a follow-up (CR-048 phase 2).
Why "Wh per correct answer" is the headline: mWh/token rewards verbose models — a model that "thinks out loud" for 100 tokens to reach a 1-token answer looks more efficient per token than a model that answers in 1 token, even though it burned 100× the energy. The headline metric is the energy cost of correctness. mWh/token stays as a supporting column because it's the canonical inference-cost figure operators recognise.
6-model panel: Qwen3 1.7B (1.7B) · Qwen3 4B (4B) · Qwen3 8B (8B) · Mistral Nemo 12B (12B) · Phi-4 (14B) · GPT-OSS 20B (20B). All on GoS1 (Ryzen 9 7900 + NVIDIA GeForce RTX 5080, Ollama 0.20.2). Panel reflects the current llm.MODELS dict; the showcase tabs are frozen on the older 5-model probe and will be re-baselined when a fresh probe of the new panel runs.
Grading: Tolerant — substring match (case-insensitive) or leading-integer match. Same rule used by the 2026-05-26 probe to select these prompts.